Our AI agent can ring you back now.Get a call back
Blog

What to actually ask an AI voice agent vendor before you buy

Every demo works. These are the questions that separate a platform still working in month nine from one that demoed well — including the ones we find uncomfortable.

14 August 2026 · 6 min readBuying guideEvaluation

Every AI voice agent demo works. It is built by people who know exactly what to say, on a good network, in the language the platform is best at. The twelve questions below separate a platform still working in month nine from one that demoed well. They are deliberately vendor-neutral, and several are questions we find uncomfortable to answer — those are included precisely because you should ask them anyway.

For each one: why it matters, and what a bad answer sounds like. Ask them live, and ask for the artefact — a recording, a screen, an invoice line — rather than a statement.

Language and speech

1. Which languages do you support for speech recognition and for speech synthesis, separately?

These are two different models doing two different jobs, and a platform can easily understand a language it cannot speak. Ask for language codes, and ask which of them have been run on a real phone call rather than merely accepted by an API without returning an error.

A bad answer is “we support over 90 languages”, which is a claim about an upstream vendor's marketing page. A good one names the languages tested end to end on the public telephone network and admits the rest are untested.

2. Which vendor runs each stage, and can I change one without rebuilding the agent?

Speech recognition, the language model, speech synthesis and telephony are separable in some architectures and welded together in others. Which one you have determines whether you can replace a bad voice, react to a price increase, or meet a data residency requirement without migrating. It also caps your language list at whatever a single vendor supports.

A bad answer treats it as an implementation detail you should not worry about. If they will not name the vendors, you cannot evaluate the concentration risk you are inheriting.

3. What happens when I interrupt the agent mid-sentence?

Barge-in is the difference between a conversation and a voicemail menu. Interrupt the agent in the demo, hard, twice in a row. Then ask what the recording and transcript look like afterwards: an agent that stops speaking but logs the full sentence has a transcript that does not match what the caller heard, and every quality score derived from it is wrong.

A bad answer is “yes, we support interruptions”. Stopping the audio is the easy half.

Latency and reliability

4. What is your median and 95th-percentile response latency on a real phone call?

Measured from the end of the caller's speech to the first audio of the reply, on a call over the public telephone network rather than a browser demo on the vendor's own network. The median is the demo; the 95th percentile is the experience. Ask which stage dominates it, and ask them to show the per-stage breakdown for one specific call.

A bad answer is a single number with no percentile and no measurement point. “Sub-second” is not a measurement.

5. When one of your upstream vendors has an outage, what does my caller hear?

Every voice platform depends on a speech vendor, a model vendor and a carrier, and all three have outages. The question is whether the platform fails over to a second vendor mid-call, fails over between calls, or simply fails. This is a question we find uncomfortable to answer precisely, and so should everyone else: treat a confident, unhedged answer as a warning rather than a reassurance.

A bad answer is one that has never been tested. Ask for the date of the last incident and what callers actually experienced during it.

6. How many concurrent calls can I run, and how much production traffic do you handle today?

Concurrency is the constraint that bites during a campaign, an outage recovery or a product recall — the exact moments you bought the agent for. Ask what your ceiling is, how it is enforced, and what happens to call 51 when your limit is 50. The volume question is a fair proxy for how many failure modes the platform has already met.

A bad answer is “unlimited”. Nothing is unlimited. The honest version is a number, a queueing behaviour, and how quickly it can be raised.

Money

7. What is the billing increment, when does the meter start, and what do I pay for a call that reached voicemail?

Whole-minute rounding on short outbound calls is a large, quiet premium, and a platform without answering-machine detection bills you for the voicemail and spends model tokens talking to it. Ask for the increment, the minimum billable call, whether the clock starts at dial or at answer, and what a detected machine costs.

A bad answer is “usage-based pricing”. That phrase covers per-second and per-rounded-up-minute equally well, and the gap between them is tens of percent.

8. What does one call cost you, itemised by stage?

Recognition, model, synthesis and telephony are four separate bills in four separate units — audio seconds, tokens, characters, minutes. A platform that cannot itemise them for a specific call either is not measuring them or does not want you to see them. Their margin does not have to be zero. You do need to know which stage to change when the bill grows.

A bad answer converts everything into one blended per-minute figure and stops there.

What happens when it is not enough

9. What happens when the agent cannot answer the question?

There are three real designs: the agent gives up and ends the call, it transfers the caller into a queue, or it consults someone and resumes the same conversation. Ask which. Then ask whether the caller's verified identity and retrieved records survive it, because everything the caller has to repeat is value the agent already destroyed.

A bad answer is “it hands off to a human” with no detail about what the human receives at the moment of handoff.

10. Play me a call your agent handled badly.

This is the most informative question on the list and the one most likely to be deflected. Any platform with real production traffic has bad calls. A vendor who can pull one up, explain what went wrong and show what changed afterwards is telling you they have a quality process. One who cannot either has no production traffic or does not review it.

A bad answer is a highlight reel, or a hypothetical about what would happen if a call went wrong.

Data, and the day you leave

11. Where does my call data live, how long is it kept, who can read it, and can it be redacted?

Call recordings and transcripts contain names, account numbers, health details and card fragments, depending on your business. Ask for the storage region, the retention period and whether it is configurable, whether vendor staff can access recordings and under what controls, whether personal data can be redacted, and whether your audio is used to train anybody's models. If you are regulated, ask in writing.

A bad answer is “it is encrypted”. Encryption at rest is table stakes and answers none of those questions.

12. If I leave in a year, what do I take with me?

Ask for the export: transcripts, recordings, prompts, agent configuration, tool definitions, scorecards. Ask whether your phone numbers sit in your carrier account or theirs, because a number you cannot port is the strongest lock-in in this category. Ask whether the platform can be self-hosted and whether you can bring your own vendor credentials, which turns an exit into a redirection.

A bad answer is a CSV of transcripts. That is the least valuable thing you built with them.

One more, if you have time

How does the agent get better after launch, and whose job is that? An agent that is a static prompt is at its best on day one and decays as the business changes around it. Ask whether every call is transcribed and scored against a rubric you control, whether the platform surfaces what worked on the calls that succeeded, and whether improving the agent is your team's work, their services team's work, or automatic. Then ask to see the rubric: quality assurance with no visible scorecard is a claim about a feature nobody has looked at.

Put an agent on your line this week

A walkthrough on a real call — voice pipeline, numbers, workflows, QA and human handoff working together.