Twelve questions to ask any AI phone agent vendor
The questions that separate a product built for production from a demo with a pricing page, and what a good answer sounds like.
Every vendor in this category demos well. Demos are chosen, quiet and short. These twelve questions are harder to stage, and the answers tell you whether a product has met production or only a slide deck.
For each one, what a good answer sounds like is included — because knowing the question is only half of it.
Reliability
- What happens when your model provider goes down mid-call? Good: it fails over to a second provider, and you can see in the dashboard which calls failed over. Bad: any answer implying this does not happen.
- What is your median time to first response, measured on real calls? Good: a number, with the method. Bad: 'instant', 'real-time', or 'sub-second' with nothing behind it.
- What happens to a call in progress when you deploy? Good: in-flight calls finish on the old version. Bad: a blank look.
Accuracy and honesty
- What does it do when asked something outside its knowledge? Good: a description of a refusal-and-escalate path you can test. Bad: 'it uses AI to figure it out'.
- How do I find out it got something wrong? Good: transcripts, outcome tagging, search, and alerts on escalations. Bad: 'you can listen to recordings' and nothing more.
- What does accuracy look like in week one versus week eight? Good: a candid answer that week one is worse and why. Bad: a claim that it is accurate immediately.
Control
- Can I stop it calling at certain hours, and is that enforced by the system or by my configuration discipline? Good: enforced server-side, per employee. Bad: 'you just schedule your campaigns carefully'.
- How do I suppress a number across everything, immediately? Good: one action, applies to every agent, takes effect on the next attempt.
- Can I see who changed what, and when? Good: an audit log, including changes made by their own staff. Bad: no answer.
Commercials
- What is the worst-case bill if something goes wrong at my end? Good: a hard spending stop that halts work rather than accruing charges. Bad: 'we'd work with you on it'.
- What do I own, and what happens to it if I leave? Good: you own recordings, transcripts and contacts, exportable, deleted on request. Bad: vagueness about the training question.
- Do you train models on my conversations? Good: no, and their providers are contractually barred from it too. Bad: 'only anonymised data', which is not a meaningful commitment for voice.
One more, and it is the real one
Can I run it on a secondary number for a week and read every transcript before it touches my main line?
A vendor who says yes is confident that a week of unedited reality supports the demo. A vendor who resists is telling you the demo was the product.
Next, read the readiness test you would run during that week, and what AI agents get wrong so you know what to look for in the transcripts.