TeamAIOps
All insights
Buying guideJuly 29, 2026 · 6 min read

How to tell if an AI agent is ready for your main line

A practical readiness test you can run in an afternoon, before you forward the number that your revenue arrives through.

Putting software on your main line is not a software decision. It is a decision about the number your revenue arrives through, and it deserves the same scrutiny you would give a new hire who answers the phone unsupervised from day one.

Here is a test you can run in an afternoon. It is deliberately unkind, because your customers will be.

1. Call it the way a real customer would

Not from a quiet office on a good handset. From a mobile, outside, near traffic, at walking pace. Talk over it. Change your mind halfway through a sentence. Give a postcode with a letter that sounds like another letter.

Demos are recorded in studios. Your customers ring from vans. If the agent only performs under studio conditions, you have not seen the product.

2. Ask it something it cannot know

Ask for a price on a service you never gave it. There are two acceptable answers: it says it does not know and offers to take a message, or it transfers you. Any third answer — a confident, invented number — tells you what it will do to a real customer, and what you will be honouring when they hold you to it.

3. Time the handover

Ask for a human. Count the seconds until you are actually speaking to one, or until you are told clearly what will happen next. Then check whether the person who picked up already knows what you said, or whether you have to start again.

Making the caller repeat themselves is the moment goodwill is lost. It is also entirely avoidable, so its presence tells you something about how carefully the product was built.

4. Break it on purpose

  • Interrupt it mid-sentence. Does it stop, or talk over you?
  • Go silent for eight seconds. Does it wait, prompt, or hang up?
  • Change your mind: book an appointment, then cancel it in the same call.
  • Give a wrong number, correct it, and check which one is recorded.
  • Say something upsetting. It should escalate, not counsel you.

5. Read a week of transcripts, not a demo

A demo shows you the best call. A week of transcripts shows you the distribution. Ask any vendor whether you can run a shadow week — the agent answering a secondary number, or a diverted overflow line — before it touches the main one.

A vendor confident in their product will say yes. That answer is itself part of the test.

6. Check what happens when something fails

Ask what happens if their model provider has an outage mid-call. The honest answers are that it fails over to a different provider, or that calls route to a fallback number. The dishonest answer is that outages do not happen.

Then ask where you would see that it happened. If failures are not visible in a dashboard, you will find out from an angry customer instead.

A realistic rollout

  1. Week one: overflow only. It answers what you miss, nothing else.
  2. Week two: read every transcript daily and fix what it got wrong.
  3. Week three: out of hours, when the alternative is voicemail anyway.
  4. Week four: main line during business hours, with escalation live.

That order means every failure happens against a low baseline. An agent that mishandles an overflow call is competing with voicemail, not with you.

Before you shortlist anyone, read what AI agents get wrong and the questions worth asking a vendor.