How fast should an AI voice agent answer a call?
Latency is the difference between a receptionist and an obvious robot. Here are the numbers that matter and how to test them yourself.
There are two speeds people confuse. One is how quickly the phone gets picked up. The other is how quickly the agent replies once the caller stops talking. The second one is what makes a call feel human or makes it feel like a kiosk.
Pickup speed: one ring, no exceptions
There is no technical reason for an AI agent to ring more than once. If a vendor's system takes three or four rings, something is queuing. Callers start abandoning around the fourth ring, and by the sixth a meaningful share are gone.
Response latency: the sub-second threshold
In natural English conversation the gap between one person finishing and the next starting is roughly 200 milliseconds. That is the target. Below about 500ms a caller registers the reply as normal conversation. Between 500ms and one second it reads as thoughtful. Past 1.5 seconds it reads as broken, and callers start repeating themselves — which then confuses the agent and compounds the problem.
| Response delay | How it feels | Verdict |
|---|---|---|
| Under 300ms | Indistinguishable from a person | Excellent |
| 300 – 600ms | Natural, slightly considered | Good |
| 600ms – 1.2s | Noticeably slow | Acceptable at the edge |
| Over 1.5s | Caller repeats themselves | Unusable |
Interruption handling matters more than raw speed
Real callers talk over the receptionist constantly. They blurt out the address mid-sentence. A fast agent that cannot be interrupted is worse than a slightly slower one that stops instantly when the caller starts speaking. Ask any vendor to demonstrate barge-in, and interrupt the demo yourself — mid-word, not politely at the end of a sentence.
Call the demo line and, while the agent is greeting you, immediately say your address. If it keeps talking over you or loses the address, the latency numbers on the marketing page do not matter.
What actually causes lag
- Telephony hop — the carrier path before audio reaches the model at all
- Speech recognition — turning your words into text
- The language model producing a reply
- Speech synthesis — turning that reply back into audio
- A calendar or CRM lookup mid-conversation, which is the sneaky one
That fifth item is where most real-world stalls happen. When the agent checks live availability, a slow calendar API can add a full second of dead air. Well-built agents cover it conversationally — 'let me check that for you' — instead of going silent.
How to test a vendor properly
- Call the demo line from a mobile on cellular, not from a laptop on office wifi
- Interrupt mid-sentence at least three times
- Give a deliberately messy answer to a simple question
- Ask something outside the script and see whether it admits ignorance
- Call again at 11pm and confirm nothing changed
Frequently asked questions
Is low latency worth paying more for?+
Up to a point. Going from 1.5 seconds to 600ms transforms call quality. Going from 400ms to 250ms is largely invisible to callers, so do not pay a premium for the last hundred milliseconds.
Does latency get worse with more simultaneous calls?+
On a well-architected platform, no — capacity scales horizontally. Ask specifically what happens at your peak concurrency, because that is where weak systems degrade.
Why does the agent sometimes pause before booking?+
It is reading live calendar availability. A good agent narrates that pause so the caller knows something is happening rather than sitting in silence.
See it answer a real call in 30 seconds
Test the AI receptionist live, or book a 20-min strategy walkthrough.
Related articles
Get the weekly playbook
One field-tested tip for AI-voice teams, every Monday.