An AI receptionist that sounds robotic or talks over callers usually isn't a voice problem — it's a timing problem. Human conversation runs on turn-taking gaps of roughly 0 to 300 milliseconds (PNAS, 2009). Many production AI voice systems take 600 milliseconds to over a second to respond, and that gap is what a caller actually notices.
It's not the voice. It's the silence before it.
Text-to-speech has gotten good enough that voice quality alone rarely gives an AI receptionist away anymore. What still gives it away is timing: how long the system takes to register that a caller stopped talking, work out what they said, and start responding. Researchers who measured real conversations across ten languages found human turn-taking gaps clustering tightly around zero, with a cross-linguistic median of about 100 milliseconds (PNAS, 2009). That's the bar a caller is unconsciously measuring against, whether they've ever thought about it or not.
The benchmark numbers vendors don't put on the same page
Telnyx's own 2026 comparison of voice-AI platforms — a vendor in this market itself, worth noting up front — cites independent production testing that shows wide gaps in real, full-turn response time: measured from the end of a caller's sentence to the start of the AI's reply, across hundreds of live calls. Tested Media's March 2026 study put Retell at 680 milliseconds median (920 ms at the 95th percentile), Vapi at 720 ms median (1,050 ms p95), and Bland at 850 ms median (1,180 ms p95) across 500 production calls per platform. A separate fixed-stack benchmark from Cekura measured ElevenLabs' full conversational turn at 1.73 seconds median (Telnyx, 2026). None of that is a bad number by current industry standards — it's just a different number than what a sales demo tends to showcase.
Ask this before you sign: what exactly is being measured?
Vendor marketing pages tend to lead with whichever number looks best, and the two most common ones measure different things. Telnyx also publishes a carrier-network test showing 71 milliseconds median and 118 ms at the 95th percentile — the fastest of three carriers tested in a June 2026 study (Telnyx, 2026). That's a real result, but it only covers the network hop between phone systems. It says nothing about how long the AI itself takes to understand a caller and generate a reply. A vendor quoting a sub-100-millisecond number without saying what's included is quoting the easy part.
| Platform | What was measured | Result |
|---|---|---|
| Retell | Full turn, end of caller speech to start of AI reply (500 live calls) | 680 ms median / 920 ms p95 |
| Vapi | Same full-turn measurement (500 live calls) | 720 ms median / 1,050 ms p95 |
| Bland | Same full-turn measurement (500 live calls) | 850 ms median / 1,180 ms p95 |
| ElevenLabs | Full turn on a fixed third-party tech stack | 1.73 sec median |
| Telnyx / Twilio / Vonage | Carrier network leg only — not a full AI response | 71–94 ms median |
Why this matters more on a Montana phone line
Every one of those benchmark numbers assumes a clean network path. Montana doesn't reliably offer one. The state ranked 50th of 51 in a broadband analysis published in August 2026, with 83% wired or fixed-wireless coverage and an overall grade of F, a gap driven largely by mountainous terrain and long distances between towers (BroadbandNow, 2026). A caller on a weak rural connection, or a business running its own AI receptionist over the same rural line, adds variability on top of whatever latency the vendor already has. That's exactly the situation where a vendor's best-case demo number matters least — and where how the system behaves on a bad connection matters most.
Three questions that actually predict call quality
- What's your full-turn latency — end of caller speech to start of reply — not just network latency? Ask for the number and how it was measured.
- What happens on a degraded connection: does the system pause gracefully, or does it produce a false interruption?
- Can I hear it handle a real, unscripted call before I sign — not just a rehearsed demo script?