Most AI voice agent demos only have to work once. Cekura's live October 2026 benchmark runs the same call scenario three times per platform and only counts a pass if all three succeed — and even Retell, the best-scoring platform tested, clears that bar just 75.61% of the time. **The real reliability question for an AI phone system isn't whether it can complete a call once; it's whether it completes the same call the same way every time.**
What Is Cekura's 'Pass³' Voice AI Benchmark?
Cekura, a voice-AI testing and monitoring company, runs a public, continuously updated benchmark site that tests leading voice AI platforms across four categories: Agent Workflow (full voice-agent platforms), Speech-to-Speech (realtime models), Voice Quality (production call sound and interruption handling), and Speech-to-Text (transcription accuracy). The Agent Workflow category — the one that matters most for a business buying an AI receptionist — tested eight platforms (LiveKit, Pipecat, Vapi, Retell, ElevenLabs, and others) across 82 real scenarios placed over actual telephony, not a text-only simulation (Cekura, 2026). Cekura's own definition of its headline metric is specific: "Pass³" is "the share of scenarios where all three retained test calls passed" — meaning each scenario was run three separate times, and a platform only gets credit if it succeeded on all three attempts, not just one (Cekura, 2026).
| Platform (Agent Workflow category) | Pass³ repeatability score | What it means |
|---|---|---|
| Retell | 75.61% | Top scorer — still fails all-three-correct about 1 in 4 scenarios |
| LiveKit | 70.73% | Second — roughly 3 in 10 scenarios aren't repeatable |
| ElevenLabs | 69.51% | Third of eight platforms tested |
That's the top three of eight platforms tested in Cekura's October 1, 2026 update, and no platform in the test cleared 80% (Cekura, 2026). A business comparing vendor one-pagers, where a 95%+ accuracy claim is standard marketing language, is looking at a different number entirely than what this benchmark measures.
Why Does Testing a Call Once Hide the Real Failure Rate?
A single successful demo call tells you almost nothing about whether the next call will go the same way, because these systems are probabilistic, not deterministic — the same input can produce a slightly different response each time. A 2026 Princeton-affiliated research paper on AI agent reliability, "Towards a Science of AI Agent Reliability" (Rabanser, Kapoor, Kirgis, Liu, Utpala, and Narayanan), makes the underlying point directly: judging an agent on a single run (what researchers call "pass@1") produces an artificially optimistic picture, because an agent can succeed occasionally while still failing on repeated attempts of the identical task — exactly the gap Cekura's three-run methodology is built to expose (arXiv, 2026). Run the arithmetic on Retell's own number: a platform that independently succeeds on a given call type around 91% of the time will, by chance alone, pass three back-to-back attempts only about 75% of the time (0.91 cubed) — which is almost exactly where Cekura's benchmark lands. A number that looks excellent framed as a single-call success rate looks a lot less reassuring once you ask for it three times in a row.
This isn't unique to voice AI. Gartner surveyed 5,728 customers and found that while 73% had used a self-service channel, only 14% fully resolved their issue that way — and even for issues customers called "very simple," only 36% resolved completely (CX Today, 2024). Vendor-reported containment or completion numbers and what a customer actually experiences on a retry have a documented habit of diverging across customer-service automation generally, not just on the phone.
When Is a One-Shot Demo Actually Good Enough?
Honestly: for a lot of calls, it is. If an AI receptionist's only job on a given call is to state a fact that doesn't change — business hours, an address, whether a service is offered — there's no multi-step state to get wrong, and a system that nails that once will nail it almost every time. Repeatability matters most exactly where Cekura's scenarios are weighted: appointment booking, data capture, and scenarios with branching logic, where a name, date, or address has to be heard, confirmed, and written down correctly in sequence. A business whose AI phone system only fields simple info requests doesn't need to interrogate a vendor's Pass³ score. A business routing bookings, service calls, or payment-adjacent requests through it does.
What Should a Business Actually Ask an AI Phone Vendor Before Buying?
Ask to see the same call scenario run back-to-back, more than once, not a highlight reel of one clean demo call. A Whitefish ski-shop or a Kalispell contractor fielding a booking call during a busy stretch doesn't get to pick which attempt the AI gets right — the caller only calls once, and if the address or callback number comes out wrong on that one call, there's no second take. That's the practical version of what Cekura's benchmark is measuring: not "can this system do the job," which nearly every vendor can demonstrate, but "does it do the job the same way every single time," which is the actual bar for trusting it with a call that matters.