HomeInsightsThe Best AI Voice Agent…

The Best AI Voice Agent Still Fails the Same Call 1 in 4 Times

Vendor demos only have to work once. A live benchmark that retests the same call three times tells a very different story about what an AI phone system actually does on a bad day.

By Alex RiveraPublished October 4, 2026

Most AI voice agent demos only have to work once. Cekura's live October 2026 benchmark runs the same call scenario three times per platform and only counts a pass if all three succeed — and even Retell, the best-scoring platform tested, clears that bar just 75.61% of the time. **The real reliability question for an AI phone system isn't whether it can complete a call once; it's whether it completes the same call the same way every time.**

What Is Cekura's 'Pass³' Voice AI Benchmark?

Cekura, a voice-AI testing and monitoring company, runs a public, continuously updated benchmark site that tests leading voice AI platforms across four categories: Agent Workflow (full voice-agent platforms), Speech-to-Speech (realtime models), Voice Quality (production call sound and interruption handling), and Speech-to-Text (transcription accuracy). The Agent Workflow category — the one that matters most for a business buying an AI receptionist — tested eight platforms (LiveKit, Pipecat, Vapi, Retell, ElevenLabs, and others) across 82 real scenarios placed over actual telephony, not a text-only simulation (Cekura, 2026). Cekura's own definition of its headline metric is specific: "Pass³" is "the share of scenarios where all three retained test calls passed" — meaning each scenario was run three separate times, and a platform only gets credit if it succeeded on all three attempts, not just one (Cekura, 2026).

Platform (Agent Workflow category)Pass³ repeatability scoreWhat it means
Retell75.61%Top scorer — still fails all-three-correct about 1 in 4 scenarios
LiveKit70.73%Second — roughly 3 in 10 scenarios aren't repeatable
ElevenLabs69.51%Third of eight platforms tested

That's the top three of eight platforms tested in Cekura's October 1, 2026 update, and no platform in the test cleared 80% (Cekura, 2026). A business comparing vendor one-pagers, where a 95%+ accuracy claim is standard marketing language, is looking at a different number entirely than what this benchmark measures.

Why Does Testing a Call Once Hide the Real Failure Rate?

A single successful demo call tells you almost nothing about whether the next call will go the same way, because these systems are probabilistic, not deterministic — the same input can produce a slightly different response each time. A 2026 Princeton-affiliated research paper on AI agent reliability, "Towards a Science of AI Agent Reliability" (Rabanser, Kapoor, Kirgis, Liu, Utpala, and Narayanan), makes the underlying point directly: judging an agent on a single run (what researchers call "pass@1") produces an artificially optimistic picture, because an agent can succeed occasionally while still failing on repeated attempts of the identical task — exactly the gap Cekura's three-run methodology is built to expose (arXiv, 2026). Run the arithmetic on Retell's own number: a platform that independently succeeds on a given call type around 91% of the time will, by chance alone, pass three back-to-back attempts only about 75% of the time (0.91 cubed) — which is almost exactly where Cekura's benchmark lands. A number that looks excellent framed as a single-call success rate looks a lot less reassuring once you ask for it three times in a row.

This isn't unique to voice AI. Gartner surveyed 5,728 customers and found that while 73% had used a self-service channel, only 14% fully resolved their issue that way — and even for issues customers called "very simple," only 36% resolved completely (CX Today, 2024). Vendor-reported containment or completion numbers and what a customer actually experiences on a retry have a documented habit of diverging across customer-service automation generally, not just on the phone.

When Is a One-Shot Demo Actually Good Enough?

Honestly: for a lot of calls, it is. If an AI receptionist's only job on a given call is to state a fact that doesn't change — business hours, an address, whether a service is offered — there's no multi-step state to get wrong, and a system that nails that once will nail it almost every time. Repeatability matters most exactly where Cekura's scenarios are weighted: appointment booking, data capture, and scenarios with branching logic, where a name, date, or address has to be heard, confirmed, and written down correctly in sequence. A business whose AI phone system only fields simple info requests doesn't need to interrogate a vendor's Pass³ score. A business routing bookings, service calls, or payment-adjacent requests through it does.

What Should a Business Actually Ask an AI Phone Vendor Before Buying?

Ask to see the same call scenario run back-to-back, more than once, not a highlight reel of one clean demo call. A Whitefish ski-shop or a Kalispell contractor fielding a booking call during a busy stretch doesn't get to pick which attempt the AI gets right — the caller only calls once, and if the address or callback number comes out wrong on that one call, there's no second take. That's the practical version of what Cekura's benchmark is measuring: not "can this system do the job," which nearly every vendor can demonstrate, but "does it do the job the same way every single time," which is the actual bar for trusting it with a call that matters.

Skyline scopes AI receptionists around exactly this distinction — which calls can run fully automated and which need a tighter script, a confirmation step, or a live handoff. Book a free AI audit to see where that line falls for your business.

Sources

  1. Cekura (2026)
  2. CX Today (2024)
  3. arXiv (2026)
[ 05 ]Questions

Related questions

Clear answers to the questions operators ask most. Still not sure if AI fits your business? Talk to us — no pitch, just a straight read on where it pays off.

What is Cekura's Pass³ score for AI voice agents?

Pass³ is Cekura's benchmark metric measuring the share of test scenarios where a voice AI platform succeeded on all three repeated attempts, not just one. In its Agent Workflow category — 82 scenarios across 8 platforms, updated October 1, 2026 — Retell scored highest at 75.61%, followed by LiveKit at 70.73% and ElevenLabs at 69.51% (Cekura, 2026).

Is a 95% accuracy claim from an AI receptionist vendor realistic?

It depends what's being measured. A 95%+ figure is usually a single-call or per-word accuracy claim, measured once. Cekura's repeated-run benchmark, which only counts a pass if the same scenario succeeds three times in a row, puts even the top platform at 75.61% (Cekura, 2026) — a reminder to ask a vendor exactly what their number measures before comparing it to another vendor's.

Does a more expensive AI receptionist mean it's more reliable?

Not necessarily. Cekura's Agent Workflow rankings mix commercial platforms with open-source frameworks like LiveKit and Pipecat, and price didn't track cleanly with the repeatability scores (Cekura, 2026). Ask a vendor for repeat-test evidence on your actual call types rather than assuming a higher price tag buys more consistency.

Which AI voice agent platform scored highest for reliability in 2026?

Retell scored highest in Cekura's Agent Workflow repeatability test at 75.61%, out of eight platforms tested as of the benchmark's October 1, 2026 update (Cekura, 2026). That's one methodology on one set of scenarios — worth confirming against a vendor's performance on your own call types rather than treating it as a permanent ranking.

Related questions

Related services

More from Insights

No Pitch, No Obligation

See exactly where AI pays off in your business

Book a free AI audit. We'll map your biggest leak — missed calls, slow follow-up, manual admin — and show you the system that fixes it. No pitch, no obligation.

Free · no obligation~30 minutesYou own everything