Cekura Bench
Cekura Bench provides an open benchmark evaluating realtime speech-to-speech AI models across 82 simulated phone call scenarios spanning clinic scheduling and Medicare intake.
Cekura Bench provides an open benchmark evaluating realtime speech-to-speech AI models across 82 simulated phone call scenarios spanning clinic scheduling and Medicare intake.
What the product does and how it is positioned
Cekura Bench is a benchmark that evaluates realtime speech-to-speech voice models using standardized phone call scenarios to measure conversational reliability, latency, and data accuracy.
The platform tests models across 82 distinct scenarios three times each using an open-source Pipecat agent framework, publishing full call transcripts, tool execution records, and performance metrics.
Source-supported ways to use the product
Teams compare realtime speech models on consistency, response times, and data accuracy before implementing voice agents.
Developers assess native speech-to-speech models against a standard cascaded pipeline of speech-to-text, language model, and speech synthesis.
The benchmark standardizes test conditions by running each speech-to-speech model through the same open-source Pipecat voice agent framework. Prompts, integration tools, and connections remain identical across runs, isolating model performance as the sole variable under test. An unranked cascade baseline combining Flux, GPT-4.1, and ElevenLabs Flash serves as an architectural reference.
Evaluation encompasses 82 test scenarios conducted over three live runs, totaling 2,952 calls across clinic appointment handling and Medicare qualification. A call passes only when the agent achieves the required outcome and correctly saves all expected data fields, with consistency measured by the proportion of scenarios that succeed across every run.
Checks to run with your own material and workflow
What was checked and when
Answers based on the source-checked product record
The benchmark evaluates 82 scenarios across two suites: a 59-scenario clinic receptionist suite for appointment booking and a 23-scenario Medicare intake suite covering consent, qualification, and routing.
Reliability is measured using a pass-cubed metric, which calculates the percentage of scenarios that successfully pass across three consecutive live-call runs.
All models run the same open-source Pipecat agent configuration with identical prompts, tool definitions, and connection setups to isolate model performance.
Yes, an unranked reference cascade combining Flux for speech-to-text, GPT-4.1 for reasoning, and ElevenLabs Flash for text-to-speech is included for comparative context.
Each model entry links directly to Cekura reports containing complete transcripts, tool call executions, and outcome scores for verification.