Cekura Bench favicon

Cekura Bench

Cekura Bench provides an open benchmark evaluating realtime speech-to-speech AI models across 82 simulated phone call scenarios spanning clinic scheduling and Medicare intake.

Research AssistantEvaluating native speech-to-speech…Reviewing public call reports…Measuring operational metricsStandardized Voice Evaluation Suites
Cekura Bench product interface screenshot
Estimated monthly visits
66K
Data period:
Listed on AIToolly

What Is Cekura Bench? Product Overview

What the product does and how it is positioned

Cekura Bench is a benchmark that evaluates realtime speech-to-speech voice models using standardized phone call scenarios to measure conversational reliability, latency, and data accuracy.

The platform tests models across 82 distinct scenarios three times each using an open-source Pipecat agent framework, publishing full call transcripts, tool execution records, and performance metrics.

What Can You Use Cekura Bench For?

Source-supported ways to use the product

Voice Model Evaluation

Teams compare realtime speech models on consistency, response times, and data accuracy before implementing voice agents.

Architecture Benchmarking

Developers assess native speech-to-speech models against a standard cascaded pipeline of speech-to-text, language model, and speech synthesis.

Benchmark Methodology and Evaluation Metrics

The benchmark standardizes test conditions by running each speech-to-speech model through the same open-source Pipecat voice agent framework. Prompts, integration tools, and connections remain identical across runs, isolating model performance as the sole variable under test. An unranked cascade baseline combining Flux, GPT-4.1, and ElevenLabs Flash serves as an architectural reference.

Evaluation encompasses 82 test scenarios conducted over three live runs, totaling 2,952 calls across clinic appointment handling and Medicare qualification. A call passes only when the agent achieves the required outcome and correctly saves all expected data fields, with consistency measured by the proportion of scenarios that succeed across every run.

  • Short-form healthcare receptionist suite covers 59 scenarios testing patient lookups, availability checks, and booking modifications.
  • Long-form Medicare intake suite assesses 23 scenarios with 42 fields covering caller consent, qualification, and routing handoffs.
  • Reliability ranking uses a pass-cubed standard requiring successful completion across all three consecutive runs.
  • Individual call transcripts, tool calls, latency figures, and outcome scores are available for inspection.

What to Test Before Choosing Cekura Bench

Checks to run with your own material and workflow

  • Verify whether the benchmark scenarios for clinic appointments and Medicare intake reflect the conversational patterns required by your voice application.
  • Check model reliability across specific challenge categories such as background noise, packet loss, interruptions, and caller corrections.
  • Review the public call reports and tool logs to observe how models manage complex multi-field data entry and handoff notes.

Cekura Bench Sources and Last Checked

What was checked and when

Last checked

Cekura Bench Frequently Asked Questions

Answers based on the source-checked product record

What scenarios are tested in the Cekura Bench speech-to-speech evaluation?

The benchmark evaluates 82 scenarios across two suites: a 59-scenario clinic receptionist suite for appointment booking and a 23-scenario Medicare intake suite covering consent, qualification, and routing.

How does Cekura Bench calculate reliability?

Reliability is measured using a pass-cubed metric, which calculates the percentage of scenarios that successfully pass across three consecutive live-call runs.

What conditions are held constant across evaluated models?

All models run the same open-source Pipecat agent configuration with identical prompts, tool definitions, and connection setups to isolate model performance.

Does Cekura Bench compare native speech models against cascaded architectures?

Yes, an unranked reference cascade combining Flux for speech-to-text, GPT-4.1 for reasoning, and ElevenLabs Flash for text-to-speech is included for comparative context.

Where can test logs and conversation transcripts be inspected?

Each model entry links directly to Cekura reports containing complete transcripts, tool call executions, and outcome scores for verification.

Explore other recently added tools in the same category.