Evaluation and benchmarking datasets
Voice agent evaluation sets: real call scenarios with outcomes
Quick answer
A voice agent evaluation dataset is a held-out set of call scenarios, each pairing a caller goal, the policy that applies, the expected end state and, ideally, real caller audio, so you can score task success, policy adherence, interruption handling and latency together. Public benchmarks give you a harness and a common yardstick. The scenarios that predict production behavior usually come from your own domain's real calls: their intents, dispositions, escalations and messy acoustics.
By SourceX Editorial · Updated
What a voice agent test set has to measure
A useful voice agent test set scores four things at once: whether the task got done, whether policy was followed, whether turn-taking held up, and how fast the agent responded. Text agent benchmarks already cover the first two well. τ-bench pairs a simulated user with tool APIs and a written domain policy, then grades by comparing the final database state with an annotated goal state rather than judging the transcript [1]. Its pass^k metric, the share of tasks an agent completes in all k repeated trials, exposes agents that succeed once and fail on the retry [1].
Voice adds failure modes that a text replay never triggers. Cascaded stacks chain voice activity detection (VAD), ASR, a text model and TTS, and each hop adds delay and discards prosody and other non-linguistic cues [3]. Full-duplex systems such as Moshi listen while speaking, which makes barge-in, backchannels ("mm-hm") and overlapping speech part of the behavior under test [3]. In March 2026, the τ-bench family released τ³-bench, which adds full-duplex voice evaluation, a sign that benchmark builders now treat voice-specific scoring as a distinct need [2].
| Dimension | What you check | Typical failure it catches |
|---|---|---|
| Task success | Final system state matches the goal state (record updated, appointment booked, refund issued) | Agent says "done" but the API call never fired |
| Policy adherence | Required disclosures, identity verification, prohibited actions | Refund granted without verifying the account holder |
| Turn-taking | Barge-in recovery, end-of-turn detection, backchannel tolerance | Agent talks over a caller reading out a card number |
| Latency | Time from caller end-of-speech to first agent audio, per turn | Long silences after tool calls that prompt callers to hang up |
| Recognition robustness | Entity capture (names, IDs, addresses) under accent, noise and codec loss | Misheard digits that silently corrupt a booking |
Why real calls beat scripted scenarios for voice
Real calls supply the distribution of intents, phrasing and acoustics that scripted or synthetic scenarios miss. Most public speech corpora are read or scripted speech; spontaneous conversational speech, with false starts, self-corrections and fillers, is comparatively scarce [5]. A scenario writer rarely invents a caller who changes their mind mid-sentence or reads an order number with a dropped digit, yet those are the turns that break production agents.
Historical call records also carry the label you need most: what actually happened. A contact center platform typically stores a disposition code, transfer and escalation flags, after-call notes and linked CRM or ticket changes. Those let you set the expected end state from the real outcome rather than from a writer's guess. The Action-Based Conversations Dataset showed the value of this structure in text, with agents required to follow company guidelines and take specific actions rather than just fill slots [4].
Scripted scenarios still have a job. Use them for rare, high-risk paths you cannot wait to observe, such as a caller disclosing self-harm or a fraud script, and keep them flagged so they do not inflate headline success rates. Our guide on stratified evaluation sets for rare and high-risk cases covers allocation, and synthetic evaluation data limits covers where generated scenarios mislead.
Anatomy of a scenario record
Each scenario should be a self-contained, machine-gradable record that a harness can replay as audio or text. The core is the caller goal, the policy excerpt that governs it, the starting backend state and the expected end state; audio and timing annotations sit alongside. Grading against state, as τ-bench does, avoids rewarding agents that sound confident but change nothing [1].
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"scenario_id": "billing-dispute-0147",
"source": {"type": "real_call", "call_date_month": "2025-11", "channel": "PSTN", "codec": "G.711 u-law 8 kHz", "redaction": "PII spans replaced with tone + tag"},
"intent": "dispute_duplicate_charge",
"secondary_intents": ["update_email"],
"caller_profile": {"speech": "spontaneous", "noise": "car_cabin", "accent_tag": "en-US-southern", "barge_ins": 3},
"policy_refs": ["verify_last4_and_zip_before_account_changes", "refunds_over_threshold_require_supervisor"],
"initial_state": {"account_status": "active", "open_charges": ["chg_A", "chg_A_dup"]},
"expected_end_state": {"refund_issued": ["chg_A_dup"], "email_updated": true, "escalated": false},
"historical_disposition": "RESOLVED_REFUND",
"turn_annotations": [
{"turn": 4, "event": "caller_barge_in", "t_ms": 18420, "expect": "agent_yields_within_ms<=500"},
{"turn": 7, "event": "entity_readout", "slot": "zip", "gold": "[REDACTED_ZIP]"}
],
"audio": {"caller_track": "caller.flac", "stereo_split": true, "reference_transcript": "verbatim_v2"},
"grading": {"state_match": "exact", "policy_checks": ["identity_verified_before_refund"], "latency_p95_ms_target": "team-defined"}
}
Three fields deserve attention when you source data. Stereo-split (dual-channel) recordings let you replay the caller track alone against your agent; mono mixes force separation and degrade turn-taking tests. The historical disposition anchors the expected outcome. The transcript style version matters because word error rate shifts with transcription conventions for fillers, numbers and casing, so pin one style and normalize both sides [6].
Text replay, audio playback and live simulation
There are three ways to run a voice scenario, and each measures something different. Text replay feeds the transcript to the dialogue layer; it is cheap and isolates reasoning and policy, but it hides ASR errors, endpointing and latency. Audio playback streams the real caller track into the full stack, which tests recognition and turn-taking but cannot react when your agent diverges from the original call's path.
Live simulation uses a user simulator, driven by the scenario's goal and persona, that speaks through TTS or a speech model and responds to whatever the agent says. This is the τ-bench pattern extended to audio [1][2]. Ground the simulator in real call behavior, including interruptions, hesitations and partial answers, or it will be more cooperative than real callers; see user simulator scenarios grounded in real conversations.
A practical stack runs all three: text replay on every build, audio playback on release candidates, and live simulation on a stratified subset. Track the gap between text and audio success on the same scenarios; that gap is your speech-modality penalty, discussed further in evaluation data for voice models and the modality gap.
Label quality, coverage and contamination checks
Outcome labels from call systems are noisy and need an audit before they become ground truth. Agents pick the wrong disposition code, after-call notes are inconsistent, and "resolved" sometimes means "caller gave up." Audits of widely used test sets, audio among them, estimate average label error rates of at least 3.3% [7], and operational dispositions are usually rougher than curated benchmark labels.
Before accepting a set, ask for:
- Disposition validation: a reviewer re-labels a random sample against the backend record, and you get the agreement rate.
- Coverage report: counts by intent, outcome, escalation path, channel, codec and noise condition, so you can see which strata are thin.
- Time window: when calls were recorded, since policies, products and IVR menus change and stale scenarios encode retired rules.
- Hold-out guarantee: written confirmation that the scenarios were not also licensed or released as training data to you or publicly; see contamination-resistant evaluation design.
- Redaction map: how spoken PII was masked in audio (tone, silence, synthetic replacement) and transcripts, because masking style can itself confuse ASR and entity tests.
Consent, voiceprints and redaction in call audio
Call audio is personal data twice over: what callers say and the sound of their voice. Recording-consent rules differ by jurisdiction, with some US states requiring all parties to consent, so confirm that the original call notice covered recording and that the supplier's terms permit reuse for model evaluation. Treat "this call may be recorded for quality and training purposes" as a starting point for counsel's review, not a conclusion.
Illinois BIPA lists voiceprints among biometric identifiers and requires notice and a written release before collection [8]. If your harness extracts speaker embeddings, for example for diarization or caller verification tests, scope that use explicitly. Health-related calls handled by HIPAA covered entities or business associates, such as provider-to-payer or pharmacy calls, can contain protected health information, and HIPAA's Safe Harbor de-identification standard lists biometric identifiers including voiceprints [9]; the provider-to-payer calls guide covers that segment.
Redaction is never perfect in audio. Spelled-out names, digits split across turns and background speech slip past detectors. Plan to measure residual PII on a sample and decide whether some evaluation runs should use transcripts only.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Request template for a call-scenario evaluation set
The fastest way to source a usable set is to describe scenarios and outcomes, not vendors. A tight request names the intents, the outcome labels you will grade against, the audio properties your stack depends on and the use you need licensed.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example entry |
|---|---|
| Use | Held-out evaluation of an inbound billing voice agent; no training use |
| Call types | Inbound customer service, US English, billing and account changes |
| Intents needed | Duplicate charge, plan downgrade, payment arrangement, address change |
| Outcome labels | Disposition code, escalation flag, linked backend change (refund, plan change) |
| Policy material | Agent scripts or knowledge-base articles in force during the recording window |
| Audio | Dual-channel, original codec preserved, caller-side track separable |
| Annotations | Verbatim transcript with timestamps; barge-in and overlap markers |
| Privacy | Spoken PII masked in audio and transcript; redaction method documented |
| Coverage | Minimum share of escalated and failed calls, not only resolved ones |
| Recency | Recorded within a window matching current policies |
How SourceX fits
SourceX sources operational datasets, including support and sales histories, from US companies on request, and manages the licensing and ongoing purchases. Nothing is held in stock, so a described call-scenario set is searched for, not pulled from a catalog, and a match is not guaranteed. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and health records require HIPAA de-identification. Delivery happens through private, access-controlled workflows only after an executed agreement and supplier approval; SourceX does not train models. If that matches your evaluation plan, you can describe your call-scenario requirements to SourceX.
For neighboring needs, the owner pages on voice agent training data from real conversational audio and licensing voice and audio data cover training use, while contact center call recordings and evaluation sets built from real business work cover adjacent sourcing. The evaluation datasets hub maps the rest of the cluster, including policy-following evaluation for support agents and agent task suites with state-based grading.
Source real call scenarios for your voice agent evaluation
Describe the intents, outcomes and audio conditions your voice agent must handle, and SourceX looks for US businesses that hold matching call data. Every release is approved by the supplying company and delivered under a license defining records, uses, term and delivery. Start a buyer request at SourceX.
Sources
- Sierra Research (arXiv:2406.12045), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- Sierra Research (GitHub), "τ³-bench 1.0.0 — Voice, Knowledge, Task Quality" (2026). https://github.com/sierra-research/tau2-bench/releases/tag/v1.0.0
- Kyutai (arXiv:2410.00037), "Moshi: a speech-text foundation model for real-time dialogue" (2024). https://arxiv.org/html/2410.00037v1
- ASAPP / NAACL 2021 (arXiv:2104.00783), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
- arXiv:2506.00267, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
- arXiv:2412.07937, "Style-agnostic evaluation of ASR using multiple reference transcripts" (2024). https://arxiv.org/pdf/2412.07937
- Northcutt, Athalye, Mueller (NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://ilga.gov/Legislation/ILCS/Articles?ActID=3004&ChapterID=57&Print=True
- eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.