Skip to content

Speech and audio data

Evaluation Data for Voice Models: Measuring the Speech Modality Gap

Quick answer

A speech LLM evaluation dataset should pair every task in two forms, text and real human speech, so you can measure how much accuracy the model loses when the input arrives as audio. Published benchmarks already show reasoning drops when the same task is spoken rather than typed [1]. Build the set from held-out, privately licensed recordings with varied speakers and acoustics, score pipelines stage by stage, and judge synthesized speech with a documented listener protocol rather than a single MOS number.

By SourceX Editorial · Updated

What the speech modality gap is and why text evals miss it

The modality gap is the difference between a model's score on a task given as text and its score on the identical task given as speech. A recent benchmark built voice and text versions of the same reasoning items and found performance falls when tasks are delivered by voice [1]. Text leaderboards cannot reveal that loss, because they never exercise the audio encoder, the speech tokenizer or the cascade's ASR front end.

The gap has several sources that a good eval set separates. Recognition errors on numbers, names and negations propagate into wrong answers. Prosody, disfluency and self-correction ("ship it Tuesday, no, Thursday") carry meaning that a clean transcript flattens. End-to-end models also have to hold spoken context across turns without a text scratchpad.

Cascaded systems chain voice activity detection, ASR, a text dialogue model and TTS, while full-duplex speech-text models such as Moshi process audio directly [4]. Your eval set must work for both architectures, so the scoring target is the task outcome, not the intermediate transcript. For broader context on benchmark choice, see the LLM evaluation datasets buyer's map.

How to build paired spoken and text tasks

Paired construction means each item has a canonical text prompt, one or more spoken renditions and a single reference answer, all sharing an item ID. Score the text version and the spoken version with the same grader, then report the per-item delta, not just two aggregate accuracies. Item-level pairing lets you attribute failures to specific phenomena such as spoken digits or long-range references.

Prefer real human speech for the spoken side. TTS-rendered prompts are cheap and perfectly aligned, but they tend to be fluent, evenly paced and acoustically clean, which understates the gap a deployed system sees; this is a working hypothesis that you should test on your own data by running both versions. The trade-offs mirror those on our page on real vs synthetic speech data.

Stratify the spoken renditions deliberately. Useful axes include accent and dialect, speaking rate, channel (8 kHz telephony, 16 kHz wideband, far-field microphone), background noise, and read versus spontaneous delivery. Record these as metadata fields so you can slice the gap; a coverage approach is outlined in coverage gap analysis.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "sqa-0412",
  "task_type": "spoken_qa_multi_hop",
  "text_prompt": "My order was 3 units at $14.50 and I returned 1. What is my refund?",
  "reference_answer": "$14.50",
  "audio_renditions": [
    {"file": "sqa-0412_spk07.flac", "speaker_id": "spk07", "source": "human_read",
     "accent": "en-US-South", "channel": "telephony_8k", "snr_db": 18, "duration_s": 6.2},
    {"file": "sqa-0412_spk22.flac", "speaker_id": "spk22", "source": "human_spontaneous",
     "accent": "en-IN", "channel": "far_field_16k", "snr_db": 9, "disfluency": true},
    {"file": "sqa-0412_tts01.flac", "speaker_id": "tts_voiceA", "source": "tts",
     "channel": "clean_24k"}
  ],
  "verbatim_transcript": "uh my order was three units at fourteen fifty and I sent one back",
  "phenomena": ["spoken_currency", "arithmetic", "disfluency"],
  "split": "held_out_private",
  "license_scope": "evaluation_only",
  "consent_record_id": "cr-88213"
}

Keep the verbatim transcript alongside each recording, following a documented convention such as the ones compared in verbatim vs clean transcription standards. It lets you compute an oracle-transcript score, which isolates how much of the gap is recognition error versus reasoning over audio.

Scoring cascades and voice agents stage by stage

Score a voice pipeline at every stage and at the outcome, because a good end score can hide a broken stage and vice versa. A health-conversation benchmark evaluates diarization, ASR and downstream summaries together on the same audio, which is the pattern to copy [3]. For multi-speaker audio that means diarization error rate (DER), a speaker-attributed WER such as time-constrained minimum-permutation WER (tcpWER), and a task metric on the final output.

Illustrative example: invented to show structure; it does not describe an available dataset.

StageMetricWhat a failure looks likeData the eval set needs
Segmentation and diarizationDER (missed speech, false alarm, speaker confusion)Agent attributes the caller's refusal to itselfMulti-speaker audio with time-stamped speaker turns
RecognitionWER, tcpWER, entity error rate on numbers and names"fifteen" heard as "fifty"Verbatim reference transcripts with normalization rules
Understanding and reasoningTask accuracy, text-vs-speech delta per itemCorrect with typed prompt, wrong when spokenPaired text and spoken items with one reference answer
Dialogue behaviorTurn-taking latency, barge-in handling, interruption recoveryTalks over the user or waits too longFull-duplex recordings with overlaps and backchannels
Speech outputListener MOS, CMOS, intelligibility via ASR-WER on outputMispronounced drug name, flat prosody on a questionText inputs with pronunciation-critical terms
OutcomeResolution, policy compliance, summary factualitySummary invents a dosageGround-truth outcomes or expert-written references

For speech-to-speech and full-duplex models, the dialogue-behavior row matters most and needs natural overlapping conversation, covered in full-duplex conversation data. Voice agents that complete transactions also need outcome labels from real calls; see voice agent evaluation sets.

How synthesized speech should be judged

TTS evaluation should report the listener protocol, the listener pool and the provenance of both the test text and any reference recordings, not just a headline MOS. A recent position paper argues for more responsible TTS evaluation and specifically flags data provenance as an under-reported factor [2]. Without that detail, two MOS figures from different labs are not comparable.

Specify the listening test before you collect ratings. State the method (absolute category rating on a five-point scale, comparative CMOS, or forced-choice preference), the number of listeners per stimulus, how listeners were screened, which playback devices were allowed, and which gold or trap stimuli were used to reject inattentive raters. Report confidence intervals and prefer comparative tests when two systems are close, because absolute MOS compresses small differences.

Hold the test sentences out of training, and include pronunciation-critical strings from your domain: SKUs, drug names, addresses, mixed-language names. When comparing against a natural reference, the reference voice must come from consented speakers; release terms are discussed in voice talent consent and release terms.

Keeping spoken eval sets clean and uncontaminated

An eval set only measures generalization while it stays out of every model's training data. LiveBench documents how public test data leaks into newer models' training sets and responds by refreshing questions [5]; audio benchmarks published to public hubs face the same risk. Keep the audio private, license it for evaluation only, restrict access by named team, and rotate a share of items each cycle. Our glossary entry on benchmark contamination and the guide to contamination-resistant evaluation design go further.

Audit labels before trusting deltas. A NeurIPS 2021 audit of test sets across vision, NLP and audio datasets estimated an average label error rate of at least 3.3% [6], and a small modality gap can sit inside that noise. Double-transcribe a sample, adjudicate disagreements and publish the normalization rules.

Treat the eval set as a governed asset. The NIST AI RMF organizes risk work into GOVERN, MAP, MEASURE and MANAGE functions [8]; for the MEASURE work, record speaker demographics, channels and collection dates alongside results.

Rights and privacy checks specific to evaluation audio

Evaluation audio still carries the same rights and privacy obligations as training audio, because recordings of real people are being copied, stored and processed. Confirm speaker consent covers evaluation by a third party, that the license defines records, uses, term and delivery, and that personal details in transcripts are removed or replaced. Redacted audio can shift metrics, a problem covered in how audio redaction affects speech model training.

Speaker-verification or voice-cloning evals can create voiceprints. In Illinois, the 2024 BIPA amendment treats repeated collection of the same biometric identifier from the same person by the same method as a single violation [7], but the consent requirement itself remains; see voiceprints and BIPA.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Illustrative example: invented to show structure; it does not describe an available dataset.

Spoken eval set request checklist

  • Task families and the text-side reference for each item (QA, multi-hop reasoning, instruction following, slot filling)
  • Speech source: human read, human spontaneous, real calls; TTS only as a comparison condition
  • Stratification targets: accents, channels (8 kHz, 16 kHz, far-field), SNR bands, speaking rate
  • Reference artifacts: verbatim transcripts, speaker turns with timestamps, outcome labels
  • Audio spec: FLAC or WAV, sample rate and channel layout per audio file specs
  • License scope: evaluation only, no training, named-team access, refresh cadence
  • Consent and de-identification method recorded per recording

Where licensed real-world audio fits

Real operational audio, such as recorded support and sales calls, is a strong source of spoken eval items because it contains the accents, crosstalk, hold music and domain vocabulary your deployment will meet. SourceX sources operational datasets, including support and sales histories and new recordings of hands-on work, from US companies on request; categories are not inventory and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and no method is perfect. You can describe the eval data you need at the SourceX buyer page, and related context is on call center audio datasets, evaluation datasets built from real business work and the speech and audio data hub.

Request held-out spoken evaluation data

Describe the spoken tasks, speakers, channels and outcome labels your speech LLM, voice agent or TTS evaluation needs, and SourceX looks for US businesses that hold that data. Each release is approved by the supplying company and delivered under a license defining records, uses, term and delivery, and nothing is contracted until a supplier agrees. Start at sourcex.si/buyers.

Sources

  1. arXiv, "Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap" (2025). https://arxiv.org/pdf/2509.26542
  2. arXiv, "Position: Towards Responsible Evaluation for Text-to-Speech" (2025). https://arxiv.org/pdf/2510.06927
  3. arXiv, "Benchmarking Speech Systems for Frontline Health Conversations: The DISPLACE-M Challenge" (2026). https://arxiv.org/pdf/2603.02813
  4. Kyutai (arXiv:2410.00037), "Moshi: a speech-text foundation model for real-time dialogue" (2024). https://arxiv.org/html/2410.00037v1
  5. arXiv (ICLR 2025), "LiveBench: A Challenging, Contamination-Free LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
  6. Northcutt, Athalye, Mueller (NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. Illinois General Assembly, "SB 2979 (103rd General Assembly), BIPA amendment, engrossed text" (2024). https://www.ilga.gov/documents/legislation/103/SB/PDF/10300SB2979eng.pdf
  8. NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data