Skip to content

Speech and audio data

Real vs Synthetic Speech Data for ASR and Voice Models: Where TTS Augmentation Falls Short

Quick answer

Synthetic speech is a reliable tool for coverage problems: wake-word hard negatives, rare words and entity names, and room or noise simulation matched to a known microphone array. It is a weak substitute for behavior problems: spontaneous disfluency, overlapping turns, emotion, and the artifacts of real telephony and device channels. Use TTS and simulation to fill measured gaps, keep real recordings as the core of training for conversational ASR, and always score on real held-out audio.

By SourceX Editorial · Updated

Where synthetic speech reliably helps ASR and wake-word training

Synthetic audio pays off when the gap is lexical or acoustic coverage that you can specify precisely, not human behavior you cannot. Three uses are well established.

Hard negatives for wake words and keyword spotting. A wake-word model trained naively on positives can detect the keyword every time and still fire on unrelated words; Deepgram describes a model triggering on words like "app" and "salad" and fixing it by generating the confusers it kept accepting with TTS across many voices, then retraining with a large share of negatives (about 1:10 positives to negatives) [2]. TTS is efficient here because you need breadth of near-miss phonetics, not natural conversation.

Rare words, names and numbers. When real recordings rarely contain your drug names, SKUs, street names or account-number formats, TTS can render them in many voices so the decoder sees the token sequences at all. This is a vocabulary fix; it does not teach how real callers mumble, spell out, or correct those terms. For specifying the term lists themselves, see domain vocabulary coverage in speech data.

Room and noise simulation for known hardware. Convolving clean speech with room impulse responses (for example from image-source simulators such as pyroomacoustics) and mixing noise at controlled SNRs is standard for far-field front ends. One far-field study trained on simulated reverberant, noisy audio matched to the device's microphone geometry [3]. That works because the array geometry is known and fixed; it is weaker when the device, placement or room population is not. Our far-field real vs simulated rooms guide covers this trade in depth.

Where TTS-generated speech falls short

TTS output is too clean and too regular to stand in for natural speech, and models trained on it struggle to generalize. A 2025 paper on speech language models for emotional conversation notes that approaches relying on TTS-synthesized speech face challenges generalizing to natural human speech [1]. The failure modes are predictable once you list what TTS does not produce.

  • Spontaneity and disfluency. Real conversation contains filled pauses, false starts, repairs, trailing sentences and backchannels ("mm-hm"). TTS renders written text, so these appear only if you script them, and then they sound scripted. The CASPER authors built a spontaneous conversation corpus precisely because most existing speech datasets are read or scripted [5]; see spontaneous conversational speech data.
  • Prosody and emotion. Frustration, hesitation, sarcasm and urgency change pitch, rate and energy in ways expressive TTS approximates but does not sample from real distributions. Emotion recognition and voice agents that detect escalation need labeled real emotion, not prompted styles.
  • Speaker and accent distribution. A TTS voice bank is a small, curated population. Accent-related accuracy gaps in ASR are documented [9], and you cannot close them with voices that were never recorded from the under-served groups. See accented English speech data.
  • Turn-taking and overlap. Crosstalk, interruptions and simultaneous speech are interactional; synthesizing two TTS streams and summing them produces overlap without the timing behavior of real conversation. See overlapping speech data.
  • Vocoder fingerprints. Neural vocoders leave spectral regularities. A model can learn to rely on those cues, which inflates scores on synthetic dev sets and disappears on real audio.

Why real channel conditions are hard to simulate

Production audio passes through a chain of codecs, gain stages and network effects that simple augmentation does not reproduce. A vendor in the call-center data market argues that ASR trained mainly on clean call audio breaks on live calls with overlap, hold music, transfers and background noise [4]; treat that as a self-reported market view, but the failure pattern is familiar to anyone who has deployed contact-center ASR.

You can down-sample to 8 kHz and pass audio through G.711 mu-law or AMR-NB to approximate narrowband telephony, and you can add MUSAN-style noise. What you cannot easily simulate is the joint distribution: VoIP packet-loss concealment, automatic gain control pumping, echo-canceller residue, speakerphone pickup, IVR prompts bleeding into the caller channel, and agent-side headsets in a noisy floor, all at once and correlated with the conversation content. The telephony 8 kHz vs wideband guide and noisy speech SNR coverage guide cover how to specify these conditions when you buy real audio.

Decision table: synthetic, simulated or real per gap

Match the data source to the specific gap your error analysis shows, rather than setting a global synthetic percentage.

Illustrative example: invented to show structure; it does not describe an available dataset.

Gap found in error analysisSynthetic/simulated is a good fitReal recordings requiredNotes for the spec
Wake word false accepts on near-miss wordsYes: TTS confusers in many voices [2]For final false-accept rate per hour on real ambient audioTrack false accepts per hour on real household or vehicle audio
Missing product names, drug names, IDsYes: TTS renders of term listsFor how real speakers spell, correct and abbreviatePair with real utterances containing the terms
Reverberation on a fixed device arrayYes: RIR simulation matched to geometry [3]For placement and room diversity you did not modelRecord device-captured audio in target rooms
Telephony codec mismatchPartly: codec transcodingYes, for packet loss, AGC, echo residueAsk for native 8 kHz call audio, not re-encoded wideband
Disfluencies, repairs, backchannelsNoYes [5]Verbatim transcripts that keep fillers and repairs
Emotion and escalationNoYesReal labeled segments with annotator agreement
Accent and dialect gapsNoYes [9]Speaker metadata per segment, self-reported where possible
Overlap and turn-takingWeakYesSeparate channels or diarization with overlap regions

How to evaluate synthetic-augmented ASR on real held-out audio

The only trustworthy test of a synthetic-augmented model is word error rate, and task metrics, on real recordings the model and the synthetic pipeline never touched. Build the test set before you generate anything, freeze it, and keep it out of any prompt, voice-cloning or term-list workflow.

Report results per slice, not as one number: channel (8 kHz call, wideband app, far-field device), accent group, speaking style (read, spontaneous), SNR band and domain-term subset. Compare three arms: real-only, real plus synthetic, synthetic-heavy. If the synthetic arm wins on a synthetic dev set but not on the real test set, the model is learning TTS cues. For voice agents and speech-LLMs, add end-to-end checks on real calls, since text-to-speech probes understate the speech modality gap; our guide to using a licensed real-data holdout to validate synthetic data covers holdout design across modalities.

A practical acceptance checklist:

  • Real test set built and frozen before synthetic generation, with documented source and collection dates.
  • No speaker in the test set is used as a voice-cloning reference or TTS fine-tuning voice.
  • WER reported per channel, accent, style and SNR slice, with confidence intervals.
  • Domain-term error rate reported separately from overall WER.
  • Wake-word models report false accepts per hour on real ambient audio and false rejects on real speakers.
  • Synthetic share per training batch recorded so results can be reproduced.

Synthetic speech is not automatically rights-free; it inherits questions from the voices and corpora used to build the TTS system. A 2025 position paper observes that many TTS systems are trained on speech gathered from the internet, where consent and licensing are hard to verify at scale, and that some technical reports describe training data only as "in-house" [6]. Research on voice actors frames the exposure as privacy, reputation, accountability, consent, credit and compensation [7].

Before using a commercial or open TTS model to generate training audio, ask: does the model's license allow its output to train a competing or commercial speech model, were the source voices recorded with consent for synthesis, and are any voices clones of identifiable people. Cloned voices can also raise voiceprint questions; see voiceprints and BIPA. For TTS built on consented studio recordings, see consented TTS training data. Record the TTS engine, version, voice IDs and license in your dataset documentation, using a structure such as Data Cards [10].

Speaker anonymization is a different technique with a similar trade: VoicePrivacy Challenge evaluations measure both how well identity is hidden and how much ASR utility is kept [8], so anonymized real speech is not equivalent to the original for training.

Specifying real speech data to pair with synthetic augmentation

Buy real audio for the gaps synthesis cannot close, and write the request around those gaps. A useful real-data spec names channel and sample rate (native 8 kHz mono call legs or stereo agent/customer channels, 16 kHz wideband, device far-field multichannel), speaking style (spontaneous, task-oriented), domains and term lists, accent and demographic metadata fields, transcript standard (verbatim with fillers, overlap tags, redaction tags), and how personal details were removed. Redaction choices matter for training; see audio redaction artifacts.

Operational audio held by businesses, such as support and sales call histories and new recordings of hands-on work, is the type of real-world speech that covers channel realism and spontaneity. SourceX sources these datasets from US companies on request rather than holding stock, so a request does not guarantee a match; each supplying company approves every release, and SourceX does not source scraped web content. Teams outlining a request can start at the SourceX buyer page, and the contact-center call recordings and voice agent training data pages describe the categories. For how licensed and synthetic data combine across modalities, see combining licensed and synthetic data and the synthetic data glossary entry; the speech and audio data hub covers the full cluster.

Request real-world speech data for ASR

SourceX finds US businesses that hold the speech data you describe, reviews ownership and consents, and delivers under a license that defines records, uses, term and delivery. Names, phone numbers and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked, though no method is perfect. Describe the channels, styles and gaps you need at sourcex.si/buyers.

Sources

  1. arXiv, "Dual Information Speech Language Models for Emotional Conversations" (2025). https://arxiv.org/pdf/2508.08095
  2. Deepgram, "Naively training a wake word model from scratch". https://deepgram.com/learn/naively-training-a-wake-word-model-from-scratch
  3. arXiv, "Spatial Attention for Far-field Speech Recognition with Deep Beamforming Neural Networks" (2019). https://arxiv.org/pdf/1911.02115
  4. AIxBlock, "Call Center Audio Dataset: What Makes It Production-Ready". https://aixblock.io/blogs/call-center-audio-dataset-what-makes-it-production-ready
  5. arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
  6. arXiv, "Position: Towards Responsible Evaluation for Text-to-Speech" (2025). https://arxiv.org/pdf/2510.06927
  7. arXiv, "PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy" (2025). https://arxiv.org/pdf/2507.16247
  8. arXiv, "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677
  9. DeepAI (publication listing), "Performance Disparities Between Accents in Automatic Speech Recognition". https://deepai.org/publication/performance-disparities-between-accents-in-automatic-speech-recognition
  10. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data