Speech and audio data
Spontaneous Conversational Speech Data: Why Read and Scripted Speech Fall Short
Quick answer
Spontaneous conversational speech is unplanned talk between people: fillers, restarts, self-corrections, overlaps, backchannels and real turn-taking. Read and scripted corpora contain almost none of this, which is why models trained mainly on them degrade on live calls and meetings. To buy it well, specify the speech style explicitly, require turn-level timestamps and overlap marking, decide how disfluencies are transcribed, and audit consent for every recorded party before you license a single hour.
By SourceX Editorial · Updated
How spontaneous speech differs from read, scripted and elicited speech
Spontaneity is a property of how the words were produced, not of the topic, so four collection styles give four very different datasets. Most public speech data is read or scripted; the CASPER authors note that most existing datasets contain scripted dialogue, which is the gap their elicited-conversation corpus targets [1]. Crowdsourced read corpora such as Common Voice, where contributors read prompted sentences, are valuable for accent and vocabulary coverage but capture no conversational behavior [4].
| Style | How it is produced | What it captures | What it misses |
|---|---|---|---|
| Read | Speaker reads a prompted sentence | Pronunciation, accents, clean phonetics | Fillers, restarts, overlap, prosody of real interaction |
| Scripted dialogue | Actors perform a written two-party script | Turn structure, domain terms | Hesitation, interruptions, unplanned repair |
| Elicited or improvised | Participants get a task or topic but no script | Much natural disfluency and turn-taking | Real stakes; some performance effects remain |
| Operational (naturally occurring) | Recorded during real work: support, sales, scheduling, meetings | Real intents, overlap, noise, domain vocabulary | Controlled balance; needs consent and redaction review |
Elicited and improvised designs sit in the middle. One Interspeech study had actors improvise agent and user roles precisely because found conversational recordings were too disfluent to use directly for TTS [2]. That is a reasonable choice for synthesis, and a poor one for an ASR model that must handle the messiness the actors removed.
Why read-speech training fails on live conversation
Read-speech models fail on conversation because the acoustic and linguistic distribution shifts on several axes at once. A model can score well on clean read benchmarks and still mis-segment a call where a caller says "I, uh, no wait, the, the second one" while the agent says "mm-hm" over them.
The common failure modes buyers see in production:
- Disfluency hallucination or deletion: fillers ("um", "uh") and repetitions are dropped inconsistently, so word error rate (WER) depends on transcript convention rather than model quality.
- Overlap collapse: when two speakers talk at once, single-stream models emit one merged or truncated transcript. Natural conversation contains frequent overlaps, backchannels and laughter [3].
- Endpointing errors: voice agents trained on read utterances treat mid-sentence pauses as end of turn and interrupt the user.
- Prosody mismatch in TTS: synthesis trained on read speech sounds like reading, which is noticeable in agent replies.
- Speech-language model turn-taking: full-duplex and speech-to-speech models need examples of when to yield, backchannel or barge in; scripted data rarely contains them. See full-duplex conversation data for the channel requirements.
Keep, tag or normalize disfluencies: decide before you buy
Disfluencies should be kept and tagged in the source transcript, and normalized only in derived views, because you cannot restore what a supplier deleted. Spontaneous recordings carry frequent disfluencies, restarts, fillers and overlaps [2], and every one is a labeling decision.
A workable policy separates three layers:
- Audio: untouched apart from agreed redaction.
- Verbatim transcript: fillers, repetitions, false starts, partial words, laughter and noise events tagged with a fixed tag set (for example
<filler>,<partial>,<laugh>,<overlap>). - Normalized view: generated by script for training targets that need clean text, such as LLM text or summarization pairs.
Ask suppliers which convention they follow and request a written style guide. Our guide on verbatim vs clean transcription standards covers tag sets and how they affect WER comparisons; speech transcripts as LLM training text covers the normalization side.
Timestamps, channels and overlap marking to require
Require turn-level start and end times per speaker, explicit overlap regions and, where possible, one channel per speaker. Recent full-duplex dataset work records natural conversations with frequent overlaps, backchannels and laughter [3]; timing and overlap annotation are what let a model learn turn-taking from them.
Specific requirements to put in the request:
- Channel layout: dual-channel (stereo, one speaker per channel) telephony or separate meeting microphones beat a mixed mono track for overlap work. Confirm sample rate and codec; telephony is often 8 kHz narrowband. See audio file specs.
- Diarization: speaker labels stable across the whole session, with a role field (agent, customer, participant).
- Segment timing: turn boundaries to at least 10 ms resolution; word-level alignment if you train endpointing or streaming ASR (forced alignment guide).
- Overlap and backchannel tags: marked as events, not silently merged into the dominant speaker.
- Packaging: a manifest with per-segment offsets and RTTM-style diarization files (manifest packaging).
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"session_id": "sess_000412",
"style": "operational_spontaneous",
"setting": "inbound_support_call",
"channels": {"ch0": "agent", "ch1": "customer"},
"sample_rate_hz": 8000,
"codec": "pcm_s16le",
"turns": [
{"spk": "customer", "ch": 1, "start": 12.34, "end": 15.02,
"text": "yeah so I <filler>uh</filler> I tried the, the reset", "events": ["restart"]},
{"spk": "agent", "ch": 0, "start": 14.71, "end": 15.10,
"text": "mm-hm", "events": ["backchannel", "overlap"]}
],
"redaction": {"method": "tone_replace", "tags": ["<CARD_NUMBER>", "<PERSON>"]},
"consent_basis": "all_party_notice_and_consent"
}
Where spontaneous speech comes from: research corpora vs operational recordings
There are three practical sources, and each has a different rights and quality profile. Research corpora are well documented but often small, dated or licensed for research only. Commissioned elicitation is controllable but costs real collection effort. Operational recordings are spontaneous by construction but need the heaviest consent review.
Classic telephone conversation corpora. Older two-party telephone collections are short, narrowband and decades old, so check recording age, channel and license before relying on them. Corpora distributed through the Linguistic Data Consortium carry membership and license terms that distinguish for-profit from non-profit use [5]; confirm which applies before commercial training. Our open speech corpora license audit walks through the checks.
Elicited conversation datasets. Recent work shows elicitation can produce large volumes of casual speech: CASPER reports 200 hours recorded, 158 hours of speech, with the first 102 hours publicly released [1]. Check the license of any public release separately from the paper.
Operational recordings. Support, sales, scheduling and meeting audio captures real intents, accents, background noise and domain vocabulary that no script anticipates. The cost is review: who consented, whether payment card numbers or health details were spoken, and how redaction was applied. Redaction itself changes the audio, as covered in how audio redaction affects training. For comparisons of collection routes, see off-the-shelf vs custom vs licensed speech data.
Consent and recording-law checks for conversational audio
Every voice in a conversation is a data subject, so consent must cover all parties, not just the employee who pressed record. California, for example, requires consent of all parties to record cellular or cordless telephone communications [6]; many other US states follow one-party rules, and calls across state lines raise the question of which rule applies.
Questions to put to any supplier:
- What notice did each party hear, and is the notice text available?
- Does the original consent or customer agreement allow use for AI training, or only for quality and training of staff?
- Were spoken payment card numbers, account numbers and health details redacted in audio and transcript?
- Does the license allow voice cloning or speaker identification, or exclude them? See speech and voice license terms.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Spontaneous speech request checklist
A good request describes the speech style and setting, not a vendor or a corpus name. Use this checklist to write one.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
| Speech style | Naturally occurring, unscripted two-party calls | Excludes read and acted data |
| Setting | Inbound support and appointment scheduling | Sets intent and vocabulary distribution |
| Channel | Dual-channel, 8 kHz or wider | Enables overlap and per-speaker training |
| Transcript convention | Verbatim with filler, partial and overlap tags | Preserves disfluencies |
| Timing | Turn-level timestamps, overlap regions | Turn-taking and endpointing |
| Speaker metadata | Role, accent region, device type where known (metadata fields) | Coverage and bias analysis |
| Redaction | Method named, tags listed, sample QA | Avoids training on noise you cannot explain |
| Rights | All-party consent basis documented | Defensible use |
| Documentation | Data statement describing speakers and situation [7] | Precise generalization claims |
Accent coverage is a separate axis from spontaneity; see accented English speech data. For voice-agent use specifically, see voice agent training data, and for meeting audio, licensed meeting transcripts.
How SourceX approaches conversational speech requests
SourceX sources operational datasets from US companies, including support and sales histories, and manages the licensing process; datasets are sourced on request, not held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, and personal details such as names, phone numbers and account numbers are removed or replaced before delivery, with the method recorded and a sample checked. No de-identification method is perfect. Buyers can describe the conversational audio they need without naming businesses; every release is approved by the supplying company. More context is on the speech and audio cluster hub, the voice and audio data page and the AI data hub.
Find spontaneous conversational speech for your model
SourceX looks for US businesses that hold the conversational recordings you describe, reviews data and licensing permissions, and agrees allowed uses in a license before anything is delivered; nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows after an executed agreement. Describe your spontaneous speech requirements to SourceX.
Frequently asked questions
Is elicited speech good enough for ASR training?
Elicited conversation is far closer to real speech than read prompts and can be collected at scale [1]. It still lacks real stakes, such as a frustrated caller or a time-pressured agent, so test on held-out operational audio before assuming it transfers.
Should I strip fillers before training ASR?
Keep them in the source transcript and decide per model. Streaming ASR and endpointing models benefit from seeing fillers; a summarization target may not. A normalized view built by script keeps both options open.
Can synthetic TTS speech replace spontaneous recordings?
Synthetic speech is usually generated from text and inherits a read-speech style unless the generator itself was trained on conversation. See real vs synthetic speech data for where augmentation helps.
Sources
- arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
- Adigwe and Klabbers, "Strategies for developing a conversational speech dataset for Text-To-Speech Synthesis," Interspeech 2022. https://www.isca-archive.org/interspeech_2022/adigwe22_interspeech.pdf
- arXiv, "Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis" (2025). https://arxiv.org/pdf/2509.04093
- arXiv (Ardila et al., Mozilla), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
- Linguistic Data Consortium, "LDC For-Profit Membership". https://Catalog.Ldc.Upenn.Edu/license/ldc-for-profit-membership.pdf
- California Legislative Information, "California Penal Code section 632.7". . https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632.7
- TACL (Bender and Friedman), "Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science" (2018). https://aclanthology.org/Q18-1041/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.