Speech and audio data
Training Data for Speech-Language Models: Audio-Text Instruction and Dialogue Pairs
Quick answer
Speech-language models need three kinds of data that text corpora and read-speech ASR sets do not supply: long-form conversational audio with aligned transcripts for pre-training, spoken instructions paired with responses for instruction tuning, and multi-turn dialogue with real turn-taking for speech-to-speech behavior. TTS-synthesized prompts fill gaps cheaply but generalize poorly to natural speech [1]. Buyers should specify speaker diversity, channel, transcript standard, timing alignment and voice consent before evaluating any source.
By SourceX Editorial · Updated
What a speech-language model actually consumes
A speech-language model consumes paired audio and text at several granularities, and each training stage wants a different pairing. Unlike an ASR model, which only maps audio to a transcript (see speech-to-text), a speech LLM has to understand spoken requests, reason over them and often answer in speech. That makes it a multimodal data consumer whose text side carries intent, not just words.
In practice, teams source four data types:
- Transcribed conversational audio for continued pre-training or audio-text alignment: hours of two-party or multi-party speech with time-aligned transcripts, speaker diarization and channel metadata.
- Spoken instruction-response pairs for instruction tuning: a user utterance in audio, plus a text or audio response that is correct for the request.
- Multi-turn spoken dialogue for conversational and speech-to-speech behavior: full sessions with overlaps, backchannels, interruptions and silences preserved.
- Paralinguistic labels such as emotion, speaking rate, hesitation or frustration, used for empathetic or style-aware responses. One empathetic speech-LLM effort assembled about 70,000 emotion-labeled utterances from several corpora to get enough coverage [3].
Why TTS-synthesized prompts are a stopgap
Synthetic speech is useful for bootstrapping instruction data, but models trained mostly on it break on real callers. Recent speech-LLM work notes that reliance on TTS-synthesized speech limits generalization to natural human speech [1]. TTS output lacks disfluencies, false starts, background noise, codec damage, accent variation and the prosodic cues people use to signal uncertainty or urgency.
A second gap shows up at evaluation. Researchers have diagnosed a modality-induced performance gap: the same reasoning tasks score worse when posed by voice than by text [2]. Real spoken task data, where people actually ask for things out loud, is one lever for closing that gap. For a broader comparison, see real vs synthetic speech data for ASR and voice models.
The defensible pattern is a mix: real conversational audio anchors the acoustic and pragmatic distribution, and synthetic speech adds controlled coverage of rare intents or languages. Keep synthetic share as a tracked field so you can ablate it.
Where real spoken dialogue comes from
Real, licensable spontaneous dialogue is scarce, which is why public options cluster around read speech or small conversational releases. The CASPER authors make this point directly: their English conversational dataset records 200 hours, of which 158 hours are actual speech, and the public release covers 102 hours [4]. Crowdsourced corpora such as Common Voice are large and public domain (CC0), but they consist mostly of read prompts, not conversations [11].
Open full-duplex conversational datasets have started to appear because interactive speech synthesis needs overlaps and turn-taking that read speech never contains [5]. Before you train on any open corpus, audit its terms. A large audit of text datasets on popular hosting sites found license omissions above 70% and license error rates above 50% [10]. Our license audit for open speech corpora walks through that check.
The deepest pool of natural dialogue sits inside businesses: contact-center recordings, sales calls, field-service calls, dictation and recorded hands-on work. See call center audio datasets and voice agent training data for those categories. Operational recordings come with real intents and outcomes, which makes them convertible into instruction pairs, similar to turning business records into instruction-response pairs.
Specifying audio-text pairs a model can learn from
A usable specification fixes the audio format, the transcript standard, the alignment granularity and the label schema before any sample arrives. Ambiguity in any of these turns into rework or silent label noise.
Audio. State sample rate, bit depth, channel layout and codec. Telephony audio at 8 kHz mu-law behaves differently from 16 kHz or 48 kHz wideband capture, and stereo with one speaker per channel makes diarization and overlap modeling far easier than a mixed mono track. Details are in audio file specs for speech datasets.
Transcripts. Choose verbatim or clean transcription explicitly. Speech LLMs that must handle hesitation usually need verbatim tags for fillers, repairs and non-speech events; see verbatim vs clean transcription standards.
Alignment. Ask for segment-level timestamps at minimum and word-level timing where available, in a declared format such as JSONL with start and end seconds, CTM or TextGrid. Turn boundaries and overlap regions should be explicit, not inferred from silence.
Labels. If you combine sources, expect to harmonize label sets. Emotion taxonomies, intent schemas and speaker-role labels differ between corpora, and the empathetic speech-LLM work above had to reconcile labels across several datasets [3]. Write a mapping table before mixing.
Illustrative instruction-pair record
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"pair_id": "sess_0412_t07",
"session_id": "sess_0412",
"turn_index": 7,
"audio_uri": "audio/sess_0412.flac",
"channel": "caller_left",
"sample_rate_hz": 16000,
"segment": {"start_s": 183.42, "end_s": 191.10},
"speaker_role": "customer",
"transcript_verbatim": "um so the the invoice shows two charges for, uh, March",
"transcript_clean": "The invoice shows two charges for March.",
"instruction_text": "Explain why the customer sees two March charges and what to do.",
"response_text": "One charge is the prorated upgrade; the other is the regular monthly fee.",
"response_audio_uri": "audio/sess_0412.flac#agent_right:191.4-204.9",
"paralinguistic": {"emotion": "frustrated", "rate": "fast", "overlap_with_prev": true},
"synthetic": false,
"deid": {"method": "transcript_tag_and_audio_tone", "pii_spans_replaced": 2},
"license_ref": "LIC-REF-PLACEHOLDER"
}
The synthetic flag, the de-identification block and the license reference are the fields buyers most often forget, and they are the ones auditors ask about later.
Voice consent, biometrics and de-identification
Speech-LLM training uses voices at scale, so every source needs a consent and biometric review, not just a copyright check. Illinois BIPA lists voiceprints among biometric identifiers and requires written notice and a written release before collecting them [6]. A 2024 amendment treats repeated collection from the same person as a single violation, which reduces but does not remove exposure [7].
Beyond statute, researchers catalog long-tail risks for speakers whose voices end up in AI data: privacy, reputation, accountability, consent, credit and compensation [8]. For recordings created for training, consent language should name model training, speech generation and retention explicitly; see consent language for commissioned collection.
Spoken audio also carries personal details: names, account numbers, addresses and card digits said aloud. Redaction in audio and transcript must stay in sync, and the choice of silence, tone or tag changes what the model learns; see how audio redaction affects speech model training. Health conversations, such as clinical calls or dictation, need HIPAA de-identification by Safe Harbor or Expert Determination [9].
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Buyer checklist for speech-LLM data requests
A request that answers the questions below lets a supplier say quickly whether they hold matching data and lets your reviewers approve it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Area | What to specify | Failure mode if skipped |
|---|---|---|
| Training stage | Pre-training alignment, instruction tuning, speech-to-speech or eval | Wrong pairing granularity; unusable for SFT |
| Speech type | Spontaneous dialogue, task requests, dictation; synthetic share cap | Model fits TTS prosody, fails on real callers |
| Turn structure | Full sessions, overlaps kept, backchannels marked | Model cannot handle interruptions |
| Audio spec | Sample rate, channels per speaker, codec, source channel | Channel mismatch at inference |
| Transcript | Verbatim or clean, tag set, word or segment timing | Silent label noise across vendors |
| Labels | Intent, emotion, role taxonomies with mapping table | Unharmonized labels across corpora |
| Coverage | Languages, accents, age bands, speaker counts, hours per speaker | A few heavy speakers dominate the set |
| Rights | Ownership, speaker consent scope, biometric notices | Unlicensable voices discovered post-training |
| Privacy | PII method in audio and text, sample check, health data rules | Spoken account numbers leak into weights |
| Holdout | Speaker-disjoint eval split | Inflated scores from speaker overlap |
For multilingual targets, also read multilingual ASR training data. For speech-to-speech systems that must speak while listening, see full-duplex conversation data. The speech and audio data hub covers the rest of the cluster.
How SourceX fits a speech-LLM data request
SourceX sources operational datasets from US companies on request, including support and sales histories, documents and new recordings of hands-on work, and it does not hold inventory, so a request does not guarantee a match. It does not source scraped web content. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. You can describe the spoken data you need to SourceX without naming specific businesses.
Source real spoken dialogue for your speech-language model
SourceX looks for US businesses that hold the conversational audio and transcripts you describe, then runs assessment of data and licensing permissions, agreement on pricing and allowed uses, and delivery through private, access-controlled workflows after an executed agreement. Nothing is contracted until a supplier agrees. Start a buyer request at SourceX.
Sources
- arXiv, "How to Leverage Synthetic Speech for LLM-Based ASR Systems?" (2026). https://arxiv.org/abs/2606.29031v2
- arXiv, "Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap" (2025). https://arxiv.org/pdf/2509.26542
- arXiv, "BLSP-Emo: Towards Empathetic Large Speech-Language Models" (2024). https://arxiv.org/pdf/2406.03872
- arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
- arXiv, "Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis" (2025). https://arxiv.org/pdf/2509.04093
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14)". https://ilga.gov/Legislation/ILCS/Articles?ActID=3004&ChapterID=57&Print=True
- Davis Wright Tremaine, "Illinois Revises Biometrics Law To Reduce the Prospect of Ruinous Damage Awards" (2024). https://dwt.com/blogs/privacy--security-law-blog/2024/08/illinois-bipa-biometrics-law-amended-for-damages
- arXiv, "PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data" (2025). https://arxiv.org/pdf/2507.16247
- U.S. Department of Health and Human Services, "Guidance Regarding Methods for De-identification of Protected Health Information". https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Ardila et al., Mozilla (arXiv), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.