Skip to content

Speech and audio data

Training Data for Speech-Language Models: Audio-Text Instruction and Dialogue Pairs

Quick answer

Speech-language models need three kinds of data that text corpora and read-speech ASR sets do not supply: long-form conversational audio with aligned transcripts for pre-training, spoken instructions paired with responses for instruction tuning, and multi-turn dialogue with real turn-taking for speech-to-speech behavior. TTS-synthesized prompts fill gaps cheaply but generalize poorly to natural speech [1]. Buyers should specify speaker diversity, channel, transcript standard, timing alignment and voice consent before evaluating any source.

By SourceX Editorial · Updated

What a speech-language model actually consumes

A speech-language model consumes paired audio and text at several granularities, and each training stage wants a different pairing. Unlike an ASR model, which only maps audio to a transcript (see speech-to-text), a speech LLM has to understand spoken requests, reason over them and often answer in speech. That makes it a multimodal data consumer whose text side carries intent, not just words.

In practice, teams source four data types:

  • Transcribed conversational audio for continued pre-training or audio-text alignment: hours of two-party or multi-party speech with time-aligned transcripts, speaker diarization and channel metadata.
  • Spoken instruction-response pairs for instruction tuning: a user utterance in audio, plus a text or audio response that is correct for the request.
  • Multi-turn spoken dialogue for conversational and speech-to-speech behavior: full sessions with overlaps, backchannels, interruptions and silences preserved.
  • Paralinguistic labels such as emotion, speaking rate, hesitation or frustration, used for empathetic or style-aware responses. One empathetic speech-LLM effort assembled about 70,000 emotion-labeled utterances from several corpora to get enough coverage [3].

Why TTS-synthesized prompts are a stopgap

Synthetic speech is useful for bootstrapping instruction data, but models trained mostly on it break on real callers. Recent speech-LLM work notes that reliance on TTS-synthesized speech limits generalization to natural human speech [1]. TTS output lacks disfluencies, false starts, background noise, codec damage, accent variation and the prosodic cues people use to signal uncertainty or urgency.

A second gap shows up at evaluation. Researchers have diagnosed a modality-induced performance gap: the same reasoning tasks score worse when posed by voice than by text [2]. Real spoken task data, where people actually ask for things out loud, is one lever for closing that gap. For a broader comparison, see real vs synthetic speech data for ASR and voice models.

The defensible pattern is a mix: real conversational audio anchors the acoustic and pragmatic distribution, and synthetic speech adds controlled coverage of rare intents or languages. Keep synthetic share as a tracked field so you can ablate it.

Where real spoken dialogue comes from

Real, licensable spontaneous dialogue is scarce, which is why public options cluster around read speech or small conversational releases. The CASPER authors make this point directly: their English conversational dataset records 200 hours, of which 158 hours are actual speech, and the public release covers 102 hours [4]. Crowdsourced corpora such as Common Voice are large and public domain (CC0), but they consist mostly of read prompts, not conversations [11].

Open full-duplex conversational datasets have started to appear because interactive speech synthesis needs overlaps and turn-taking that read speech never contains [5]. Before you train on any open corpus, audit its terms. A large audit of text datasets on popular hosting sites found license omissions above 70% and license error rates above 50% [10]. Our license audit for open speech corpora walks through that check.

The deepest pool of natural dialogue sits inside businesses: contact-center recordings, sales calls, field-service calls, dictation and recorded hands-on work. See call center audio datasets and voice agent training data for those categories. Operational recordings come with real intents and outcomes, which makes them convertible into instruction pairs, similar to turning business records into instruction-response pairs.

Specifying audio-text pairs a model can learn from

A usable specification fixes the audio format, the transcript standard, the alignment granularity and the label schema before any sample arrives. Ambiguity in any of these turns into rework or silent label noise.

Audio. State sample rate, bit depth, channel layout and codec. Telephony audio at 8 kHz mu-law behaves differently from 16 kHz or 48 kHz wideband capture, and stereo with one speaker per channel makes diarization and overlap modeling far easier than a mixed mono track. Details are in audio file specs for speech datasets.

Transcripts. Choose verbatim or clean transcription explicitly. Speech LLMs that must handle hesitation usually need verbatim tags for fillers, repairs and non-speech events; see verbatim vs clean transcription standards.

Alignment. Ask for segment-level timestamps at minimum and word-level timing where available, in a declared format such as JSONL with start and end seconds, CTM or TextGrid. Turn boundaries and overlap regions should be explicit, not inferred from silence.

Labels. If you combine sources, expect to harmonize label sets. Emotion taxonomies, intent schemas and speaker-role labels differ between corpora, and the empathetic speech-LLM work above had to reconcile labels across several datasets [3]. Write a mapping table before mixing.

Illustrative instruction-pair record

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "sess_0412_t07",
  "session_id": "sess_0412",
  "turn_index": 7,
  "audio_uri": "audio/sess_0412.flac",
  "channel": "caller_left",
  "sample_rate_hz": 16000,
  "segment": {"start_s": 183.42, "end_s": 191.10},
  "speaker_role": "customer",
  "transcript_verbatim": "um so the the invoice shows two charges for, uh, March",
  "transcript_clean": "The invoice shows two charges for March.",
  "instruction_text": "Explain why the customer sees two March charges and what to do.",
  "response_text": "One charge is the prorated upgrade; the other is the regular monthly fee.",
  "response_audio_uri": "audio/sess_0412.flac#agent_right:191.4-204.9",
  "paralinguistic": {"emotion": "frustrated", "rate": "fast", "overlap_with_prev": true},
  "synthetic": false,
  "deid": {"method": "transcript_tag_and_audio_tone", "pii_spans_replaced": 2},
  "license_ref": "LIC-REF-PLACEHOLDER"
}

The synthetic flag, the de-identification block and the license reference are the fields buyers most often forget, and they are the ones auditors ask about later.

Speech-LLM training uses voices at scale, so every source needs a consent and biometric review, not just a copyright check. Illinois BIPA lists voiceprints among biometric identifiers and requires written notice and a written release before collecting them [6]. A 2024 amendment treats repeated collection from the same person as a single violation, which reduces but does not remove exposure [7].

Beyond statute, researchers catalog long-tail risks for speakers whose voices end up in AI data: privacy, reputation, accountability, consent, credit and compensation [8]. For recordings created for training, consent language should name model training, speech generation and retention explicitly; see consent language for commissioned collection.

Spoken audio also carries personal details: names, account numbers, addresses and card digits said aloud. Redaction in audio and transcript must stay in sync, and the choice of silence, tone or tag changes what the model learns; see how audio redaction affects speech model training. Health conversations, such as clinical calls or dictation, need HIPAA de-identification by Safe Harbor or Expert Determination [9].

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Buyer checklist for speech-LLM data requests

A request that answers the questions below lets a supplier say quickly whether they hold matching data and lets your reviewers approve it.

Illustrative example: invented to show structure; it does not describe an available dataset.

AreaWhat to specifyFailure mode if skipped
Training stagePre-training alignment, instruction tuning, speech-to-speech or evalWrong pairing granularity; unusable for SFT
Speech typeSpontaneous dialogue, task requests, dictation; synthetic share capModel fits TTS prosody, fails on real callers
Turn structureFull sessions, overlaps kept, backchannels markedModel cannot handle interruptions
Audio specSample rate, channels per speaker, codec, source channelChannel mismatch at inference
TranscriptVerbatim or clean, tag set, word or segment timingSilent label noise across vendors
LabelsIntent, emotion, role taxonomies with mapping tableUnharmonized labels across corpora
CoverageLanguages, accents, age bands, speaker counts, hours per speakerA few heavy speakers dominate the set
RightsOwnership, speaker consent scope, biometric noticesUnlicensable voices discovered post-training
PrivacyPII method in audio and text, sample check, health data rulesSpoken account numbers leak into weights
HoldoutSpeaker-disjoint eval splitInflated scores from speaker overlap

For multilingual targets, also read multilingual ASR training data. For speech-to-speech systems that must speak while listening, see full-duplex conversation data. The speech and audio data hub covers the rest of the cluster.

How SourceX fits a speech-LLM data request

SourceX sources operational datasets from US companies on request, including support and sales histories, documents and new recordings of hands-on work, and it does not hold inventory, so a request does not guarantee a match. It does not source scraped web content. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. You can describe the spoken data you need to SourceX without naming specific businesses.

Source real spoken dialogue for your speech-language model

SourceX looks for US businesses that hold the conversational audio and transcripts you describe, then runs assessment of data and licensing permissions, agreement on pricing and allowed uses, and delivery through private, access-controlled workflows after an executed agreement. Nothing is contracted until a supplier agrees. Start a buyer request at SourceX.

Sources

  1. arXiv, "How to Leverage Synthetic Speech for LLM-Based ASR Systems?" (2026). https://arxiv.org/abs/2606.29031v2
  2. arXiv, "Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap" (2025). https://arxiv.org/pdf/2509.26542
  3. arXiv, "BLSP-Emo: Towards Empathetic Large Speech-Language Models" (2024). https://arxiv.org/pdf/2406.03872
  4. arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
  5. arXiv, "Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis" (2025). https://arxiv.org/pdf/2509.04093
  6. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14)". https://ilga.gov/Legislation/ILCS/Articles?ActID=3004&ChapterID=57&Print=True
  7. Davis Wright Tremaine, "Illinois Revises Biometrics Law To Reduce the Prospect of Ruinous Damage Awards" (2024). https://dwt.com/blogs/privacy--security-law-blog/2024/08/illinois-bipa-biometrics-law-amended-for-damages
  8. arXiv, "PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data" (2025). https://arxiv.org/pdf/2507.16247
  9. U.S. Department of Health and Human Services, "Guidance Regarding Methods for De-identification of Protected Health Information". https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  10. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  11. Ardila et al., Mozilla (arXiv), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data