Text and language data
Using Speech Transcripts as LLM Training Text: Quality and Normalization
Quick answer
Speech transcripts make good LLM training text when you treat them as a derived artifact, not as plain prose. Before training, decide between verbatim and clean-read style, keep speaker turns, timestamps and disfluency flags as structured fields, and record how each segment was produced: ASR engine and version, estimated error rate, and whether a human verified it. Then de-identify names and numbers and deduplicate. A transcript with that metadata can feed pre-training, SFT or summarization. Without it, you inherit recognition errors you cannot measure.
By SourceX Editorial · Updated
Why transcripts are worth using as text at all
Transcripts capture registers that written corpora under-represent: turn-taking, clarification, interruption, negotiation and spoken reasoning. Openly licensed pretraining collections already include audio transcripts as a source type [1]. For an applied team, the attraction is domain dialogue that never existed as edited text: support calls, sales discovery calls, engineering stand-ups and expert interviews.
The catch is that a transcript has two authors, the speakers and the transcription process. Every substitution, dropped negation or misheard product name from the ASR system becomes a "fact" in your training text. If you train on audio and text together, see training data for speech-language models. This page covers transcript text alone. For the record types themselves, the owner pages on meeting transcripts and sales call transcripts describe what those collections contain.
Verbatim or clean read: pick per use, keep both if you can
The right transcription style depends on the downstream task, so the best practice is to keep a verbatim layer and derive a clean layer from it. Verbatim transcripts keep fillers ("um", "uh"), repetitions, false starts, self-corrections and non-speech tags. Clean-read (or "intelligent verbatim") transcripts remove them and repair grammar lightly.
Style is not a cosmetic choice. Research on ASR evaluation shows that word error rate shifts substantially with transcription conventions, so two transcripts of the same audio in different styles look "wrong" against each other [3]. Normalization rules alone can move measured error rates on a benchmark [4].
| Downstream use | Preferred layer | Why |
|---|---|---|
| Pre-training on conversational text | Light normalization of verbatim | Keeps the distribution of real spoken language; drop only non-speech tags |
| SFT for a dialogue or voice agent | Verbatim with disfluency flags | Model must handle restarts and repairs in user turns |
| Summarization or action-item extraction | Clean read as input; verbatim retained for audit | Fillers inflate tokens without adding content |
| Assistant-turn targets in SFT | Clean read | You rarely want the model to imitate "um, so, like" in its own output |
| Evaluation sets | Both, with documented style guide | Lets you score against either convention |
Disfluencies are signal for some tasks and noise for others. A self-correction such as "the renewal is in March, sorry, May" carries the true answer only in the repair. Naive filler removal keeps "March" and loses the correction, which teaches the model a wrong fact.
The structured record to require for each segment
Transcripts should arrive as segment-level records, not as flat text files, because the metadata is what lets you filter, weight and audit. A JSON Lines file with one row per utterance, or WebVTT and SRT with a sidecar manifest, both work; a .txt dump with speaker names in prose does not. The general field set for licensed text is covered in metadata fields to require with licensed text corpora. Transcripts need several more.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"conversation_id": "conv_8f21",
"segment_id": "conv_8f21_0042",
"speaker_turn": 42,
"speaker_role": "agent",
"speaker_pseudonym": "SPK_A",
"start_ms": 183420,
"end_ms": 189110,
"text_verbatim": "uh so the the renewal is in March, sorry, May",
"text_clean": "So the renewal is in May.",
"disfluency_spans": [[0, 2, "filler"], [6, 13, "repetition"], [28, 33, "reparandum"]],
"overlap": false,
"transcript_source": "asr",
"asr_engine": "vendor-model-x",
"asr_version": "2025-11-03",
"segment_confidence": 0.87,
"human_verified": false,
"est_wer_bucket": "10-20%",
"language": "en-US",
"deid_method": "ner_plus_regex_v3",
"deid_reviewed_sample": true,
"channel": "telephony_8khz"
}
The fields that matter most and are most often missing:
- transcript_source and human_verified. Separates ASR output, human transcription and human-corrected ASR. Weight or filter on it.
- asr_engine and asr_version. Error profiles change between model versions. Without the version, you cannot tell whether a batch came from an older, noisier engine.
- Confidence or estimated WER. Per-segment confidence lets you drop the worst tail. A corpus-level WER measured on a human-verified sample tells you what to expect.
- Speaker turns and roles. Diarization errors that merge two speakers corrupt dialogue structure; role labels (agent, customer, interviewer) are what make SFT formatting possible.
- Timestamps and overlap. Needed to rebuild turn order and to flag crosstalk, where ASR quality degrades most.
- channel or audio condition. Narrowband telephony, far-field meeting rooms and studio podcasts produce very different error rates.
Measuring transcript quality before you buy
Ask for word error rate on a human-verified sample drawn from the same collection, with the normalization rules stated. WER compares a hypothesis to a reference after both are normalized; casing, punctuation, number formats and contractions all change the score [5]. The Whisper authors built a dedicated text normalizer for exactly this reason, so that benchmark WER reflected recognition errors rather than formatting differences [2].
Three checks catch most problems in a sample of a few hundred segments:
- Domain term accuracy. Pull every product name, part number, drug name or legal term from the sample and check it against the audio or a human reference. General WER hides the fact that rare domain terms fail disproportionately, and those are often the tokens you are buying the data for.
- Negation and number fidelity. Search for "not", "can't", "no" and digit strings; compare against the reference. A dropped "not" inverts meaning while counting as a single word error.
- Hallucinated text in silence or noise. Some ASR systems emit fluent text over music, hold tones or silence. Look for repeated phrases at segment boundaries and segments with text but very low audio energy.
Also check for machine-generated "human" transcripts. The Whisper team found that weakly supervised transcripts on the internet included ASR output and used heuristics to remove it from training data [2]. The same risk applies when a supplier labels a collection "human transcribed": ask what share was ASR-first with human review, and how much review each segment received.
Normalization pipeline for transcript text
Normalize in a fixed, logged order so the transformation is reproducible and reversible to the verbatim layer. A practical sequence:
Illustrative example: invented to show structure; it does not describe an available dataset.
| Step | Operation | Failure mode it prevents |
|---|---|---|
| 1 | Validate schema, timestamps and segment order | Shuffled or truncated conversations |
| 2 | Drop or tag non-speech markers ([music], [crosstalk], [inaudible]) | Bracket tokens leaking into model output |
| 3 | De-identify (see next section) before any text leaves the secure zone | Personal data copied into derived layers |
| 4 | Filter by confidence, estimated WER bucket and channel | Training on the noisiest tail |
| 5 | Mark disfluencies; produce clean layer with repairs applied correctly | Keeping the reparandum instead of the correction |
| 6 | Restore casing and punctuation if ASR output lacks it | Run-on lowercase text that hurts tokenization and readability |
| 7 | Standardize numbers, dates and currency with one documented convention | "twenty twenty six", "2026" and "20 26" as three variants |
| 8 | Exact and near-duplicate removal across conversations | Hold music scripts, IVR prompts and compliance disclosures repeated thousands of times |
| 9 | Format for the target: plain text, chat template turns, or document plus summary pairs | Speaker roles lost in final formatting |
Deduplication matters more for transcripts than for most text. Call centers repeat scripted greetings, recorded disclosures and hold messages across every call. Near-duplicate examples and long repeated substrings are common in LM datasets, and removing them reduces verbatim memorization [6]. Remove or downsample templated segments before computing token counts for licensing, or you will pay for and train on boilerplate.
De-identifying spoken names and numbers
Transcripts leak personal data in forms written-text tools miss, so de-identification needs speech-specific rules. People spell names letter by letter, read account numbers digit by digit, say email addresses as "jane dot doe at" and give dates of birth in words. A regex for a 16-digit card number will not match "four one one one, two two two two" split across two ASR segments.
Run NER and pattern detection on both layers, merge adjacent segments before matching spelled-out sequences, and replace with consistent pseudonyms per conversation so dialogue coherence survives. Open-source tools such as Presidio help, but the project itself warns that ML-based detection cannot guarantee it finds all sensitive information [7]. Test a reviewed sample and record the method. For health-related calls, such as provider-to-payer or patient lines, recordings that are protected health information held by a covered entity or business associate must meet HIPAA de-identification under 45 CFR 164.514(a)-(b), by Safe Harbor or Expert Determination [8]. Residual risk in de-identified text is rising as LLMs get better at linkage; see LLM-assisted re-identification of de-identified text.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Rights and provenance questions specific to recorded speech
Transcripts carry consent questions from the original recording, so rights review has to reach back to how the audio was captured. Ask who recorded the call or meeting, under what notice, and whether the participants' consent or the company's agreement covers derivative text used for model training. Two-party consent laws, employee monitoring policies and customer terms of service all shape what the holder can license.
Dataset documentation is often thin. An audit of more than 1,800 text datasets found license information omitted in over 70% of cases on popular hosting sites [10]. For transcripts, document speaker demographics, dialect, speech situation and recording conditions; the Data Statements schema provides fields for exactly these speaker and speech-situation properties [9]. Whether the source company holds rights to license the material is a separate question from whether the text is clean.
Where SourceX fits
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Relevant categories include support and sales histories, engineering records and new recordings of hands-on work. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.
Before delivery, personal details such as names, emails, phones and account numbers are removed or replaced, the method is recorded and a sample is checked; no method is perfect. Health records require HIPAA de-identification by Safe Harbor or Expert Determination. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Buyers can describe the transcript data they need, including the segment fields above, and SourceX looks for US businesses that hold it. More context on the broader category sits in the text and language data hub, and on terse internal writing in operational free-text notes.
Sourcing transcripts as LLM training data
If you need conversational transcripts from US businesses with the segment metadata and de-identification described here, describe the data rather than the companies. SourceX assesses data and licensing permissions with the holder, and nothing is contracted until a supplier agrees. Start with a request at sourcex.si/buyers.
Frequently asked questions
Should I keep fillers like "um" and "uh" in pre-training text?
Keep them for user-side turns if you want robustness to spoken input, and tag them so you can ablate. For assistant targets and summaries, remove them. Measuring both variants on a held-out conversational eval is cheaper than arguing about it.
Is ASR output good enough, or do I need human transcripts?
It depends on error profile, not headline WER. Mixed corpora work when each segment carries a humanverified flag and a confidence score, so you can upweight verified data and filter the noisy tail. Glossary background: speech-to-text and call transcript.
Which quality metric should I put in the purchase spec?
Specify WER on a stratified human-verified sample with named normalization rules, plus domain-term accuracy and negation fidelity on the same sample. Tie these to the dimensions in training data quality metrics.
Sources
- arXiv, "The Common Pile v0.1 (arXiv:2506.05209)" (2025). https://arxiv.org/html/2506.05209v1
- Radford et al., OpenAI (arXiv:2212.04356), "Robust Speech Recognition via Large-Scale Weak Supervision" (2022). https://arxiv.org/pdf/2212.04356
- arXiv (arXiv:2412.07937), "Style-agnostic evaluation of ASR using multiple reference transcripts" (2024). https://arxiv.org/pdf/2412.07937
- arXiv (arXiv:2508.11771), "Investigating Transcription Normalization in the Faetar ASR Benchmark" (2025). https://arxiv.org/pdf/2508.11771
- Speechmatics documentation, "Calculating WER". https://docs.speechmatics.com/tutorials/calculating-wer
- Lee et al. (arXiv:2107.06499; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/
- Longpre et al. (arXiv:2310.16787), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.