Speech and audio data
Verbatim vs Clean Transcription Standards for Speech Training Data
Quick answer
Use full verbatim when the model must hear and reproduce what was actually said: disfluency detection, conversational ASR, eval references that score fillers, and TTS alignment. Use clean verbatim when the target is readable text and you will score after normalization. Either works for training; mixing them silently does not. Write one guideline that fixes fillers, false starts, numbers, casing and tags, version it, and ship it with every delivered transcript so the error floor stays measurable.
By SourceX Editorial · Updated
What verbatim, clean verbatim and edited transcripts actually contain
The three styles differ in what happens to disfluencies, not in accuracy: a perfect clean verbatim transcript and a perfect verbatim transcript of the same call are different strings. Industry practice distinguishes a verbatim style that keeps what was spoken from a clean style that produces readable text [1]. Edited transcripts go further and rewrite grammar, which makes them a poor ASR target unless you deliberately train a speech-to-written model (see training ASR on non-verbatim transcripts).
| Phenomenon | Full verbatim | Clean verbatim | Edited |
|---|---|---|---|
| Fillers (um, uh, er) | Kept, from a closed list | Removed | Removed |
| Discourse markers (like, you know) | Kept | Kept when meaningful, removed as filler | Usually removed |
| False starts ("I wa- I want") | Kept, with truncation marker | Removed, final attempt kept | Removed |
| Repetitions ("the the") | Kept | Collapsed | Collapsed |
| Backchannels (mm-hmm, yeah) | Kept as separate turns | Often dropped | Dropped |
| Ungrammatical speech | Kept | Kept | Corrected |
| Non-speech events | Tagged | Tagged selectively or dropped | Dropped |
The practical decision is which column your downstream task needs. A disfluency detection model needs the left column plus span labels; a call summarization pipeline may be happier with the middle one.
Why mixed conventions put a floor under your error rate
Inconsistent conventions create errors no model can remove, because the reference itself disagrees with itself. A study of the Faetar ASR benchmark examined how inconsistent transcription conventions, such as phonetic versus lexical spellings, add error that better modeling cannot remove [3]. Research from Rev transcribed the same audio under opposing style parameters and showed that the style choice alone changes the evaluation result, and that datasets differ in style and formality enough to distort scores [2].
The failure mode in a licensed corpus is usually quieter: three supplier teams, one keeping "uh" and two dropping it, all labeled "verbatim." The model learns to emit fillers at roughly the average rate, and every eval reference punishes it either way. The same authors argue that standard single-reference WER likely overstates the contentful errors of strong current systems [2], which is why a style mismatch can look like a model regression.
Choosing a standard by use case
Pick the style from the job the transcript does, then hold every supplier and annotator to it. Use this as a starting point, not a rule.
- ASR training, conversational domains (contact center, meetings): full verbatim for fillers and false starts, so the acoustic model learns to align them rather than hallucinate words into them. Normalize at training time if you want clean output.
- Eval references: match the style your production output targets, or keep verbatim and normalize both hypothesis and reference before scoring. Tools such as BenchmarkSTT apply configurable normalization before computing metrics [4].
- TTS alignment and voice data: verbatim, including breaths and lengthened words where you want them synthesized; clean text misaligns forced alignment.
- Disfluency detection data: verbatim plus span annotation (reparandum, interregnum, repair) as a separate layer, never inline edits to the transcript.
- LLM text from transcripts: clean verbatim is often enough; see speech transcripts as LLM training text.
Budget for the choice. Accurate verbatim reference transcripts are time-consuming and expensive to produce [4], so many teams buy verbatim for a held-out eval set and clean verbatim for bulk training hours, with the mapping between them documented.
Rules for numbers, dates, currency, acronyms and casing
Decide whether transcripts are spoken-form or written-form, and say so in the first line of the guideline. Spoken form ("twenty twenty six", "three fifty") preserves what was said and avoids ambiguity about "350" versus "three hundred fifty"; written form matches what most products display but needs an inverse text normalization step that the transcript cannot verify. Whichever you pick, write examples for each entity type your domain uses (see domain vocabulary coverage).
- Numbers and dates: one form per entity type; never both in the same corpus without a tag.
- Currency and units: "five dollars" or "$5", plus a rule for "five bucks."
- Acronyms and spelled letters: "IBM" vs "I B M" when spelled out letter by letter; a rule for initialisms said as words.
- Names and products: a project lexicon with canonical spellings; unknown names marked for review rather than guessed.
- Punctuation and casing: either lowercase and unpunctuated, or truecased with sentence punctuation; mixed corpora force you to strip both before scoring anyway.
- Hyphenation, contractions, compounds: "gonna" vs "going to" is a style decision; list the allowed reduced forms.
Tags for non-speech, overlap, redaction and language switches
A tag set is part of the standard, and every tag needs a written trigger, a closed spelling and a rule for how scoring treats it. Keep tags in a reserved syntax (square brackets are common) that never collides with spoken words, and publish a mapping so evaluators can strip or keep them consistently.
- Unintelligible:
[unintelligible]with a minimum duration rule, distinct from[inaudible]for low signal. - Uncertain words: a best guess wrapped in a marker so it can be excluded from references.
- Non-speech:
[noise],[music],[laughter],[cough], with a rule for speech over noise (see noise and SNR coverage). - Overlap: per-speaker segments with timestamps rather than inline interleaving (see overlapping speech data).
- Redaction placeholders: typed tokens such as
[NAME]or[ACCOUNT_NUMBER]aligned to the muted or toned audio span (see audio redaction artifacts). - Language switches: a language tag per segment, plus a rule for borrowed words.
A guideline header you can hand to a supplier
The guideline should fit on a page at the top and point to a longer appendix of examples. The header below is what reviewers and annotators should read before transcribing their first file.
Illustrative example: invented to show structure; it does not describe an available dataset.
guideline_id: tx-std-contactcenter-en-us
version: 1.3.0
effective_date: 2026-10-01
style: full_verbatim
text_form: spoken # numbers, dates, currency spelled as spoken
casing: lowercase
punctuation: none
fillers:
keep: true
closed_list: [um, uh, er, hmm]
false_starts: keep_with_marker # "i wa- i want"
repetitions: keep
backchannels: separate_turn
reduced_forms_allowed: [gonna, wanna, kinda]
tags:
unintelligible: "[unintelligible]"
noise: "[noise]"
laughter: "[laughter]"
redaction: "[NAME] [PHONE] [EMAIL] [ACCOUNT_NUMBER]"
language_switch: "<lang=es>...</lang>"
segmentation: per_speaker_with_timestamps
lexicon_file: lexicon_v7.tsv
scoring_normalization: normalize_v2.yaml
change_log: "1.3.0: backchannels moved to separate turns"
And one illustrative record showing the same utterance in both styles, so annotators can see the delta:
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"segment_id": "c0042_s017",
"speaker": "agent",
"start": 132.48,
"end": 137.91,
"verbatim": "um so i wa- i want to move the the payment to [NAME]'s card uh ending in four two",
"clean_verbatim": "So I want to move the payment to [NAME]'s card ending in 42.",
"guideline_version": "1.3.0"
}
Documenting the standard so the dataset stays usable
Ship the guideline with the data, versioned, and record which version produced each file. The Rev authors make the case that evaluation should account for transcription style explicitly rather than assume one canonical reference [2], and that is only possible if the style is written down. Put the guideline ID and version in the dataset card, which exists to promote responsible use and inform users of potential biases [7], and in a data statement covering speakers, varieties and annotator background [5][6].
A buyer checklist for any transcribed speech delivery:
- The guideline document and change log, not a one-line style label.
- A per-file
guideline_versionfield, with re-transcribed files flagged. - The normalization script used for scoring, so your WER matches the supplier's.
- Inter-annotator agreement on a double-transcribed sample, broken out by tag.
- The project lexicon and the rule for unknown names.
- A note on how redaction tags map to audio edits.
Acceptance testing against the guideline is its own process; see acceptance testing supplier transcripts. For the wider buying picture, start at the speech and audio data hub. Definitions of a call transcript and of data annotation are in the SourceX glossary.
Where licensed operational transcripts fit
Recorded support and sales calls already carry transcripts, but they were produced to whatever standard the business used, often a vendor's clean or edited style. If you describe the data you need to SourceX, including the transcription standard, SourceX looks for US businesses that hold it and manages the licensing process; datasets are sourced on request, and a request does not guarantee a match. Personal details such as names, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.
Sourcing speech data with a documented transcription standard
SourceX sources operational datasets, including support and sales histories, from US companies and manages licensing and ongoing purchases. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and nothing is contracted until a supplier agrees. Tell SourceX what speech data and transcription standard you need.
Frequently asked questions
Should fillers be kept in ASR training transcripts?
Keep them when the model must handle conversational speech or detect disfluencies, because dropping them leaves acoustic frames with no target text. If your product shows clean text, keep fillers in training data and remove them in post-processing or through normalization at scoring time.
Can I combine verbatim and clean verbatim corpora?
Yes, if every segment carries a style label and you either convert one to the other with a documented rule or condition the model on the label. Unlabeled mixing makes the reference inconsistent, which is the kind of error floor described above.
Does clean verbatim lower WER?
It changes WER rather than improving the model. A clean hypothesis scored against a verbatim reference loses points for every dropped filler, and the reverse is also true [2]. Normalize both sides with the same script before comparing systems.
Sources
- 3Play Media, "The 2024 State of Automatic Speech Recognition Report" (2024). https://go.3playmedia.com/hubfs/WP%20PDFs/The%202024%20State%20of%20Automatic%20Speech%20Recognition%20Report_Final_Remediated-1.pdf
- arXiv, "Style-agnostic evaluation of ASR using multiple reference transcripts" (2024). https://arxiv.org/pdf/2412.07937
- arXiv, "Investigating Transcription Normalization in the Faetar ASR Benchmark" (2025). https://arxiv.org/pdf/2508.11771
- BenchmarkSTT, "BenchmarkSTT 1.0.0 tutorial". https://benchmarkstt.readthedocs.io/en/1.0.0/tutorial.html
- Transactions of the ACL, "Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science" (2018). https://aclanthology.org/Q18-1041/
- University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/
- Hugging Face, "Create a dataset card". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.