Speech and audio data
Training ASR on Non-Verbatim Transcripts: Using Professional and Clean Transcripts as Targets
Quick answer
Yes, you can train ASR on non-verbatim transcripts, but not by feeding them in raw. Edited transcripts such as medical reports, broadcast captions and clean-read meeting notes drop fillers, fix grammar and sometimes rephrase, so the text no longer matches the audio word for word [1]. Treat them as light supervision: normalize the text, align it to the audio, keep only segments where the transcript and a biased decode agree, and measure the result on a separate verbatim test set [2][3].
By SourceX Editorial · Updated
This page covers reusing transcripts that already exist. If you are commissioning new transcription, the style decision is covered in verbatim vs clean transcription standards. For the wider cluster, start at the speech and audio data buyer's guide.
What changes between the audio and an edited transcript
An edited transcript differs from the audio in predictable, classifiable ways, and the kind of edit decides whether a segment is salvageable. Medical transcription is the textbook case: transcriptionists produce a formatted report, not a literal record, because verbatim work costs more and the report is what the clinic needs [1]. Captions and meeting minutes follow the same economics.
The edits fall into a short list. Each one produces a different failure in training:
- Deletions of disfluencies. "Um", "uh", false starts and repetitions are removed. A model trained on this learns to suppress fillers, which is often desirable, but alignment sees extra audio with no text.
- Rewrites and reordering. Dictated corrections ("scratch that, make it 20 milligrams") are applied silently. Rephrased sentences and moved clauses create text that never occurs in the audio.
- Inserted boilerplate. Report headers, section titles ("ASSESSMENT:"), templated normal findings and signature blocks were never spoken.
- Formatting conversions. "Twenty milligrams" becomes "20 mg"; dates, currency, drug names and abbreviations are written in house style.
- Truncation and condensing. Captions are often shortened to fit reading speed, and summaries drop entire turns.
- Speaker and timing loss. Captions may carry cue timestamps offset from speech; clean transcripts often lack timestamps entirely.
Found captions on web video are a weaker case still: HowTo100M's authors note that subtitles paired with clips are not manually annotated, are often ASR output and are frequently incomplete or out of sync with the visual content [5]. Ask whether the text you are buying was written by a person at all.
Which transcript layers are worth training on
A transcript layer is usable for ASR targets when most of its words were spoken in order and the edits are local. Summaries and templated reports are useful for language-model adaptation, but rarely as acoustic targets.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Layer a supplier may hold | Typical edits | Expected word-level match | Best use |
|---|---|---|---|
| Strict verbatim (with fillers, false starts) | None beyond spelling | High | Training targets and test references |
| Clean verbatim | Fillers and repetitions removed | High after alignment | Training targets after filtering |
| Broadcast or live captions | Condensing, paraphrase, timing drift | Medium, varies by program | Light supervision with segment selection |
| Edited professional report (dictation, legal, minutes) | Corrections applied, boilerplate, formatting | Low to medium | Text-only LM or biased-decode LM; select islands for acoustics |
| Summary or notes | Most content rewritten | Low | Not for acoustic targets |
| Prior ASR output | Recognizer errors | Unknown | Pseudo-labels only, never references |
The last row matters. If a "transcript" archive is really machine output from a previous recognizer, you are doing self-training on someone else's errors, which is a different and riskier technique.
How to align and filter found transcripts
The standard recipe is lightly supervised training: use the existing text to bias a decoder, compare, and keep what agrees. Large corpora have been built from audio that already had transcriptions by aligning and filtering it rather than retranscribing it [2]. The method dates to closed-caption work on broadcast news, where caption text was used to steer decoding rather than trusted as ground truth.
A practical pipeline for an incoming archive:
- Normalize text both ways. Expand numerals and units to spoken form for alignment ("20 mg" to "twenty milligrams"), strip headers and templated sections, and keep a reversible map so you can emit written form later.
- Build a biased language model. Interpolate a small n-gram or a prompt-conditioned model trained on the document's own transcript with a general background model, so the decoder favors, but is not forced to, the provided words.
- Decode and align. Run the biased decode on long-form audio, then align hypothesis to transcript with a word-level edit-distance alignment (or a CTC segmentation on a pretrained model) to get anchor words with timestamps.
- Cut at anchors. Split into segments of roughly 5 to 30 seconds at runs of matching words, so a rewrite in one sentence does not contaminate the next.
- Score each segment. Compute matched error rate between transcript and hypothesis, mean alignment confidence, and the share of transcript words that are boilerplate or formatting tokens.
- Filter and tier. Keep segments under a matched-error threshold as targets, route borderline segments to pseudo-labeling or human review, and drop the rest. Tune the threshold on a held-out set, not by intuition.
- Log the drop rate by stratum. Report retained hours by speaker, channel, domain and transcriber or caption vendor, because filtering silently removes the hardest audio.
That last step is where light supervision usually goes wrong. Accented speakers, noisy rooms and fast talkers produce worse baseline hypotheses, so a naive filter keeps clean, easy audio and throws away the coverage you were paying for. Compare the retained distribution against noise and SNR coverage targets and the speaker mix you intended.
Training on clean targets without breaking evaluation
Clean targets teach a model to write clean output, which is fine as long as you measure it against references with the same conventions. Research on transcript style shows that style parameters such as filler handling, number formatting and punctuation materially change measured error, so a model trained on clean text will look worse against verbatim references even when it heard the speech correctly [3].
Two controls keep the numbers honest. First, apply a text normalizer to both hypothesis and reference before scoring; the Whisper authors built one specifically so harmless formatting differences were not counted as errors, and checked it against an independently written normalizer to avoid tuning it to their own model [4]. Second, keep a small, separately produced verbatim test set from the same domain, with its own audit, because test references carry label errors too [6].
Other training choices to settle early:
- Mixing ratio. Combine filtered clean-target segments with any strictly verbatim data you have rather than training on clean targets alone; tune the mix on the verbatim test set.
- Style conditioning. If you need both verbatim and clean output, tag each segment with its source style so the model can be prompted for either at inference.
- Disfluency handling. Decide whether fillers should appear in output. If downstream consumers are summarizers or note generators, clean output may be what you want.
- Written-form output. If the transcripts use house style for numbers and drug names, an inverse text normalization step trained on that style may be more valuable than the acoustic data itself.
For sizing the archive against your fine-tuning goal, see how many hours of audio you need to fine-tune ASR.
Questions to put to a supplier before you license an archive
A supplier's archive of edited transcripts is only assessable if you know how each layer was produced. Send a short questionnaire before reviewing samples, and request a stratified sample of paired audio and text, not just text.
Illustrative example: invented to show structure; it does not describe an available dataset.
transcript_archive_questionnaire:
layers_held: # verbatim | clean_verbatim | captions | report | summary | asr_output
production_method: # human transcriptionist, editor pass, live captioner, ASR + edit
style_guide_available: # yes/no; attach filler, number, abbreviation and tag rules
timestamps: # none | per document | per caption cue | per word
speaker_labels: # none | role | diarized per turn
boilerplate_fields: # headers, templates, macros inserted by the editor
audio_format: # codec, sample rate, channels (mono, dual-channel)
audio_transcript_link: # file-level ID join; any orphaned audio or text
redaction: # audio masking method; transcript tags used, e.g. [NAME]
rights_basis: # who owns audio and text; recording consent and notices
sample_request: # 2-5 hours stratified by speaker, channel, transcriber
Per-segment output from your own alignment pass should then look like this:
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"segment_id": "arch01_doc0457_seg012",
"audio_uri": "audio/doc0457.flac",
"start_s": 312.40,
"end_s": 329.85,
"source_layer": "clean_verbatim",
"target_text": "patient reports intermittent chest pain over the last two weeks",
"matched_error_rate": 0.06,
"mean_alignment_confidence": 0.91,
"boilerplate_token_share": 0.0,
"tier": "train_target",
"redaction_tags_present": false
}
Two items deserve particular attention. Redaction changes the match: if names are masked with silence or tones in audio but replaced with tags in text, alignment must treat those spans consistently, as covered in how audio redaction affects speech model training. And the file-level join between audio and transcript is often the weakest part of an archive; mismatched or missing pairs should be counted before any price discussion.
Domains where existing edited transcripts are common
Edited transcripts cluster in workflows where the document, not the recording, was the product. Physician dictation paired with final reports is the best-studied example [1], and has its own considerations in physician dictation audio for medical ASR. Contact centers often keep QA transcripts or summaries alongside recordings; a call transcript produced for QA review is frequently clean rather than verbatim. Meeting platforms and corporate minute-taking produce notes that range from near-verbatim to summary, which is why meeting transcripts for AI training need the layer question answered first.
Legal depositions and court proceedings are the opposite case: certified transcripts are close to verbatim by design, though they still apply formatting conventions. Broadcasters hold captions, which are useful for light supervision but vary by whether they were prepared offline or captioned live.
For any of these, run the acceptance checks in acceptance testing supplier transcripts on the retained segments, not only on the raw archive.
Rights and privacy questions specific to reused transcripts
Reusing an existing transcript raises a question that new transcription does not: who owns the text, and was it created under terms that allow model training. A caption file may belong to a broadcaster or captioning vendor rather than the audio owner, and a dictation report is part of a medical record. Confirm the rights basis for both audio and text separately, and for health data confirm the de-identification method before any transfer.
SourceX sources operational datasets, including support and sales histories, documents and new recordings of hands-on work, from US companies on request; nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery with the method recorded and a sample checked, though no method is perfect, and health records require HIPAA de-identification by Safe Harbor or Expert Determination. If you are scoping such an archive, you can describe the audio and transcript layers you need.
Source audio with existing transcripts for ASR training
Describe the audio, the transcript layers you can use and the domain, and SourceX looks for US businesses that hold it, assesses data and licensing permissions, and agrees pricing and allowed uses in a license only once the supplying company approves the release. Diligence materials on source, rights, preparation and allowed use are prepared per dataset, and delivery runs through private, access-controlled workflows after an executed agreement. Start a buyer request at sourcex.si/buyers.
Sources
- Association for Computational Linguistics, "Generating training data for medical dictations (NAACL 2001, N01-1017)" (2001). https://preview.aclanthology.org/fix_video/N01-1017.pdf
- arXiv, "The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage (arXiv:2111.09344)" (2021). https://arxiv.org/pdf/2111.09344
- arXiv, "Style-agnostic evaluation of ASR using multiple reference transcripts (arXiv:2412.07937)" (2024). https://arxiv.org/pdf/2412.07937
- arXiv (OpenAI), "Robust Speech Recognition via Large-Scale Weak Supervision (arXiv:2212.04356)" (2022). https://arxiv.org/pdf/2212.04356
- arXiv, "HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips (arXiv:1906.03327)" (2019). https://arxiv.org/pdf/1906.03327
- arXiv (NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (arXiv:2103.14749)" (2021). https://arxiv.org/abs/2103.14749
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.