Skip to content

Speech and audio data

Acceptance Testing Supplier Transcripts: Auditing Transcript Accuracy in Licensed Speech Data

Quick answer

To audit transcript accuracy in a licensed speech dataset, fix the transcription standard in the contract, draw a stratified random sample of segments, have an independent transcriber re-transcribe them to that standard without seeing the supplier text, normalize both sides identically, and compute word error rate with the supplier transcript as the hypothesis. Then adjudicate disagreements, and separately verify label coverage, timestamps, speaker labels, redaction tags and file metadata before you sign acceptance.

By SourceX Editorial · Updated

Why supplier transcripts need their own acceptance test

Supplier transcripts are labels, and labels in shipped datasets carry errors at rates that matter. One audit of widely used test sets, including audio datasets, estimated an average label error rate of at least 3.3% [5]. For ASR, a transcript error is a wrong training target: the model learns to emit the mistake, and if the delivery doubles as an evaluation reference, your WER numbers inherit it.

Generic supplier checks, such as those in our guide to evaluating data supplier quality, confirm that a vendor has a QA process. They do not tell you whether this specific delivery meets the accuracy you paid for. That requires a measurement you run yourself, on audio the supplier did not pick, against a standard both parties signed. The general framework lives in acceptance criteria for licensed training data; this page covers the speech-specific mechanics.

Pin the transcription standard before you measure anything

You cannot score accuracy without an agreed style guide, because most "errors" in a naive comparison are style differences. Datasets vary in style and formality, and research using multiple reference transcripts shows that single-reference WER likely overstates the contentful errors in a transcript [2]. Write the standard into the order form or statement of work, not just the supplier's onboarding deck.

The style guide should resolve, at minimum:

  • Verbatim level: whether fillers (um, uh), false starts, repetitions and stutters are transcribed. See verbatim vs clean transcription standards.
  • Numerals and entities: "twenty five" vs "25", currency, dates, account-like strings, spelled-out letters.
  • Tags: the exact inventory for noise, laughter, crosstalk, unintelligible speech and foreign words, with their spelling (for example [noise], [unk]).
  • Redaction tokens: the placeholder used where personal data was removed, and whether it replaces one word or a span. The training effects are covered in audio redaction artifacts in speech model training.
  • Overlap and speaker turns: how simultaneous speech is represented and attributed.
  • Casing and punctuation: whether they are part of the deliverable or stripped.

Draw a stratified sample the supplier does not choose

The sample should be random within strata that mirror where errors concentrate, and it should be drawn by you from the full delivery manifest. Supplier-curated "golden" files tell you about their best work, not the batch. Stratify by the variables that move transcription difficulty: channel (telephony at 8 kHz vs wideband), speaker accent or language, background noise, overlap density, segment duration, and, if the supplier discloses it, the transcriber or vendor team that produced the batch.

For a continuing stream of monthly or per-batch deliveries, an attribute sampling scheme such as ANSI/ASQ Z1.4 gives you normal, tightened and reduced inspection with switching rules tied to an acceptable quality level [6]. Adapting it to datasets means defining the unit (a segment or a file) and what counts as a defective unit, which is worked through in acceptance sampling for dataset deliveries. Measure WER on the pooled sample, but also record a per-file pass or fail so you can apply lot-level rules.

Re-transcribe blind, normalize identically, then adjudicate

The core test is an independent human re-transcription of the sampled audio, compared word by word with the supplier's transcript. Run it in four steps.

  1. Blind reference. Give your transcribers the audio and the contracted style guide, not the supplier text. Post-editing the supplier transcript anchors the reviewer and hides errors.
  2. Identical normalization. Apply one normalization script to both texts: casing, punctuation, numeral expansion, contraction handling, tag mapping. Normalization choices are a studied source of variation in benchmark error rates [3], and tools such as BenchmarkSTT expose them as configurable steps so the pipeline is reproducible [4].
  3. Compute WER with the supplier transcript as hypothesis. WER is (substitutions + deletions + insertions) divided by reference word count. Report substitution, deletion and insertion counts separately, since a deletion-heavy profile often signals skipped speech rather than mishearing.
  4. Adjudicate disagreements. Normalization is never perfect, so spot-check the diff rather than trusting the number alone [1]. A third reviewer labels each disagreement as supplier error, reference error, or legitimate style ambiguity; for ambiguous spans, accept either rendering, which is the logic of scoring against multiple references [2].

Keep the adjudicated reference set. It becomes the regression suite for every later batch and, after it is frozen, a clean evaluation slice.

Check what WER misses: coverage, timing, speakers and metadata

A delivery can pass on WER and still fail on structure, so run deterministic checks on every file, not just the sample. Start with label coverage: confirm that the share of hours carrying transcripts and diarization matches the contract.

CheckWhat to computeTypical failure mode
Label coverageTranscribed and diarized hours ÷ delivered hours, per language and stratumUntranscribed tail of hard audio silently excluded
Audio-transcript pairingEvery audio ID has exactly one transcript; no orphansOff-by-one file naming after a re-export
Timestamp integritySegment start < end; end ≤ file duration; no overlaps within a speaker unless allowedOffsets drift after resampling or trimming leading silence
Alignment spot-checkListen to segment boundaries on the sampleWords cut at boundaries; segments shifted by a fixed lag
Speaker labelsSpeaker count per file vs audio; label consistency across a fileAgent and customer labels swapped mid-call
Overlap annotationOverlap regions marked per the style guideCrosstalk flattened into one speaker's turn
Redaction tagsTag spelling and density; audio masked where tags appearTag present in text but PII audible in audio
Empty and junk segmentsZero-length text, tag-only segments, non-UTF-8 bytesSilence segments counted toward billed hours
File metadataSample rate, channels, codec and duration match the spec and the manifestStereo calls delivered as downmixed mono

Overlap handling has its own pitfalls, covered in overlapping speech data for multi-talker ASR, and required header and speaker fields are listed in speech dataset metadata fields. Residual personal data is a separate audit with its own sampling plan; see residual PII audit sampling.

Write acceptance criteria and remedies the supplier can price

Criteria work only when they are measurable, tied to a method, and paired with a remedy. Agree the thresholds, the sampling plan and the normalization script before delivery, ideally after a small paid pilot, as described in how to run a data pilot with a supplier. Data quality management standards such as ISO/IEC 5259-3 expect this kind of defined, repeatable process but leave the specific metrics to you [7].

Illustrative example: invented to show structure; it does not describe an available dataset.

acceptance_spec:
  standard: "Contract style guide v2, verbatim with fillers, tags per Annex B"
  unit: segment
  sample:
    method: stratified_random
    strata: [channel, language, overlap_density, duration_bucket]
    drawn_by: buyer
  transcript_accuracy:
    metric: WER (supplier transcript = hypothesis; adjudicated reference)
    normalization: "buyer_norm.py, shared with supplier before delivery"
    batch_threshold: "<= 5.0% pooled; <= 10% for any single stratum"
    style_ambiguity: "counted as correct if either rendering matches the guide"
  structural_checks:
    label_coverage_min: "98% of delivered hours transcribed and diarized"
    orphan_files: 0
    timestamp_violations: "0 hard; boundary drift <= 200 ms on listened sample"
    speaker_label_swaps: "<= 1 file per 100 sampled"
  remedies:
    on_fail: "batch rejected; supplier re-transcribes failed strata and redelivers"
    retest: "tightened sample on redelivery"
    inspection_window_days: 30

The thresholds above are placeholders: set yours from pilot results and from what the downstream use tolerates. An evaluation reference warrants a much stricter bar than bulk pre-training audio. Process terms such as inspection windows, rejection notices and cure periods are covered in dataset acceptance testing.

Where licensed real-world speech fits

Operational recordings from real businesses, such as support and sales calls, tend to be harder to transcribe than read speech, so acceptance testing matters more, not less. SourceX sources operational datasets, including support and sales histories and new recordings of hands-on work, from US companies on request; nothing is held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement. You can describe the transcription standard and acceptance checks you need when you submit a speech data request. Start from the speech and audio data buyer's guide for sourcing context.

Getting a transcription accuracy audit into your next speech data deal

SourceX looks for US businesses that hold the speech data you describe and manages the commercial process from assessment through a license that defines records, uses, term and delivery. Diligence materials on source, rights, preparation and allowed use are prepared per dataset, and nothing is contracted until a supplier agrees. Describe your speech data and acceptance criteria to SourceX.

Sources

  1. Speechmatics, "Calculating WER". https://docs.speechmatics.com/tutorials/calculating-wer
  2. arXiv, "Style-agnostic evaluation of ASR using multiple reference transcripts" (2024). https://arxiv.org/pdf/2412.07937
  3. arXiv, "Investigating Transcription Normalization in the Faetar ASR Benchmark" (2025). https://arxiv.org/pdf/2508.11771
  4. BenchmarkSTT project, "BenchmarkSTT tutorial (1.0.0)". https://benchmarkstt.readthedocs.io/en/1.0.0/tutorial.html
  5. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes". https://asq.org/quality-press/display-item?item=T1164
  7. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Data quality for analytics and ML, Part 3: Data quality management requirements and guidelines" (2024). https://www.iso.org/standard/81092.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data