Speech and audio data
How to Measure Word Error Rate Fairly: Normalization, References and Vendor Comparisons
Quick answer
Word error rate (WER) is the number of substitutions, deletions and insertions needed to turn a system's transcript into the reference, divided by the number of words in the reference [1]. The formula is simple; the hard part is making scores comparable. Run one versioned text normalizer on both reference and hypothesis, check that the reference style matches what you expect the model to output, report a formatting-aware metric where casing and numbers matter, and test on audio from your own domain [1][2][4].
By SourceX Editorial · Updated
The WER formula and what each term counts
WER is computed as (S + D + I) / N, where S, D and I come from a minimum edit-distance alignment between the reference and the hypothesis, and N is the number of reference words [1]. A substitution is a reference word replaced by a different word. A deletion is a reference word the system dropped, and an insertion is an extra hypothesis word with no reference counterpart. Because insertions are unbounded, WER can exceed 100% on hallucination-heavy output, so it is an error rate, not an accuracy percentage.
Three implementation details change the number more than most teams expect:
- Corpus-level vs. averaged WER. Corpus WER sums S, D and I across all files and divides by total reference words. Averaging per-file WER gives short utterances the same weight as long ones, so a three-word clip with one error counts as 33%. Report corpus WER and state which you used.
- Tokenization. Splitting on whitespace treats "re-enter", "re enter" and "reenter" as one, two and one tokens. Decide hyphen and apostrophe handling before you score, not after.
- Empty references. Segments where the reference is silence or non-speech have N = 0. Exclude them from WER and track their insertions separately, or the metric becomes undefined.
Keep S, D and I in your report, not only the ratio. A model whose errors are mostly deletions behaves differently in production, for example by dropping words under crosstalk, than one that inserts filler or hallucinates during silence.
Why text normalization before WER decides the result
Normalization is the step that strips formatting differences that do not change meaning, and it must be applied identically to both reference and hypothesis [1]. Without it, "Dr." versus "doctor", "3rd" versus "third" or "you're" versus "you are" each count as errors even when the speech was recognized correctly. The effect is largest on conversational references full of contractions and on read news or finance audio full of spoken numbers and currency, where a single normalizer rule can move a score by several points.
Normalization cuts both ways. A normalizer tuned to one system's output habits can flatter that system, so the normalizer itself is a variable in any comparison. Treat it like code under test: version it, pin it in the evaluation config, and never let a vendor score its own output with its own normalizer when you are comparing vendors.
Normalizers also make mistakes. Rule-based number expansion can turn "2020" into "two thousand twenty" in the reference and "twenty twenty" in the hypothesis, creating two errors where there were none. Speechmatics' guidance is to spot-check alignments after normalization rather than trust the final number [1]. Open-source tools such as BenchmarkSTT expose normalization as configurable steps, which makes the rules auditable and reproducible across runs [5].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Step | Reference | Hypothesis | S | D | I | WER |
|---|---|---|---|---|---|---|
| Raw text, case-sensitive | please send the invoice to Dr. Smith by March 3rd | please send invoice to doctor smith by march third | 4 | 1 | 0 | 5/10 = 50% |
| Lowercase, strip punctuation | please send the invoice to dr smith by march 3rd | please send invoice to doctor smith by march third | 2 | 1 | 0 | 3/10 = 30% |
| Plus abbreviation and ordinal expansion | please send the invoice to doctor smith by march third | please send invoice to doctor smith by march third | 0 | 1 | 0 | 1/10 = 10% |
Only the dropped "the" is a recognition error. The 50% raw score is almost entirely a formatting disagreement, which is exactly the distortion that makes unnormalized vendor comparisons meaningless.
A normalization spec you can pin in your evaluation config
A good normalization spec lists every rule, applies it to both sides, and stays fixed across every system you score. Write it down before the first vendor result arrives so no one tunes it after seeing rankings.
Illustrative example: invented to show structure; it does not describe an available dataset.
normalizer: wer-norm-en-us
version: 1.3.0
apply_to: [reference, hypothesis]
rules:
case: lower
punctuation: strip # keep apostrophes inside words until contraction step
contractions: expand # "you're" -> "you are"; list in contractions_en.tsv
numbers: spoken_form # "3rd" -> "third", "$68" -> "sixty eight dollars"
abbreviations: expand # "dr" -> "doctor"; list in abbrev_en.tsv
hyphens: split # "re-enter" -> "re enter"
spelling_variants: us_english # "colour" -> "color"
fillers: remove # "um", "uh", "er" removed from both sides
disfluency_tags: remove # "[crosstalk]", "<unk>" removed before alignment
partial_words: remove # "transf-" removed
non_speech_segments: exclude_from_wer
report: [corpus_wer, substitutions, deletions, insertions, n_ref_words, cer, formatted_error_rate]
Two choices in this spec deserve explicit sign-off. Removing fillers is right if your product discards them, but wrong if you are evaluating a verbatim captioning model, and the verbatim vs clean transcription standards guide covers that decision. Converting numbers to spoken form hides real errors in account numbers and dates, so pair it with a separate entity check described below.
How reference transcript style distorts scores
Reference style is a major hidden source of WER disagreement, because human transcribers make legitimate but different choices about fillers, false starts, contractions and formality [3]. Research from Rev argues that datasets vary in style and formality, and that standard WER probably overstates the contentful errors of current top systems as a result [3]. A clean-read reference scored against a verbatim-style hypothesis, or the reverse, penalizes the model for following a convention rather than for mishearing.
The paper's remedy is to score against multiple reference transcripts in different styles and count a hypothesis word as correct if it matches any acceptable variant [3]. If you cannot afford full multiple references, a practical substitute is a style guide shared with every transcriber, a lattice of allowed alternates for common variants such as "okay/OK" and "gonna/going to", and a double-transcribed subset to measure how much two humans disagree. That inter-transcriber disagreement is a floor: differences between vendors smaller than it are not meaningful.
References also contain outright errors. An audit of widely used test sets across vision, text and audio estimated an average label error rate of at least 3.3%, enough to change which model ranks first [6]. Budget for a correction pass on your reference set, and use the sample size guide for estimating a dataset's error rate to size the audit.
WER vs CER and formatted error rate: which metric to report
WER is the right headline for word-segmented languages and for spoken content, but it is blind to casing, punctuation and number formatting, which matter in captions, notes and any downstream parsing [4]. 3Play Media's ASR report adds a formatted error rate for this reason, so output that reads correctly is distinguished from output that only sounds correct [4]. Character error rate (CER) applies the same edit-distance formula at the character level and is the usual choice for languages written without spaces between words, such as Chinese, Japanese and Thai, where word segmentation is itself a modeling choice [1].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Use case | Headline metric | Report alongside | Normalization stance |
|---|---|---|---|
| Voice search, command and control | WER | Entity accuracy on product and place names | Aggressive |
| Contact center analytics | WER | Deletion rate per channel, keyword recall | Aggressive, fillers removed |
| Captioning and subtitles | Formatted error rate | WER, punctuation F1 | Minimal: keep case and punctuation |
| Clinical or legal notes | WER | Numeric and entity error rate, formatted error rate | Moderate; never collapse drug doses or dates |
| Mandarin, Japanese, Thai | CER | WER after a fixed segmenter, if needed | Character-level, width and script normalized |
| Code-switched speech | WER and CER per language | Language-ID accuracy | Per-language rules |
For regulated or numeric-heavy content, add an entity error rate: extract dates, amounts, drug names and identifiers from both sides and score them as exact matches. A 4% WER that gets one digit of a dosage wrong is worse than a 7% WER that only drops fillers. The domain vocabulary coverage guide explains how to build that term list.
A fair protocol for comparing ASR vendors with WER
A fair vendor comparison holds everything constant except the system under test: the same audio, the same references, the same normalizer and the same scoring code [1][2]. Vendor-published WER figures usually fail at least one of these conditions, so treat them as marketing context rather than evidence.
- Use your own domain audio. Public sets such as read audiobooks rarely match call audio, accents or vocabulary, and Speechmatics' benchmarking guidance recommends testing on representative data [2]. Match codec and sample rate to production; the telephony vs wideband guide covers why 8 kHz audio scores differently.
- Keep the test set private. Do not upload it for vendor "tuning" before scoring, and rotate a held-out slice so it cannot leak into anyone's training data.
- Stratify the set. Tag each file by accent, channel, noise and speaker so you can report per-slice WER. The accent-stratified evaluation set guide shows how to design the strata.
- Score everything yourself. Collect raw hypotheses, apply your pinned normalizer, and compute S, D and I with one tool.
- Quantify uncertainty. Bootstrap over files or speakers, not individual words, to get confidence intervals; word-level resampling understates variance because errors cluster within recordings.
- Inspect the alignments. Read the top 50 highest-error files for each vendor. Normalizer bugs, reference errors and silence hallucinations show up there first [1].
- Report the full picture. Corpus WER, S/D/I, per-slice WER, CER or formatted error rate as relevant, and the normalizer version.
If diarization is in scope, score it separately; mixing speaker attribution errors into WER hides both. The diarization error rate guide covers that metric.
Using WER to judge whether new training data is worth buying
WER is also a direct, practical way to estimate the value of additional speech data, provided the test set reflects the gap the data is meant to close. Before licensing a corpus, fine-tune on a sample and score the result on a held-out, in-domain test set that shares no speakers or recordings with the sample. If the sample improves WER only on slices that are already strong, the full set is unlikely to justify its cost.
Ask suppliers for a representative sample with references produced under your style guide, not theirs, so the evaluation measures the audio rather than a transcription convention [3]. The guide to requesting a training data sample lists what to specify, and the hours needed to fine-tune ASR guide helps size the full purchase. For definitions of the underlying terms, see the speech-to-text glossary entry and the model evaluation glossary entry.
When the evaluation shows a gap that public corpora cannot fill, such as real support calls or recordings of hands-on work in a specific trade, SourceX helps AI teams source that data from US companies on request, with every dataset rights-reviewed and personal details removed or replaced before delivery.
Common failure modes in WER reporting
Most misleading WER numbers come from a small set of avoidable mistakes:
- Normalizing only the hypothesis, or using different normalizer versions across runs [1].
- Comparing WER across test sets with different reference styles, then attributing the gap to the model [3].
- Averaging per-utterance WER, which overweights short clips.
- Dropping files where a vendor timed out or returned empty output, which hides the worst cases; score empty output as all deletions.
- Ignoring insertions during long silences or hold music, where some end-to-end models hallucinate text.
- Reporting a single public-benchmark number for a domain the benchmark does not represent [2].
For the broader speech data picture, including licensing and sourcing, start at the speech and audio data hub or the AI data guides index.
Get speech evaluation or training data that matches your domain
SourceX sources operational datasets from US companies on request, including support and sales histories and new recordings of hands-on work, and manages the licensing and purchase process. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery; a request does not guarantee a match. Describe the speech data you need on the SourceX buyers page.
Sources
- Speechmatics, "Calculating Word Error Rate". https://docs.speechmatics.com/tutorials/calculating-wer
- Speechmatics, "Accuracy Benchmarking". https://docs.speechmatics.com/tutorials/accuracy-benchmarking
- arXiv (Rev), "Style-agnostic evaluation of ASR using multiple reference transcripts" (2024). https://arxiv.org/pdf/2412.07937
- 3Play Media, "The 2024 State of Automatic Speech Recognition Report" (2024). https://go.3playmedia.com/hubfs/WP%20PDFs/The%202024%20State%20of%20Automatic%20Speech%20Recognition%20Report_Final_Remediated-1.pdf
- BenchmarkSTT project, "BenchmarkSTT tutorial (v1.0.0)". https://benchmarkstt.readthedocs.io/en/1.0.0/tutorial.html
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.