Speech and audio data
Domain Vocabulary Coverage in Speech Data: Terms, Entities and Numbers
Quick answer
Domain-specific ASR training data closes a vocabulary gap only if it contains your terms, entities and number formats in realistic speech. Build a weighted term list from your domain, score a baseline model on term and entity error rate rather than overall WER alone, normalize numbers before scoring, then buy in-domain audio for pronunciation and acoustic coverage and in-domain text for language-model coverage. Check term occurrence counts in candidate audio before you license it.
By SourceX Editorial · Updated
Why overall WER hides domain vocabulary failures
Overall word error rate (WER) under-reports domain failures because domain terms are a small share of tokens but carry most of the business value. A cardiology dictation model can score a respectable WER while consistently writing "metoprolol" as "met a pro lol", turning a policy number into words, or dropping a SKU prefix. Function words and common verbs dominate the denominator, so a model can improve WER on everyday speech while getting worse on the words downstream systems actually parse.
The practical fix is to report at least three numbers side by side: overall WER, a term error rate over a curated domain lexicon, and an entity error rate over tagged spans such as drug names, ICD-10 codes, account identifiers, carrier names and street addresses. Treat the split metrics as working hypotheses for your stack, not universal standards; the point is to make the vocabulary gap visible before you spend money closing it. If your team already runs named entity recognition on transcripts, reuse its span labels for scoring.
Build the domain term list before you look at any dataset
A domain term list is the measuring stick for every later decision, so it comes first. Pull candidates from the systems your model will feed: EHR order sets and formularies in healthcare, product and fund names in finance, coverage and peril codes in insurance claims, SKU masters and carrier or lane codes in logistics. Add abbreviations and the ways people say them ("A-fib", "afib", "atrial fibrillation"), because spoken variants rarely match the written canonical form.
Weight each term by downstream impact and expected frequency, then bucket it: high-impact rare (drug names, dosage units, policy numbers), high-impact common (product names), and long tail. Specialty coverage matters as much as raw size. Vendors in medical dictation advertise coverage by specialty count, for example 31 specialties in one self-reported catalog as of October 2026 [4], which signals that buyers do and should specify sub-domain targets rather than "medical" as a single bucket.
Illustrative example: invented to show structure; it does not describe an available dataset.
| term_id | canonical_form | spoken_variants | category | impact_weight | min_occurrences_audio | min_speakers | notes |
|---|---|---|---|---|---|---|---|
| T-0412 | metoprolol succinate | "metoprolol", "toprol" | drug | 5 | 30 | 10 | brand and generic both needed |
| T-0977 | ICD-10 I48.91 | "I forty-eight point nine one" | code | 4 | 20 | 8 | test letter-digit-dot format |
| T-1203 | comprehensive deductible | "comp deductible" | insurance term | 3 | 25 | 10 | clipped form common on calls |
| T-1550 | 1Z tracking prefix | "one Z", "one zed" | logistics code | 4 | 40 | 15 | spelled-out alphanumerics |
| T-2001 | 250 mg BID | "two fifty milligrams twice a day", "B I D" | dose + frequency | 5 | 30 | 10 | normalization rule required |
Numbers, codes and amounts need normalization rules before scoring
Numbers, codes and currency amounts must be normalized to one canonical form before you compute any error rate, or you will measure formatting rather than recognition. "Two fifty" can be 250 or 2:50, "$1.2M" may be spoken "one point two million dollars", and "I48.91" may be read letter by letter. Research evaluating spoken reasoning, including math, normalized spoken numbers and operators to canonical notation before computing WER for exactly this reason [2].
The Whisper authors built a dedicated text normalizer so that harmless spelling and formatting differences were not counted as errors, and checked it against an independently written normalizer to reduce the risk of tuning it to their own model's quirks [3]. Domain teams should do the same with a written normalization spec: digit versus word rules, unit spellings (mg, mcg, mL), date and time formats, currency, alphanumeric codes, and whether "dash" and "slash" are spoken tokens. Version the spec, apply it to both reference and hypothesis, and keep raw unnormalized transcripts so you can rerun scoring when the spec changes. Transcription conventions in purchased data must match it; see verbatim vs clean transcription standards for how style choices affect numbers and disfluencies.
In-domain text adapts the language model; in-domain audio fixes pronunciation
In-domain text corpora are an efficient way to teach a model which words and word sequences are likely, but they cannot teach how those words sound in your speakers' mouths. Medical ASR work has combined general speech fine-tuning with a separate medical text corpus used for language-model adaptation [1], which is a reasonable pattern when audio is scarce and text such as clinical notes or claim narratives is plentiful.
Text-only adaptation stops working where the acoustic signal is the problem: unusual pronunciations, brand names with non-English roots, clipped jargon on telephony audio, and alphanumeric codes read quickly. Those failures need paired audio and transcripts containing the terms. A workable split for most domains is text for breadth across the long tail, and audio for the high-impact terms, number formats and speaker conditions you will see in production. If you are sourcing the text side, the companion guide on domain vocabulary in text data covers abbreviation density and jargon coverage.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Gap observed in baseline | Primary fix | Data to source | Failure mode if you pick the wrong one |
|---|---|---|---|
| Rare terms substituted by common words | LM or biasing adaptation | In-domain text, term lists | Audio-only buys pay for hours that rarely contain the term |
| Terms recognized in read speech, missed on calls | Acoustic fine-tuning | Telephony in-domain audio | Text adaptation raises term prior but acoustics still fail |
| Codes and amounts garbled | Normalization plus audio | Audio with spoken codes, normalization spec | Scoring noise masks real improvement |
| Brand or drug names mispronounced by speakers | Acoustic plus lexicon | Audio from many speakers per term | Model overfits to one speaker's pronunciation |
Check term occurrence counts in candidate audio before you license it
Hours are a poor proxy for vocabulary coverage, so ask for term occurrence counts against your list before you agree to buy. Hand the supplier your lexicon (or a hashed or category-level version if terms are sensitive) and request per-term counts in the transcripts, distinct speaker counts per term, and the split by recording condition. A 500-hour corpus of general customer service calls may contain fewer target term instances than 40 hours from the right sub-domain.
Run the check on a sample you can score yourself: transcribe it with your baseline, compute term and entity error rates on the normalized output, and confirm the reference transcripts are accurate. Test sets across widely used vision, text and audio benchmarks have shown an estimated average label error rate of at least 3.3% [5], and in domain audio, reference errors on rare terms are the ones most likely to distort a term error rate. Budget a manual audit of reference transcripts on target-term spans, not on random segments.
Coverage also interacts with other axes. Accent and dialect spread affect how domain terms are pronounced (accented English speech data), and noise and channel conditions determine whether a term survives a speakerphone (noisy speech and SNR coverage). The how many hours to fine-tune ASR guide covers volume once coverage is known.
Specify vocabulary coverage in the data request
A data request for domain speech should state coverage targets per term bucket, not just total hours. Suppliers who hold real operational recordings (dictation, claims calls, dispatch radio, sales calls) may be able to report term counts if you give them the list and the counting rule; synthetic augmentation can fill gaps but has known limits, covered in real vs synthetic speech data for ASR.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: in-domain speech for ASR adaptation
domain: property and casualty insurance claims
sub_domains: [auto first notice of loss, homeowners water damage, adjuster callbacks]
audio:
channel: telephony, 8 kHz mono acceptable; wideband preferred
speech_style: spontaneous two-party calls
transcription: verbatim, speaker-labeled, timestamps per segment
normalization_spec: attached v1.3 (numbers, currency, policy and claim IDs)
vocabulary_coverage:
term_list: attached (412 terms, weighted 1-5)
weight_5_terms: >= 30 occurrences each, >= 10 distinct speakers
weight_3_4_terms: >= 15 occurrences each
report: per-term counts from reference transcripts before agreement
entities_tagged: [policy_number, claim_number, vehicle_make_model, address, dollar_amount]
privacy: names, phone numbers, account and policy numbers replaced with typed placeholders; method documented
text_side: claim notes or adjuster narratives for LM adaptation, same de-identification
Placeholder replacement interacts with vocabulary coverage: if every policy number becomes a tag, you lose the spoken digit patterns you wanted. Agree with the supplier whether identifiers are replaced with format-preserving synthetic values or silenced, and how transcripts mark the change; the trade-offs are covered in audio redaction artifacts in speech model training.
Keep a domain vocabulary eval set separate from training data
A held-out domain eval set is the only way to prove that purchased data closed the gap, so carve it out before training. Draw it from different speakers, sessions and ideally different source organizations than the training audio, and stratify it by term bucket so weight-5 terms have enough instances to measure. Freeze it alongside the normalization spec and term list version.
Track the three metrics per release and per sub-domain, and log which data batch was added between runs. This fits the MEASURE function of the NIST AI RMF, which frames measurement as an ongoing lifecycle activity rather than a one-time test [6]. When a new product line, drug or carrier code enters the business, add it to the list, check the eval set's coverage, and source more data only where the split metrics show a regression.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Term list versioned, weighted and bucketed by sub-domain
- Baseline scored on overall WER, term error rate and entity error rate
- Written normalization spec applied to reference and hypothesis
- Per-term occurrence and speaker counts received from supplier before agreement
- Reference transcripts audited on target-term spans
- De-identification method agreed so identifier patterns survive where needed
- Held-out eval set disjoint by speaker, session and source
How SourceX fits a domain ASR data request
SourceX sources operational datasets from US companies on request, including support and sales histories and new recordings of hands-on work, and manages the licensing process; nothing is held in stock and a request does not guarantee a match. Buyers describe the data they need, such as call audio rich in specific claim or product vocabulary, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. You can share a term list and coverage targets as part of a request on the buyer page, and see related capabilities under domain-specific fine-tuning data.
Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Health records such as physician dictation require HIPAA de-identification by Safe Harbor or Expert Determination. More on the wider category is in the speech and audio data buyer's guide and the AI data hub.
Request domain-specific ASR training data
If your baseline shows term or entity errors that general speech data will not fix, describe the in-domain audio and text you need, including your term list and coverage targets. SourceX assesses data and licensing permissions with candidate suppliers and agrees allowed uses in a license before anything is transacted. Start a request at sourcex.si/buyers.
Sources
- arXiv, "Automatic Speech Recognition for Greek Medical Dictation" (2025). https://arxiv.org/pdf/2509.23550
- arXiv, "Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap" (2025). https://arxiv.org/pdf/2509.26542
- OpenAI (arXiv:2212.04356), "Robust Speech Recognition via Large-Scale Weak Supervision" (2022). https://arxiv.org/pdf/2212.04356
- Shaip, "License High-quality Healthcare/Medical Data for AI & ML Models (physician dictation audio)". https://shaip.com/offerings/physician-dictation-audio-data-medical-data-catalog
- Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.