Skip to content

Speech and audio data

Domain Vocabulary Coverage in Speech Data: Terms, Entities and Numbers

Quick answer

Domain-specific ASR training data closes a vocabulary gap only if it contains your terms, entities and number formats in realistic speech. Build a weighted term list from your domain, score a baseline model on term and entity error rate rather than overall WER alone, normalize numbers before scoring, then buy in-domain audio for pronunciation and acoustic coverage and in-domain text for language-model coverage. Check term occurrence counts in candidate audio before you license it.

By SourceX Editorial · Updated

Why overall WER hides domain vocabulary failures

Overall word error rate (WER) under-reports domain failures because domain terms are a small share of tokens but carry most of the business value. A cardiology dictation model can score a respectable WER while consistently writing "metoprolol" as "met a pro lol", turning a policy number into words, or dropping a SKU prefix. Function words and common verbs dominate the denominator, so a model can improve WER on everyday speech while getting worse on the words downstream systems actually parse.

The practical fix is to report at least three numbers side by side: overall WER, a term error rate over a curated domain lexicon, and an entity error rate over tagged spans such as drug names, ICD-10 codes, account identifiers, carrier names and street addresses. Treat the split metrics as working hypotheses for your stack, not universal standards; the point is to make the vocabulary gap visible before you spend money closing it. If your team already runs named entity recognition on transcripts, reuse its span labels for scoring.

Build the domain term list before you look at any dataset

A domain term list is the measuring stick for every later decision, so it comes first. Pull candidates from the systems your model will feed: EHR order sets and formularies in healthcare, product and fund names in finance, coverage and peril codes in insurance claims, SKU masters and carrier or lane codes in logistics. Add abbreviations and the ways people say them ("A-fib", "afib", "atrial fibrillation"), because spoken variants rarely match the written canonical form.

Weight each term by downstream impact and expected frequency, then bucket it: high-impact rare (drug names, dosage units, policy numbers), high-impact common (product names), and long tail. Specialty coverage matters as much as raw size. Vendors in medical dictation advertise coverage by specialty count, for example 31 specialties in one self-reported catalog as of October 2026 [4], which signals that buyers do and should specify sub-domain targets rather than "medical" as a single bucket.

Illustrative example: invented to show structure; it does not describe an available dataset.

term_idcanonical_formspoken_variantscategoryimpact_weightmin_occurrences_audiomin_speakersnotes
T-0412metoprolol succinate"metoprolol", "toprol"drug53010brand and generic both needed
T-0977ICD-10 I48.91"I forty-eight point nine one"code4208test letter-digit-dot format
T-1203comprehensive deductible"comp deductible"insurance term32510clipped form common on calls
T-15501Z tracking prefix"one Z", "one zed"logistics code44015spelled-out alphanumerics
T-2001250 mg BID"two fifty milligrams twice a day", "B I D"dose + frequency53010normalization rule required

Numbers, codes and amounts need normalization rules before scoring

Numbers, codes and currency amounts must be normalized to one canonical form before you compute any error rate, or you will measure formatting rather than recognition. "Two fifty" can be 250 or 2:50, "$1.2M" may be spoken "one point two million dollars", and "I48.91" may be read letter by letter. Research evaluating spoken reasoning, including math, normalized spoken numbers and operators to canonical notation before computing WER for exactly this reason [2].

The Whisper authors built a dedicated text normalizer so that harmless spelling and formatting differences were not counted as errors, and checked it against an independently written normalizer to reduce the risk of tuning it to their own model's quirks [3]. Domain teams should do the same with a written normalization spec: digit versus word rules, unit spellings (mg, mcg, mL), date and time formats, currency, alphanumeric codes, and whether "dash" and "slash" are spoken tokens. Version the spec, apply it to both reference and hypothesis, and keep raw unnormalized transcripts so you can rerun scoring when the spec changes. Transcription conventions in purchased data must match it; see verbatim vs clean transcription standards for how style choices affect numbers and disfluencies.

In-domain text adapts the language model; in-domain audio fixes pronunciation

In-domain text corpora are an efficient way to teach a model which words and word sequences are likely, but they cannot teach how those words sound in your speakers' mouths. Medical ASR work has combined general speech fine-tuning with a separate medical text corpus used for language-model adaptation [1], which is a reasonable pattern when audio is scarce and text such as clinical notes or claim narratives is plentiful.

Text-only adaptation stops working where the acoustic signal is the problem: unusual pronunciations, brand names with non-English roots, clipped jargon on telephony audio, and alphanumeric codes read quickly. Those failures need paired audio and transcripts containing the terms. A workable split for most domains is text for breadth across the long tail, and audio for the high-impact terms, number formats and speaker conditions you will see in production. If you are sourcing the text side, the companion guide on domain vocabulary in text data covers abbreviation density and jargon coverage.

Illustrative example: invented to show structure; it does not describe an available dataset.

Gap observed in baselinePrimary fixData to sourceFailure mode if you pick the wrong one
Rare terms substituted by common wordsLM or biasing adaptationIn-domain text, term listsAudio-only buys pay for hours that rarely contain the term
Terms recognized in read speech, missed on callsAcoustic fine-tuningTelephony in-domain audioText adaptation raises term prior but acoustics still fail
Codes and amounts garbledNormalization plus audioAudio with spoken codes, normalization specScoring noise masks real improvement
Brand or drug names mispronounced by speakersAcoustic plus lexiconAudio from many speakers per termModel overfits to one speaker's pronunciation

Check term occurrence counts in candidate audio before you license it

Hours are a poor proxy for vocabulary coverage, so ask for term occurrence counts against your list before you agree to buy. Hand the supplier your lexicon (or a hashed or category-level version if terms are sensitive) and request per-term counts in the transcripts, distinct speaker counts per term, and the split by recording condition. A 500-hour corpus of general customer service calls may contain fewer target term instances than 40 hours from the right sub-domain.

Run the check on a sample you can score yourself: transcribe it with your baseline, compute term and entity error rates on the normalized output, and confirm the reference transcripts are accurate. Test sets across widely used vision, text and audio benchmarks have shown an estimated average label error rate of at least 3.3% [5], and in domain audio, reference errors on rare terms are the ones most likely to distort a term error rate. Budget a manual audit of reference transcripts on target-term spans, not on random segments.

Coverage also interacts with other axes. Accent and dialect spread affect how domain terms are pronounced (accented English speech data), and noise and channel conditions determine whether a term survives a speakerphone (noisy speech and SNR coverage). The how many hours to fine-tune ASR guide covers volume once coverage is known.

Specify vocabulary coverage in the data request

A data request for domain speech should state coverage targets per term bucket, not just total hours. Suppliers who hold real operational recordings (dictation, claims calls, dispatch radio, sales calls) may be able to report term counts if you give them the list and the counting rule; synthetic augmentation can fill gaps but has known limits, covered in real vs synthetic speech data for ASR.

Illustrative example: invented to show structure; it does not describe an available dataset.

request: in-domain speech for ASR adaptation
domain: property and casualty insurance claims
sub_domains: [auto first notice of loss, homeowners water damage, adjuster callbacks]
audio:
  channel: telephony, 8 kHz mono acceptable; wideband preferred
  speech_style: spontaneous two-party calls
  transcription: verbatim, speaker-labeled, timestamps per segment
  normalization_spec: attached v1.3 (numbers, currency, policy and claim IDs)
vocabulary_coverage:
  term_list: attached (412 terms, weighted 1-5)
  weight_5_terms: >= 30 occurrences each, >= 10 distinct speakers
  weight_3_4_terms: >= 15 occurrences each
  report: per-term counts from reference transcripts before agreement
entities_tagged: [policy_number, claim_number, vehicle_make_model, address, dollar_amount]
privacy: names, phone numbers, account and policy numbers replaced with typed placeholders; method documented
text_side: claim notes or adjuster narratives for LM adaptation, same de-identification

Placeholder replacement interacts with vocabulary coverage: if every policy number becomes a tag, you lose the spoken digit patterns you wanted. Agree with the supplier whether identifiers are replaced with format-preserving synthetic values or silenced, and how transcripts mark the change; the trade-offs are covered in audio redaction artifacts in speech model training.

Keep a domain vocabulary eval set separate from training data

A held-out domain eval set is the only way to prove that purchased data closed the gap, so carve it out before training. Draw it from different speakers, sessions and ideally different source organizations than the training audio, and stratify it by term bucket so weight-5 terms have enough instances to measure. Freeze it alongside the normalization spec and term list version.

Track the three metrics per release and per sub-domain, and log which data batch was added between runs. This fits the MEASURE function of the NIST AI RMF, which frames measurement as an ongoing lifecycle activity rather than a one-time test [6]. When a new product line, drug or carrier code enters the business, add it to the list, check the eval set's coverage, and source more data only where the split metrics show a regression.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Term list versioned, weighted and bucketed by sub-domain
  • Baseline scored on overall WER, term error rate and entity error rate
  • Written normalization spec applied to reference and hypothesis
  • Per-term occurrence and speaker counts received from supplier before agreement
  • Reference transcripts audited on target-term spans
  • De-identification method agreed so identifier patterns survive where needed
  • Held-out eval set disjoint by speaker, session and source

How SourceX fits a domain ASR data request

SourceX sources operational datasets from US companies on request, including support and sales histories and new recordings of hands-on work, and manages the licensing process; nothing is held in stock and a request does not guarantee a match. Buyers describe the data they need, such as call audio rich in specific claim or product vocabulary, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. You can share a term list and coverage targets as part of a request on the buyer page, and see related capabilities under domain-specific fine-tuning data.

Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Health records such as physician dictation require HIPAA de-identification by Safe Harbor or Expert Determination. More on the wider category is in the speech and audio data buyer's guide and the AI data hub.

Request domain-specific ASR training data

If your baseline shows term or entity errors that general speech data will not fix, describe the in-domain audio and text you need, including your term list and coverage targets. SourceX assesses data and licensing permissions with candidate suppliers and agrees allowed uses in a license before anything is transacted. Start a request at sourcex.si/buyers.

Sources

  1. arXiv, "Automatic Speech Recognition for Greek Medical Dictation" (2025). https://arxiv.org/pdf/2509.23550
  2. arXiv, "Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap" (2025). https://arxiv.org/pdf/2509.26542
  3. OpenAI (arXiv:2212.04356), "Robust Speech Recognition via Large-Scale Weak Supervision" (2022). https://arxiv.org/pdf/2212.04356
  4. Shaip, "License High-quality Healthcare/Medical Data for AI & ML Models (physician dictation audio)". https://shaip.com/offerings/physician-dictation-audio-data-medical-data-catalog
  5. Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data