Skip to content

Speech and audio data

How Many Hours of Audio Do You Need to Fine-Tune ASR?

Quick answer

For domain fine-tuning of a strong pretrained model such as Whisper, wav2vec 2.0 or a NeMo Conformer, tens of hours of transcribed in-domain audio is a realistic starting point, often paired with domain text for language-model adaptation; one published medical-dictation study used about 49 hours [1]. Training a general recognizer from scratch is a different problem measured in tens of thousands of hours [2]. Speaker count, acoustic conditions and vocabulary coverage decide whether those hours actually move word error rate (WER).

By SourceX Editorial · Updated

Fine-tuning and from-scratch training sit three to four orders of magnitude apart

The first sizing decision is whether you are adapting an existing acoustic model or building one, because the hours required differ by roughly three to four orders of magnitude. A pretrained encoder already models phonetics, speaker variation and common noise; fine-tuning teaches it your vocabulary, channel and speaking style. Training from scratch must learn all of that, which is why broad English corpora built for general training reach about 30,000 hours [2].

Most buyers sizing a purchase are in the first group. Their real question is not "how much speech exists" but "how much in-domain speech closes the gap between the base model's WER on our audio and our target WER." That gap, not a generic rule of thumb, sets the budget.

GoalTypical starting volumeWhat dominates the outcome
Vocabulary and entity adaptation on a strong base modelText-only LM adaptation plus a few hours of audio for evaluationTerm coverage in text; contextual biasing
Domain fine-tuning (dictation, contact center, field service)Tens of hours of transcribed in-domain audio [1]Speaker count, channel match, transcription standard
Curated conversational or spontaneous-speech modelLow hundreds of hours; curated corpora such as CASPER run about 200 hours [4]Disfluency handling, overlap, turn-taking
Narrow keyword or wake-word taskThousands of targeted utterances rather than hours [3]Positive and hard-negative coverage
General-purpose recognizer from scratchTens of thousands of hours [2]Breadth of speakers, accents, topics and channels

These rows are planning anchors from published work, not guarantees. A base model that already performs well on your audio may need far less; a low-resource language or 8 kHz telephony channel the base model rarely saw may need more.

Hours are the wrong single unit for an ASR data request

Hours alone are a poor sizing unit because fifty hours from five speakers and fifty hours from five hundred speakers produce very different models. Fine-tuned ASR overfits to the voices, microphones and rooms it sees. If a held-out test speaker is acoustically unlike every training speaker, WER on that speaker can stay close to the base model's level no matter how many hours you added.

Specify the request along at least five axes:

Public corpora show why speaker breadth matters at scale: Common Voice reported over 50,000 contributors producing about 2,500 hours as of November 2019 (the corpus has grown substantially since), an average well under an hour per voice [6]. Operational recordings such as call-center audio often have the opposite shape, with a small agent pool and many customers, so count both sides of the conversation separately.

A learning-curve pilot tells you how many hours to buy

The most reliable way to size a speech purchase is to fine-tune on nested subsets and plot WER against hours before committing to full volume. Error on held-out data tends to follow a power law in training-set size across domains including speech recognition [5], so three or four points often give a usable extrapolation and show whether you are on the steep or flat part of the curve.

Illustrative example: invented to show structure; it does not describe an available dataset.

Learning-curve pilot plan (ASR domain fine-tuning)

StepSettingNotes
1. Freeze an eval set3–5 hours of held-out in-domain audio, speakers disjoint from trainingNever draw eval speakers from the purchase pool
2. BaselineBase model WER, plus entity error rate on domain termsNormalize text (casing, numerals, punctuation) identically for every run
3. Nested subsets10, 25, 50, 100 hours, each a superset of the lastKeep speaker and channel mix proportional at every size
4. Fixed recipeSame learning rate schedule, epochs scaled to steps, SpecAugment and speed perturbation onChange one variable at a time
5. Fit the curveLog-log fit of WER versus hoursStop buying when projected gain per added 50 hours falls below your threshold
6. Slice resultsWER by speaker, channel, SNR band, accentFlat aggregate WER can hide a slice that is still failing

A worked reading of such a pilot: if WER falls sharply from 10 to 50 hours and then flattens, the remaining errors are probably not a volume problem. They usually point to transcription inconsistency, missing vocabulary, or an under-represented condition, and the next purchase should target that slice rather than add generic hours.

Text often substitutes for audio in vocabulary-heavy domains

For domains where errors concentrate on rare terms, in-domain text for language-model adaptation can deliver gains that would otherwise require many more hours of audio. The medical-dictation study cited above combined its roughly 49 hours of speech with domain text [1], a common pattern for clinical, legal and technical speech.

Practical options include shallow fusion with an n-gram or neural LM (for example KenLM with CTC decoders), contextual biasing lists of entities, and prompt conditioning in Whisper-style decoders. Text is usually cheaper to license than transcribed audio, but it does not fix acoustic mismatch. Use the pilot's error analysis to decide: substitution errors on known terms point to text; deletions in noise or overlap point to more audio.

Targeted collections beat generic volume for narrow tasks

When the task is narrow, a small collection aimed at the exact failure mode can compete with a much larger generic set. Wake-word research compared about 9.5k targeted utterances against a 375-hour production baseline [3], a reminder that the useful unit for keyword spotting is utterances and hard negatives, not hours.

The same logic applies to fine-tuning. If your base model fails on drive-thru orders over engine noise, a few hours captured in that condition can be worth more than a hundred hours of quiet read speech. Pages on noisy speech data and SNR coverage and overlapping speech for multi-talker ASR cover how to specify those slices.

Transcription quality changes how many hours you need

Inconsistent transcripts raise the hours required because the model spends capacity learning annotator noise. A corpus that mixes verbatim transcription (fillers, false starts, partial words) with clean transcription will teach the model to do both unpredictably, and WER against either standard stays high.

Before sizing, fix the standard and measure it:

  • Choose verbatim or clean transcription per verbatim vs clean transcription standards, and write normalization rules for numbers, currency, dates and acronyms.
  • Ask suppliers for inter-annotator agreement or a double-transcribed sample, reported as WER between transcribers.
  • Check how redaction appears in audio and text (silence, tones, tags like [NAME]), since inconsistent redaction markers create training artifacts; see how audio redaction affects speech model training.
  • Confirm segment boundaries and timestamps; segments longer than 30 seconds need re-chunking for Whisper-style training.

Synthetic speech stretches real hours but does not replace them

Text-to-speech augmentation can add coverage of rare terms cheaply, but it does not substitute for real speakers, channels and spontaneous disfluencies. Treat synthetic audio as a supplement for vocabulary, and keep the evaluation set entirely real. The trade-offs are covered in real vs synthetic speech data for ASR.

Turning a sizing estimate into a data request

A good speech data request states a target range, the distribution you need inside it, and the evidence that justifies it. "We need 1,000 hours" invites generic audio; "40–80 hours of 8 kHz dual-channel support calls, at least 300 distinct customer speakers, 15% with background noise below 10 dB SNR, verbatim transcripts with entity tags" invites a match.

Include in the request:

  1. Target hours as a range, tied to your pilot curve.
  2. Minimum unique speakers and maximum hours per speaker.
  3. Channel, sample rate, bit depth and codec constraints.
  4. Accent, noise and overlap distribution.
  5. Transcription standard, normalization rules and redaction format.
  6. Intended use (fine-tuning, evaluation, or both) and the rights you need for it.

For the general version of this question across data types, see how many records AI labs want; for LLM text sizing, see how much data to fine-tune an LLM. The speech and audio data buyer's guide covers licensing and consent for recordings, and licensing voice and audio data describes the categories buyers request. If your pilot points to operational recordings such as support calls, you can describe the data to SourceX, which looks for US businesses that hold it.

Sourcing in-domain speech for ASR fine-tuning

SourceX sources operational datasets, including support and sales histories and new recordings of hands-on work, from US companies on request, and manages licensing for AI teams wherever they are based. Every dataset is rights-reviewed, personal details are removed or replaced before delivery, and nothing is contracted until the supplying company agrees. Describe the speech data your ASR model needs.

Frequently asked questions

Can I fine-tune Whisper with one hour of audio?

You can run the training job, but an hour from a handful of speakers mostly teaches the model those voices. One hour is better used as an evaluation set, with vocabulary handled through prompts or text-based biasing until you have tens of hours from many speakers [1].

Should I count hours before or after silence trimming?

Count speech-active hours after voice activity detection, and ask suppliers to report both raw and speech-active duration. Call recordings with long hold music or silence can contain far less usable speech than their file length suggests.

Do I need separate hours for each language or dialect?

Yes, budget per language and treat strong dialect differences as separate slices with their own eval sets. Cross-lingual transfer helps, but per-language sizing is covered in multilingual ASR training data sourcing.

Sources

  1. arXiv, "Automatic Speech Recognition for Greek Medical Dictation" (2025). https://arxiv.org/pdf/2509.23550
  2. arXiv, "The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage" (2021). https://arxiv.org/pdf/2111.09344
  3. arXiv, "Towards Data-efficient Modeling for Wake Word Spotting" (2020). https://arxiv.org/pdf/2010.06659
  4. arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
  5. Baidu Research (arXiv), "Deep Learning Scaling is Predictable, Empirically" (2017). https://arxiv.org/abs/1712.00409v1
  6. Mozilla (arXiv), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data