Speech and audio data
Audio File Specs for Speech Datasets: Sample Rate, Bit Depth, Channels and Codecs
Quick answer
For most ASR training, ask for 16 kHz, 16-bit linear PCM in mono WAV or FLAC, one file per recording, with channels kept separate when the source had them. Telephony audio should arrive at its native 8 kHz, not upsampled. TTS and non-speech audio need higher rates, often 22.05, 24, 44.1 or 48 kHz. Above all, request the original codec or a lossless copy, plus a transcoding history, because resampling, lossy re-encoding and channel downmixing cannot be undone after delivery.
By SourceX Editorial · Updated
Which sample rate should a speech dataset use?
Choose the sample rate your model consumes at inference, and require that the source was natively captured at or above it. A sample rate of 16 kHz captures frequencies up to 8 kHz, which covers most of the speech cues ASR depends on, and it is the de facto ASR target; research releases such as CHILDES-Aligned ship as 16 kHz mono WAV [1]. Higher rates only help if the model's front end uses the extra band.
The rule of thumb by use case:
- ASR and keyword spotting: 16 kHz is standard. Delivering 44.1 or 48 kHz is fine as long as it is native, since you can downsample cleanly yourself.
- Telephony ASR: Narrowband audio only carries content up to 4 kHz, and models trained on wideband audio meet a channel mismatch on phone calls [3]. Teams often resample 8 kHz audio to 16 kHz so a single model front end handles both [2], but that resampling is your pipeline's job, not the supplier's.
- TTS and voice cloning: 22.05 kHz or higher, commonly 24 or 48 kHz, because the vocoder must reproduce the high-frequency detail listeners hear.
- Audio classification and machine sounds: Follow the task. Machine-condition datasets document multi-microphone capture setups and their recording parameters alongside the audio [5].
The failure mode to catch is the "fake wideband" file: 8 kHz call audio upsampled to 16 kHz and labeled as wideband. A spectrogram shows a hard ceiling at 4 kHz with nothing above it. Training on such files as if they were wideband teaches the model that the upper band is always empty. For the full telephony trade-off, see telephony vs wideband audio for ASR training.
What bit depth do speech recordings need?
Sixteen-bit integer PCM is enough for speech model training, and it is what most ASR toolkits expect. It offers about 96 dB of theoretical dynamic range, far more than conversational speech recorded on headsets or phones actually uses. Accept 24-bit or 32-bit float masters if that is what the supplier holds, because downconverting with dither is trivial on your side.
Bit depth problems show up in different ways than the header suggests:
- Low effective resolution: A 16-bit file made from an 8-bit G.711 mu-law or A-law telephony stream still carries 8-bit companded resolution. That is acceptable for telephony models, but it should be declared, not hidden.
- Clipping: Samples pinned at full scale (32767 or -32768) mean the recording gain was too high. Clipping cannot be repaired by changing bit depth.
- Very low level: Speech peaking at -40 dBFS wastes most of the available range and usually signals a gain or routing error at capture.
WAV, FLAC or MP3: which container and codec?
Request lossless audio: WAV (RIFF) with PCM data or FLAC, and treat lossy formats as acceptable only when they are the original capture codec. FLAC is an open, IETF-standardized lossless codec that compresses multichannel PCM at common speech bit depths, typically cutting storage substantially without changing a single sample. Plain WAV is simplest to read but its 32-bit size field caps files at 4 GB; long multichannel recordings may need RF64 or splitting.
MP3 and other lossy codecs remove information selectively. Their psychoacoustic models discard spectral detail judged inaudible to people, which is not the same as detail irrelevant to an acoustic model, and the damage grows at low bitrates and in noisy audio. Unvoiced sounds such as fricatives, already weak in telephone audio, tend to suffer first. Vendors still deliver MP3 alongside WAV [4], so ask which one is closer to the source.
The worst case is the transcoding chain: a call recorded in G.729 or Opus, stored as MP3 by a recording platform, then converted to WAV for delivery. The WAV looks lossless but inherits every earlier loss, and a second lossy pass compounds artifacts. Ask for the original codec where possible, and require a written transcoding history either way.
How should channels be delivered?
Keep multichannel recordings multichannel, and document what each channel is. Contact-center platforms often record the agent and customer on separate channels; that separation is free ground truth for speaker attribution and lets you measure crosstalk and overlap precisely. Downmixing to mono before delivery throws it away permanently.
Specify the channel layout explicitly:
- Dual-channel calls: State which channel is agent and which is customer, and whether this is consistent across the whole delivery. Inconsistent mapping is a common, silent error that corrupts diarization labels.
- Mono mixes: Acceptable for single-speaker dictation or when the source was mono. Ask whether a mono file is native or a downmix.
- Microphone arrays: Keep all channels plus the array geometry. Beamforming and far-field work depend on it, as do machine-sound tasks recorded with multiple microphones [5].
If your training pipeline wants mono, derive it yourself. Diarization guidance and RTTM conventions are covered in speaker diarization training data.
What do conversions and redaction lose?
Every conversion between capture and delivery either preserves the signal or destroys part of it, and you cannot tell which from the file header alone. The table below summarizes what each operation costs and how to detect it.
| Operation | What is lost | Reversible? | How to detect |
|---|---|---|---|
| Downsampling (48 to 16 kHz) | Content above the new Nyquist limit | No, but usually intended | Header rate; spectral ceiling |
| Upsampling (8 to 16 kHz) | Nothing, but nothing is gained | Not needed | Empty band above 4 kHz in spectrogram |
| Lossy encode (MP3, AAC, Opus) | Masked spectral detail, transients | No | Codec in metadata; spectral holes and cutoffs |
| Lossy to WAV re-wrap | Nothing further, but hides prior loss | No | Transcoding history; spectral analysis |
| Stereo to mono downmix | Speaker separation, spatial cues | No | Channel count; request source layout |
| Bit-depth reduction (24 to 16) | Low-level detail below the noise floor | No, rarely matters | Header; dither notes |
| Loudness normalization | Original level relationships | Partly, if gain is logged | Ask for the gain value applied |
| Redaction (silence, tone, bleep) | The redacted speech and the boundary context | No | Redaction log with timestamps |
Redaction deserves its own attention because it changes the waveform on purpose. Silence, beeps and noise replacements each leave different artifacts that a model can learn; see how audio redaction affects speech model training.
What should an audio acceptance check include?
An acceptance check should verify technical properties on every file and listen to a sample, before the data enters training. Run ffprobe or soundfile on the full delivery to confirm sample rate, bit depth, channel count, codec and duration against the manifest, then compute signal metrics. Loudness can be measured per ITU-R BS.1770 (LUFS) so files from different sources can be compared consistently.
Illustrative example: invented to show structure; it does not describe an available dataset.
Speech audio delivery spec and acceptance checklist
| Field | Requirement (ASR example) | Rejection trigger |
|---|---|---|
| Container / codec | WAV (PCM) or FLAC; original codec copy if lossy at source | Lossy re-encode without history |
| Sample rate | 16 kHz native; 8 kHz native for telephony | Upsampled audio labeled wideband |
| Bit depth | 16-bit PCM or higher | Undeclared 8-bit companded source |
| Channels | Source layout preserved; channel map documented | Undocumented or inconsistent agent/customer mapping |
| Duration | Matches manifest within 10 ms | Truncated or zero-length files |
| Clipping | Under 0.1% of samples at full scale per file | Sustained clipping in speech regions |
| Silence | Leading/trailing silence under 2 s; speech ratio logged | Files that are mostly silence or hold music |
| Loudness | Integrated LUFS reported per file | Speech peaking below -40 dBFS |
| Provenance | Capture device or platform, capture codec, transcoding steps | Missing transcoding history |
| Redaction | Method and per-file timestamps logged | Unlogged gaps or tones |
A per-file record in the manifest might look like this:
{"file": "call_000123.flac", "sample_rate": 8000, "bit_depth": 16, "channels": 2,
"channel_map": {"0": "agent", "1": "customer"}, "capture_codec": "G.711 mu-law",
"transcode_chain": ["G.711 mu-law -> PCM16 -> FLAC"], "duration_s": 412.36,
"integrated_lufs": -24.1, "clipped_sample_pct": 0.0, "redaction": "tone", "redaction_spans": 3}
Size the listening sample with the method in sample sizes for estimating a dataset's error rate, and pair the checks with manifest conventions from packaging speech datasets. For very large deliveries, WebDataset tar shards keep FLAC files and their JSON sidecars together.
How do file specs fit into a speech data request?
Write audio specs into the request before sourcing begins, because a supplier can only deliver what its systems actually captured. If a contact center's recorder stored mono MP3 at 8 kHz, no export setting will produce stereo 16 kHz lossless audio. State your hard requirements (native rate, channel separation, lossless or original codec) and your tolerances separately, so a near-match is not rejected for a parameter you can convert yourself.
Then validate on a sample: measure WER lift on a pilot sample before committing to a full order. For contact-center audio, the contact-center audio requirements template covers the non-technical fields as well. The broader speech and audio data buyer's guide and the general answer on what file formats AI buyers accept put these specs in context.
SourceX sources operational datasets, including support and sales histories, from US companies on request rather than from stock, so a request does not guarantee a match. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, so ask how redaction was applied to the audio. Buyers can describe the audio they need to SourceX, or browse voice and audio data and contact-center call recordings.
Source speech audio with the specs you need
SourceX looks for US businesses that hold the speech data you describe, rights-reviews each dataset and delivers it under a license defining records, uses, term and delivery. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Describe your speech audio requirements to SourceX.
Sources
- arXiv, "CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling" (2026). https://arxiv.org/pdf/2607.03670
- Hugging Face (NeMo tutorial mirror), "ASR for telephony speech (NeMo tutorial notebook)". https://huggingface.co/Respair/NeMo_Canary/blame/main/tutorials/asr/ASR_for_telephony_speech.ipynb
- arXiv, "Channel-Aware Pretraining of Joint Encoder-Decoder Self-Supervised Model for Telephonic-Speech ASR" (2022). https://arxiv.org/pdf/2211.01669
- Unidata, "Medical Conversations English". https://unidata.pro/datasets/medical-conversations-english/
- DCASE Community, "DCASE 2020 Task 2: Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring" (2020). https://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.