Skip to content

Speech and audio data

Audio File Specs for Speech Datasets: Sample Rate, Bit Depth, Channels and Codecs

Quick answer

For most ASR training, ask for 16 kHz, 16-bit linear PCM in mono WAV or FLAC, one file per recording, with channels kept separate when the source had them. Telephony audio should arrive at its native 8 kHz, not upsampled. TTS and non-speech audio need higher rates, often 22.05, 24, 44.1 or 48 kHz. Above all, request the original codec or a lossless copy, plus a transcoding history, because resampling, lossy re-encoding and channel downmixing cannot be undone after delivery.

By SourceX Editorial · Updated

Which sample rate should a speech dataset use?

Choose the sample rate your model consumes at inference, and require that the source was natively captured at or above it. A sample rate of 16 kHz captures frequencies up to 8 kHz, which covers most of the speech cues ASR depends on, and it is the de facto ASR target; research releases such as CHILDES-Aligned ship as 16 kHz mono WAV [1]. Higher rates only help if the model's front end uses the extra band.

The rule of thumb by use case:

  • ASR and keyword spotting: 16 kHz is standard. Delivering 44.1 or 48 kHz is fine as long as it is native, since you can downsample cleanly yourself.
  • Telephony ASR: Narrowband audio only carries content up to 4 kHz, and models trained on wideband audio meet a channel mismatch on phone calls [3]. Teams often resample 8 kHz audio to 16 kHz so a single model front end handles both [2], but that resampling is your pipeline's job, not the supplier's.
  • TTS and voice cloning: 22.05 kHz or higher, commonly 24 or 48 kHz, because the vocoder must reproduce the high-frequency detail listeners hear.
  • Audio classification and machine sounds: Follow the task. Machine-condition datasets document multi-microphone capture setups and their recording parameters alongside the audio [5].

The failure mode to catch is the "fake wideband" file: 8 kHz call audio upsampled to 16 kHz and labeled as wideband. A spectrogram shows a hard ceiling at 4 kHz with nothing above it. Training on such files as if they were wideband teaches the model that the upper band is always empty. For the full telephony trade-off, see telephony vs wideband audio for ASR training.

What bit depth do speech recordings need?

Sixteen-bit integer PCM is enough for speech model training, and it is what most ASR toolkits expect. It offers about 96 dB of theoretical dynamic range, far more than conversational speech recorded on headsets or phones actually uses. Accept 24-bit or 32-bit float masters if that is what the supplier holds, because downconverting with dither is trivial on your side.

Bit depth problems show up in different ways than the header suggests:

  • Low effective resolution: A 16-bit file made from an 8-bit G.711 mu-law or A-law telephony stream still carries 8-bit companded resolution. That is acceptable for telephony models, but it should be declared, not hidden.
  • Clipping: Samples pinned at full scale (32767 or -32768) mean the recording gain was too high. Clipping cannot be repaired by changing bit depth.
  • Very low level: Speech peaking at -40 dBFS wastes most of the available range and usually signals a gain or routing error at capture.

WAV, FLAC or MP3: which container and codec?

Request lossless audio: WAV (RIFF) with PCM data or FLAC, and treat lossy formats as acceptable only when they are the original capture codec. FLAC is an open, IETF-standardized lossless codec that compresses multichannel PCM at common speech bit depths, typically cutting storage substantially without changing a single sample. Plain WAV is simplest to read but its 32-bit size field caps files at 4 GB; long multichannel recordings may need RF64 or splitting.

MP3 and other lossy codecs remove information selectively. Their psychoacoustic models discard spectral detail judged inaudible to people, which is not the same as detail irrelevant to an acoustic model, and the damage grows at low bitrates and in noisy audio. Unvoiced sounds such as fricatives, already weak in telephone audio, tend to suffer first. Vendors still deliver MP3 alongside WAV [4], so ask which one is closer to the source.

The worst case is the transcoding chain: a call recorded in G.729 or Opus, stored as MP3 by a recording platform, then converted to WAV for delivery. The WAV looks lossless but inherits every earlier loss, and a second lossy pass compounds artifacts. Ask for the original codec where possible, and require a written transcoding history either way.

How should channels be delivered?

Keep multichannel recordings multichannel, and document what each channel is. Contact-center platforms often record the agent and customer on separate channels; that separation is free ground truth for speaker attribution and lets you measure crosstalk and overlap precisely. Downmixing to mono before delivery throws it away permanently.

Specify the channel layout explicitly:

  • Dual-channel calls: State which channel is agent and which is customer, and whether this is consistent across the whole delivery. Inconsistent mapping is a common, silent error that corrupts diarization labels.
  • Mono mixes: Acceptable for single-speaker dictation or when the source was mono. Ask whether a mono file is native or a downmix.
  • Microphone arrays: Keep all channels plus the array geometry. Beamforming and far-field work depend on it, as do machine-sound tasks recorded with multiple microphones [5].

If your training pipeline wants mono, derive it yourself. Diarization guidance and RTTM conventions are covered in speaker diarization training data.

What do conversions and redaction lose?

Every conversion between capture and delivery either preserves the signal or destroys part of it, and you cannot tell which from the file header alone. The table below summarizes what each operation costs and how to detect it.

OperationWhat is lostReversible?How to detect
Downsampling (48 to 16 kHz)Content above the new Nyquist limitNo, but usually intendedHeader rate; spectral ceiling
Upsampling (8 to 16 kHz)Nothing, but nothing is gainedNot neededEmpty band above 4 kHz in spectrogram
Lossy encode (MP3, AAC, Opus)Masked spectral detail, transientsNoCodec in metadata; spectral holes and cutoffs
Lossy to WAV re-wrapNothing further, but hides prior lossNoTranscoding history; spectral analysis
Stereo to mono downmixSpeaker separation, spatial cuesNoChannel count; request source layout
Bit-depth reduction (24 to 16)Low-level detail below the noise floorNo, rarely mattersHeader; dither notes
Loudness normalizationOriginal level relationshipsPartly, if gain is loggedAsk for the gain value applied
Redaction (silence, tone, bleep)The redacted speech and the boundary contextNoRedaction log with timestamps

Redaction deserves its own attention because it changes the waveform on purpose. Silence, beeps and noise replacements each leave different artifacts that a model can learn; see how audio redaction affects speech model training.

What should an audio acceptance check include?

An acceptance check should verify technical properties on every file and listen to a sample, before the data enters training. Run ffprobe or soundfile on the full delivery to confirm sample rate, bit depth, channel count, codec and duration against the manifest, then compute signal metrics. Loudness can be measured per ITU-R BS.1770 (LUFS) so files from different sources can be compared consistently.

Illustrative example: invented to show structure; it does not describe an available dataset.

Speech audio delivery spec and acceptance checklist

FieldRequirement (ASR example)Rejection trigger
Container / codecWAV (PCM) or FLAC; original codec copy if lossy at sourceLossy re-encode without history
Sample rate16 kHz native; 8 kHz native for telephonyUpsampled audio labeled wideband
Bit depth16-bit PCM or higherUndeclared 8-bit companded source
ChannelsSource layout preserved; channel map documentedUndocumented or inconsistent agent/customer mapping
DurationMatches manifest within 10 msTruncated or zero-length files
ClippingUnder 0.1% of samples at full scale per fileSustained clipping in speech regions
SilenceLeading/trailing silence under 2 s; speech ratio loggedFiles that are mostly silence or hold music
LoudnessIntegrated LUFS reported per fileSpeech peaking below -40 dBFS
ProvenanceCapture device or platform, capture codec, transcoding stepsMissing transcoding history
RedactionMethod and per-file timestamps loggedUnlogged gaps or tones

A per-file record in the manifest might look like this:

{"file": "call_000123.flac", "sample_rate": 8000, "bit_depth": 16, "channels": 2,
 "channel_map": {"0": "agent", "1": "customer"}, "capture_codec": "G.711 mu-law",
 "transcode_chain": ["G.711 mu-law -> PCM16 -> FLAC"], "duration_s": 412.36,
 "integrated_lufs": -24.1, "clipped_sample_pct": 0.0, "redaction": "tone", "redaction_spans": 3}

Size the listening sample with the method in sample sizes for estimating a dataset's error rate, and pair the checks with manifest conventions from packaging speech datasets. For very large deliveries, WebDataset tar shards keep FLAC files and their JSON sidecars together.

How do file specs fit into a speech data request?

Write audio specs into the request before sourcing begins, because a supplier can only deliver what its systems actually captured. If a contact center's recorder stored mono MP3 at 8 kHz, no export setting will produce stereo 16 kHz lossless audio. State your hard requirements (native rate, channel separation, lossless or original codec) and your tolerances separately, so a near-match is not rejected for a parameter you can convert yourself.

Then validate on a sample: measure WER lift on a pilot sample before committing to a full order. For contact-center audio, the contact-center audio requirements template covers the non-technical fields as well. The broader speech and audio data buyer's guide and the general answer on what file formats AI buyers accept put these specs in context.

SourceX sources operational datasets, including support and sales histories, from US companies on request rather than from stock, so a request does not guarantee a match. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, so ask how redaction was applied to the audio. Buyers can describe the audio they need to SourceX, or browse voice and audio data and contact-center call recordings.

Source speech audio with the specs you need

SourceX looks for US businesses that hold the speech data you describe, rights-reviews each dataset and delivers it under a license defining records, uses, term and delivery. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Describe your speech audio requirements to SourceX.

Sources

  1. arXiv, "CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling" (2026). https://arxiv.org/pdf/2607.03670
  2. Hugging Face (NeMo tutorial mirror), "ASR for telephony speech (NeMo tutorial notebook)". https://huggingface.co/Respair/NeMo_Canary/blame/main/tutorials/asr/ASR_for_telephony_speech.ipynb
  3. arXiv, "Channel-Aware Pretraining of Joint Encoder-Decoder Self-Supervised Model for Telephonic-Speech ASR" (2022). https://arxiv.org/pdf/2211.01669
  4. Unidata, "Medical Conversations English". https://unidata.pro/datasets/medical-conversations-english/
  5. DCASE Community, "DCASE 2020 Task 2: Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring" (2020). https://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data