Skip to content

Speech and audio data

Telephony vs Wideband Audio: Training ASR on 8 kHz Call Recordings

Quick answer

Phone audio sampled at 8 kHz can only carry content up to 4 kHz, while most modern ASR models are trained on 16 kHz audio that reaches 8 kHz, so a wideband model meets telephony audio with its upper band missing [1]. The practical fixes are a dedicated narrowband model, mixed training with upsampled telephony data, simulated telephony augmentation of wideband speech, or bandwidth extension. Whichever you choose, the dataset you buy must declare its native sample rate and codec history, because audio upsampled before delivery hides the problem rather than solving it.

By SourceX Editorial · Updated

Why 8 kHz call audio breaks wideband ASR models

The core mismatch is bandwidth, not file format: an 8 kHz signal stops at the 4 kHz Nyquist limit, and a model trained on 0-8 kHz features learns to rely on cues that phone audio has removed [1]. The most exposed sounds are fricatives such as /s/, /f/ and /th/, whose distinguishing energy sits largely above 4 kHz, so "fifty" versus "sixty" and plural endings become harder to resolve. Legacy telephone paths are narrower still, typically band-limiting speech to roughly 300-3,400 Hz, which also thins low-frequency pitch energy.

The damage shows up in the front end before the network sees anything. A 16 kHz model with 80 mel bins spanning 0-8 kHz will see roughly the top quarter of its filterbank (about 20 of 80 bins) near zero on upsampled phone audio, which shifts per-utterance normalization statistics and can push the decoder toward deletions. Speaker recognition suffers the same failure: wideband speaker embeddings trained on 16 kHz audio degrade on 8 kHz telephony input, which is why bandwidth expansion has been studied for speaker verification as well as ASR [4].

Codecs compound the bandwidth gap. G.711 (mu-law in North America, A-law elsewhere) is the clean narrowband baseline, but calls routed through mobile networks or VoIP often pass through lower-bitrate codecs such as G.729 or AMR-NB, then get transcoded again on recording. Each hop adds quantization artifacts, packet-loss concealment and occasional clipping that a model trained on clean wideband read speech has never encountered.

Four training strategies for narrowband telephony data

There is no single correct strategy; the right choice depends on whether your production traffic is purely telephony or mixed with wideband sources such as apps, WebRTC and meeting audio. The table below summarizes the trade-offs teams usually weigh.

Illustrative example: invented to show structure; it does not describe an available dataset.

StrategyHow it worksBest whenMain failure mode
Separate narrowband modelTrain or fine-tune a model at 8 kHz on native telephony audioTraffic is almost entirely PSTN or contact-centerTwo models to maintain; routing errors when rate metadata is wrong
Upsample and mixResample 8 kHz audio to 16 kHz and train one model on bothMixed telephony and wideband trafficModel may learn "empty upper band = phone" shortcuts and underuse wideband detail
Simulated telephony augmentationDownsample wideband speech, band-pass it and run it through codec simulationLittle native phone audio availableSimulated codecs miss real network impairments, crosstalk and hold music
Bandwidth extension or channel-aware pretrainingPredict missing upper-band features, or pretrain with channel awarenessLarge unlabeled telephony pool, research capacityAdded complexity; extension can hallucinate cues

Separate acoustic models per sampling rate are the traditional approach [2]. They give the cleanest narrowband performance, but you pay for it in duplicated training, evaluation and deployment pipelines, and the same patent describes sampling-rate-independent recognition as a way out of that duplication [2].

Upsampling and mixing is the default in many modern toolkits. NVIDIA's NeMo telephony tutorial upsamples telephone corpora to 16 kHz before training so a single model can serve both 8 kHz and 16 kHz input [3]. Upsampling adds no information above 4 kHz; it only makes the tensors compatible, so the ratio of native telephony hours to wideband hours in the mix decides how well the model handles phone calls.

Simulated telephony augmentation is the cheapest way to harden a wideband model. A typical recipe downsamples to 8 kHz, applies a 300-3,400 Hz band-pass, encodes through G.711, GSM or AMR-NB with ffmpeg or sox, then resamples back to 16 kHz. It helps, but it does not reproduce real carrier impairments, agent headsets, speakerphone echo or overlapping speech, so it complements native call audio rather than replacing it; see our guide to real vs synthetic speech data for ASR for where simulation falls short.

Bandwidth extension and channel-aware learning are the research frontier. One approach extends features in the feature domain and rebalances the spectrum as a data augmentation step [1]. Another pretrains a joint encoder-decoder self-supervised model with channel awareness so telephonic and wideband speech share representations without pretending they are the same [5].

How to upsample 8 kHz audio to 16 kHz without fooling yourself

Upsampling is safe for compatibility and unsafe for bookkeeping. Use a proper polyphase or sinc resampler (sox, torchaudio's resample, or soxr through ffmpeg) rather than naive sample repetition, which creates spectral images above 4 kHz that look like real speech energy to the model. Then record that the file was upsampled, because downstream nobody can tell a resampled 16 kHz file from a wideband one by its header.

Three checks catch most problems:

  1. Spectral cutoff check. Compute the long-term average spectrum per file; native wideband speech has energy past 4 kHz, while upsampled narrowband shows a sharp cliff near 3.4-4 kHz.
  2. Header versus content check. Compare the WAV or FLAC header rate with the detected cutoff; a "16 kHz" file with a 4 kHz cliff was upsampled somewhere upstream.
  3. Split evaluation. Report WER separately for native narrowband, native wideband and simulated telephony test sets; a blended WER hides regressions on the slice you care about.

Vendor-reported WER floors for telephony models circulate widely, including accounts of training on millions of hours of phone speech [6]. Treat such numbers as market claims, not benchmarks, and reproduce them on your own held-out calls before relying on them.

What to require from a telephony speech dataset supplier

The single most useful requirement is native sample-rate and codec metadata for every file, because it lets you choose any of the strategies above later. Audio that was upsampled or transcoded before delivery forces you to reverse-engineer its history from spectrograms, and some information, such as which lossy codec ran first, cannot be recovered at all. The general format choices are covered in audio file specs for speech datasets; this section is specific to phone audio.

Ask for the recording as captured by the call recording system, with any conversion documented rather than applied. Prefer lossless containers (WAV or FLAC) holding the original 8 kHz PCM or the decoded G.711 stream, and ask whether agent and customer legs are on separate channels; stereo legs matter for diarization, as explained in dual-channel call recordings for ASR and diarization. If the archive stores compressed audio, ask which codec and bitrate it used, and whether any segments were re-encoded during migration between platforms.

A per-file manifest record makes these requirements testable at acceptance.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"audio_id": "call_000184_agent", "path": "audio/call_000184.flac", "channel": 0, "speaker_role": "agent", "native_sample_rate_hz": 8000, "stored_sample_rate_hz": 8000, "resampled": false, "codec_chain": ["G.711 mu-law", "PCM decode"], "network_path": "PSTN", "measured_cutoff_hz": 3600, "duration_s": 412.7, "transcript_path": "transcripts/call_000184.json", "redaction_method": "spoken PAN masked with tone", "language": "en-US"}

Fields such as resampled, codec_chain and measured_cutoff_hz are the ones most often missing in practice. Segment timing and transcript alignment conventions are covered in packaging speech datasets with manifests, and the broader requirement set for call audio lives in the contact-center audio requirements template.

Redaction, card data and the narrowband channel

Redaction interacts with bandwidth more than buyers expect. Tones or silence inserted over spoken card numbers, account numbers and names change the acoustic statistics of narrowband audio, and the PCI Security Standards Council's supplement on telephone-based payment card data specifically addresses cardholder data held in call recording systems [7]. Ask suppliers how redaction was applied (silence, tone, noise or transcript-only), whether it was applied before or after any resampling, and how redacted spans are tagged in the transcript.

Redaction applied to an upsampled copy can drift out of alignment with the native-rate file when spans are stored as sample offsets rather than seconds, so request redaction at the native rate with time-based span tags. For the transcript side, see redacting spoken PII from call recordings, and for training effects, how audio redaction affects speech model training.

Buyer checklist for 8 kHz call audio

Use this checklist when scoping a telephony speech dataset request; it turns the technical points above into acceptance criteria.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Native sample rate stated per file, and no undocumented upsampling or downsampling.
  • Codec chain recorded (G.711, G.729, AMR-NB, Opus or other), including archive re-encodes.
  • Lossless delivery container for audio that was captured losslessly.
  • Agent and customer channels separated where the recording system captured them.
  • Network path labeled where known (PSTN, mobile, VoIP, softphone).
  • Redaction method and timing documented relative to any resampling.
  • A held-out test slice of native narrowband calls reserved for evaluation, kept separate from training audio; see voice agent evaluation sets.
  • Noise and SNR conditions described, using the approach in noisy speech data for robust ASR.

Hours alone are a weak specification for this problem. Fifty hours of correctly labeled native 8 kHz calls with a documented codec chain is often more useful for adaptation than a larger pool of mixed-rate audio whose history is unknown. For the broader picture of speech categories, rights and licensing, start from the speech and audio data buyer's guide.

Sourcing native telephony audio for ASR training

Native call recordings sit inside companies' contact-center and sales platforms, which is why they rarely appear in open corpora with clean commercial licenses. SourceX sources operational datasets, including support and sales call histories, from US companies, and manages the licensing process; data is sourced on request rather than held in stock, and a request does not guarantee a match. You describe the data you need, such as native 8 kHz rate, codec metadata and channel layout, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. You can see related owner pages for contact-center call recordings and sales call recordings, or describe your telephony audio requirements.

Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Personal details such as names, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.

Request native 8 kHz call audio for ASR training

If your model needs real phone-line audio with its sample rate and codec history documented, describe the rate, channels, codecs and redaction you need. SourceX looks for US companies that hold matching recordings, reviews rights, and manages the license and delivery through private, access-controlled workflows. Start a buyer request at sourcex.si/buyers.

Sources

  1. United States Patent and Trademark Office, "Feature domain bandwidth extension and spectral rebalance for ASR data augmentation (US Patent 12148437)". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/12148437
  2. United States Patent and Trademark Office, "Sampling rate independent speech recognition (US Patent 7983916)". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/7983916
  3. Hugging Face (mirror of NVIDIA NeMo tutorial), "ASR for telephony speech (NeMo tutorial notebook)". https://huggingface.co/Respair/NeMo_Canary/blame/main/tutorials/asr/ASR_for_telephony_speech.ipynb
  4. Odyssey 2020 (SuperLectures), "Speech bandwidth expansion for speaker recognition on telephony audio" (2020). https://www.superlectures.com/odyssey2020/speech-bandwidth-expansion-for-speaker-recognition-on-telephony-audio
  5. arXiv, "Channel-Aware Pretraining of Joint Encoder-Decoder Self-Supervised Model for Telephonic-Speech ASR" (2022). https://arxiv.org/pdf/2211.01669
  6. Gnani.ai, "How we train acoustic models on 14 million hours of telephonic speech". https://www.gnani.ai/resources/research/how-we-train-acoustic-models-on-14-million-hours-of-telephonic-speech
  7. destinationCRM, "PCI Council Releases Supplemental Guidance for Protecting Telephone-Based Payment Card Data" (2011). https://www.destinationcrm.com/Articles/CRM-News/Daily-News/PCI-Council-Releases-Supplemental-Guidance-for-Protecting-Telephone-Based-Payment-Card-Data-74450.aspx

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data