Speech and audio data
Noisy Speech Data for Robust ASR: Specifying Noise Types and SNR Coverage
Quick answer
A useful noisy speech dataset for ASR is specified along two axes: which noise types occur in your deployment (stationary machinery, competing voices, music, impulsive alarms, channel distortion) and what signal-to-noise ratio range each type spans, typically 0 to 25 dB in published simulation recipes [1]. Train on real noisy recordings plus clean speech mixed with matched noise, label every file with environment and estimated SNR, and evaluate only on held-out real recordings, because clean-benchmark accuracy hides noise failures [2].
By SourceX Editorial · Updated
Which noise types should a robust ASR data spec name?
A spec should name noise by its acoustic behavior, not just by venue, because "warehouse" can mean steady HVAC hum at one moment and a forklift reversing alarm the next. The taxonomy below is the working vocabulary most ASR teams converge on; it is a practitioner framing, not a formal standard. Ask suppliers to tag each file with one primary class and any secondary classes present.
- Stationary noise: HVAC, refrigeration compressors, engine and road noise at cruise, conveyor drone. Spectrally stable, so front-end enhancement handles it relatively well, but it masks low-energy consonants.
- Non-stationary noise: babble from nearby talkers, cart and pallet clatter, door slams, PA chatter. This is the hardest class for both enhancement and ASR, because babble shares the speech spectrum.
- Music and media: in-store playlists, vehicle radio, TV in a break room. Lyrics create insertion errors; a voice agent can transcribe the song.
- Impulsive noise: reversing beepers, scanner beeps, alarms, tool strikes, horn blasts. Short, loud and often clipped by automatic gain control.
- Channel and device noise: codec artifacts, packet loss, wind on a lapel mic, handling noise on a rugged handheld, radio squelch on push-to-talk. Treat this as a separate axis that interacts with sample rate and codec choices covered in audio file specs for speech datasets and telephony vs wideband audio.
Distance and reverberation are a different problem with different data. If your microphones sit meters from the talker, read far-field speech data: real vs simulated rooms alongside this page.
What SNR range should training and test audio cover?
Cover the SNR range your deployment actually produces, with deliberate weight at the low end where errors concentrate, rather than a uniform spread. SNR in decibels is 10·log10 of speech power over noise power, so 0 dB means speech and noise are equally loud and 20 dB means speech carries 100 times the noise power. A published far-field recipe mixed noise into roughly two million simulated utterances at SNRs from 0 to 25 dB [1], which is a sensible default band for added-noise augmentation.
Real deployments do not respect that band. A worker shouting over a compactor can sit below 0 dB, while a quiet car cabin at idle can exceed 25 dB. Vendor guides publish SNR ranges by environment such as home, car and industrial settings [3]; treat those numbers as self-reported starting points and measure your own sites with a sound level meter and a few hours of pilot recordings.
A practical spec splits the range into bands and sets a minimum share of hours per band for each noise class. Typical bands are below 0 dB, 0 to 5, 5 to 10, 10 to 15, 15 to 20 and above 20 dB. Report word error rate (WER) per band and per class, never just a pooled number, so a regression at 0 to 5 dB babble is visible even if the average improves.
How do you get SNR labels for real recordings?
For real recordings you cannot compute true SNR, because speech and noise were never captured separately, so you request an estimated SNR plus the method used to estimate it. Common blind estimators include WADA-SNR and approaches built on voice activity detection that compare energy in speech frames with energy in non-speech frames. Each method has known biases with babble and music, so pin one method and version across the whole corpus.
The exception is a paired setup, where the talker wears a close-talk microphone while a device microphone captures the noisy signal, as in some CHiME challenge recordings. That close-talk channel gives a near-clean reference, which makes SNR estimation and enhancement evaluation far more reliable; ask whether a supplier's capture setup can include one.
Real noisy recordings or clean speech plus added noise?
Use both: added noise gives you controllable SNR coverage and cheap scale, while real noisy recordings carry the acoustic and behavioral effects simulation misses. Real recordings capture the Lombard effect (people speak louder and change pitch in noise), clipping from automatic gain control, device placement, and the correlation between noise and speech content, such as a picker reading a SKU over a beeping scanner. Mixing studio speech with a noise file reproduces none of that.
Added noise has its own data supply chain. Teams commonly draw on public music, speech and noise banks such as MUSAN, distributed through OpenSLR, plus noise recorded at their own sites. Check the license of every noise bank as carefully as the speech itself; the open speech corpora license audit covers the attribution and commercial-use questions.
A practical hybrid design: simulate broadly across a controlled 0 to 25 dB band for training, as the far-field recipe did [1], record background noise at the target sites to mix with clean speech, and hold real recordings back for evaluation. Simulation widens coverage; real audio keeps the test honest. For broader trade-offs between generated and recorded audio, see real vs synthetic speech data for ASR.
Why clean benchmarks hide noise failures
Clean-benchmark scores predict noisy performance poorly, so the test set must contain real noise from the target environment. A study of keyword spotting found that models with strong clean-benchmark accuracy produced high false alarm rates in real noisy environments, and the authors released a real-noise benchmark to expose the gap [2]. The same failure shows up in ASR as insertions during music and deletions during babble.
Three test-set rules prevent most surprises. Hold out whole sites, devices and speakers, not random utterances, so the model cannot memorize a specific compressor hum. Keep at least one test slice that is purely real noisy audio with no added noise. Audit test transcripts, because label errors in widely used test sets average at least 3.3% across audited vision, NLP and audio benchmarks [4], and noisy audio is harder to transcribe correctly than clean audio.
For keyword spotting and voice agents, add metrics beyond WER: false accepts per hour on long noise-only recordings, and missed commands per noise class. For transcription conventions in noisy audio, such as how to tag unintelligible spans and background speech, align with your verbatim vs clean transcription standard.
Per-file metadata to request
Every file should carry enough metadata to rebuild any SNR band or noise slice without re-listening to audio. ISO/IEC 5259-3 frames data quality as a managed process with defined requirements [5], and a per-file schema is how you make SNR coverage auditable rather than asserted. The record below shows the fields worth requiring.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"file_id": "wh-site03-dev2-000418.flac",
"duration_s": 7.84,
"sample_rate_hz": 16000,
"channels": 1,
"device_class": "rugged_handheld",
"mic_position": "chest_clip",
"environment": "warehouse_pick_aisle",
"noise_primary": "non_stationary_clatter",
"noise_secondary": ["impulsive_reverse_alarm", "stationary_hvac"],
"snr_est_db": 3.6,
"snr_method": "wada_snr_v1",
"noise_origin": "real",
"added_noise_source": null,
"clipping_ratio": 0.004,
"speech_style": "spontaneous_command",
"speaker_id_hash": "spk_7c1e",
"site_id": "site03",
"split": "test",
"transcript_standard": "verbatim_v2",
"redaction": {"applied": true, "method": "tone_replace"}
}
Two fields matter most for honest evaluation: noise_origin (real, added or both) and site_id, which lets you hold out entire locations. If audio was redacted, the replacement method changes the acoustics the model learns; see how audio redaction affects speech model training.
A coverage matrix you can hand to a supplier
The fastest way to specify noisy speech data is a matrix of noise class by SNR band, with target hours and the real-versus-added split in each cell. Fill it from site measurements, then let suppliers mark which cells they can cover with real recordings.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Noise class | < 0 dB | 0–10 dB | 10–20 dB | > 20 dB | Real share required | Test slice |
|---|---|---|---|---|---|---|
| Stationary (HVAC, engine) | 5 h | 20 h | 20 h | 10 h | 50% | Real only |
| Non-stationary (babble, clatter) | 10 h | 30 h | 15 h | 5 h | 70% | Real only |
| Music and media | 2 h | 10 h | 10 h | 3 h | 40% | Real only |
| Impulsive (alarms, beeps) | 3 h | 8 h | 5 h | 2 h | 80% | Real only |
| Channel (codec, wind, squelch) | n/a | 10 h | 10 h | 5 h | 60% | Real only |
Hours here are placeholders for structure; how many you actually need depends on your base model and target WER, covered in how many hours of audio you need to fine-tune ASR. Weight the low-SNR, non-stationary cells highest, because that is where deployed systems in warehouses, retail floors and vehicles typically fail.
Sourcing real noisy operational speech
Real noisy speech usually already exists inside operating businesses: recorded dispatch and push-to-talk traffic, support calls placed from vehicles or shop floors, and voice-picking or field-service recordings. Licensing that audio means reviewing consents and ownership, removing personal details, and fixing the allowed uses before delivery; the speech and audio data buyer's guide and off-the-shelf, custom or licensed speech data cover those decisions.
SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and it does not hold this audio in stock, so a request does not guarantee a match. Buyers describe the data they need, such as the coverage matrix above, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. You can describe a noisy-speech requirement to SourceX or compare options on the voice and audio data page and the voice agent training data use case. The real-world data glossary entry explains how operational data differs from collected or synthetic sets.
Find noisy speech data recorded where your model will run
SourceX sources operational speech datasets from US companies on request, rights-reviews each one for ownership and consents, and delivers it under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, with the method recorded. Describe the noise types and SNR coverage you need.
Frequently asked questions
Can I just add MUSAN noise to clean speech and skip real recordings?
You can for a first pass, but expect a gap on deployment audio. Added noise does not reproduce the Lombard effect, device handling, gain-control clipping or talker behavior in noise, and clean-benchmark scores have been shown to overstate robustness in real noise [2].
Should SNR be uniform across the training set?
No. Weight hours toward the low-SNR bands and the noise classes your sites actually produce, and keep some clean and high-SNR audio so the model does not degrade on quiet input.
How is this different from far-field or wake-word data?
Far-field data is about distance and reverberation, and wake-word data is about false accepts on a fixed phrase. This page covers the noise-type taxonomy and SNR coverage that sit underneath both; see far-field speech data for room acoustics.
Sources
- arXiv, "Spatial Attention for Far-field Speech Recognition with Deep Beamforming Neural Networks" (2019). https://arxiv.org/pdf/1911.02115
- IOS Press, "A Real-World Dataset for Benchmarking False Alarm Rate in Keyword Spotting" (2023). https://ebooks.iospress.nl/pdf/doi/10.3233/FAIA230695
- AIxBlock, "Noisy and Far-Field Speech Data for Robust ASR (2026)" (2026). https://aixblock.io/blogs/noisy-speech-data-asr
- arXiv / NeurIPS 2021 Datasets and Benchmarks (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Data quality for analytics and machine learning - Part 3: Data quality management requirements and guidelines" (2024). https://www.iso.org/standard/81092.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.