Skip to content

Speech and audio data

Far-Field Speech Data: Real Distant-Microphone Recordings vs Simulated Rooms

Quick answer

Most far-field ASR teams should train mainly on simulated far-field audio and reserve real distant-microphone recordings for evaluation and targeted fine-tuning. Simulation (clean speech convolved with room impulse responses, plus noise at controlled SNR, rendered through your device's microphone geometry) gives scale and control [1]. Real recordings capture what simulation misses: device front-end processing, moving talkers, playback echo and the acoustics of rooms you never modeled. Source real audio for the gap your simulation cannot close, and specify the geometry metadata that makes it usable.

By SourceX Editorial · Updated

What makes far-field speech a different data problem

Far-field speech differs from close-talk speech because the room, not just the talker, shapes the signal that reaches the microphone. At distance, the direct-path energy drops relative to reflections, so the direct-to-reverberant ratio (DRR) falls and late reverberation smears phonemes into each other. Reverberation time (RT60), source-to-microphone distance, talker orientation and the microphone's position against walls or tables all change the impulse response. A model trained on headset or phone audio sees none of this, which is why distant speech recognition is treated as its own data category rather than a noise setting.

This page owns distance and reverberation as data properties. Additive noise types and SNR bands are covered in noisy speech data and SNR coverage, and the broader TTS-versus-real debate is in real vs synthetic speech data for ASR. For the full cluster, start at the speech and audio data buyer's guide.

Multichannel capture adds a second dimension. Microphone-array front ends (delay-and-sum, MVDR or neural beamformers) depend on inter-channel phase and time differences, so training data must preserve channel count, channel order and array geometry exactly. Downmixing to mono or resampling channels independently destroys the spatial cues a beamforming model learns from.

How the simulation recipe works and where it is strong

Simulated far-field data is built by convolving clean, close-talk speech with room impulse responses (RIRs) and adding noise at a chosen SNR, rendered for the exact microphone layout of the target device. A published deep-beamforming system took about two million utterances, placed them in simulated reverberant rooms, mixed noise at SNRs from 0 to 25 dB, and simulated a four-microphone array matching the evaluation device's geometry [1]. That is the canonical recipe: scale comes from the clean corpus, diversity from the RIR and noise sampling.

Simulation is strong where you need breadth and labels for free. The transcript of the clean source is the transcript of the reverberant copy, so no new annotation is needed. You can also sweep RT60, distance and SNR systematically, and you get a clean target for speech enhancement and dereverberation training.

The RIRs can be synthetic (image-source or ray-tracing models, or wave-based solvers) or measured in real rooms. Measured RIR collections such as BUT ReverbDB provide real impulse responses, room noises and retransmitted speech with recorded microphone and loudspeaker positions, released under CC BY 4.0 [4]. Measured RIRs close part of the realism gap at low cost, and their license terms are usually simpler than speech corpora, but check each one using the steps in the open speech corpora commercial license audit.

Where simulated rooms fall short

Simulation fails when the real signal chain contains effects the convolution model does not represent. Researchers releasing new far-field datasets argue that ASR, dereverberation and enhancement models must be trained on data that accurately represents complex far-field acoustic behavior, and simulation fidelity remains an active research area [2]. Common gaps buyers see in practice:

  • Device front-end processing. Production devices run acoustic echo cancellation (AEC), automatic gain control, noise suppression and sometimes vendor beamforming before audio reaches the ASR model. Simulated audio skips that chain unless you replay it through the real firmware.
  • Low-frequency and diffraction effects. Simple image-source models handle specular reflections well but approximate low-frequency modes and diffraction around furniture and people poorly.
  • Moving and turning talkers. A static RIR assumes a fixed source; real talkers walk, turn away and lean in mid-utterance.
  • Lombard and distance-aware speech. People speak louder and differently when addressing a device across a room. Re-reverberating read speech does not change the speaking style.
  • Playback and self-noise. Smart speakers and kiosks hear their own audio output, fans and motors through the chassis, which is coupled differently from ambient noise.
  • Distributed capture. Meeting systems with several devices suffer clock drift and unsynchronized channels, which matched-array simulation does not reproduce.

Each gap maps to a failure mode in production: word error rate that looks fine on simulated test sets and degrades sharply on real far-field audio. That is the strongest argument for keeping a real evaluation set.

What real far-field corpora look like

Real far-field corpora are usually small, captured in a few rooms, and paired with close-talk reference channels. DiPCo, for example, provides close-talk microphones plus multiple 7-microphone arrays, but at roughly 5 hours it is suggested for evaluation rather than primary training [3]. The AMI Meeting Corpus is larger at about 100 hours and synchronizes close-talking and far-field microphones with room-view cameras, slides and whiteboard capture on a common timeline [5].

Those design choices are worth copying when you commission or license real data. A close-talk channel recorded in parallel gives you a cleaner transcription reference and a target signal for enhancement. Multiple arrays at different positions in the same room give you distance and angle diversity from a single session.

The limitation is coverage. Public far-field corpora capture particular rooms, languages, accents and array geometries, and their licenses often restrict commercial use. If your product is a kiosk in retail floors or a conference-room bar in open-plan offices, you will usually need recordings from environments that match it, which means licensing operational audio or commissioning new capture. Meeting-room audio overlaps with licensed meeting transcripts, and broader options are on the voice and audio data page.

Deciding the real-to-simulated split

The right split depends on how far your deployment acoustics are from anything you can simulate. Use simulation for bulk training and real far-field audio for evaluation, then add real data to training only where error analysis shows a gap simulation cannot close [3]. The table below is a starting framework, not a benchmark.

Illustrative example: invented to show structure; it does not describe an available dataset.

Deployment situationSimulated data roleReal far-field data roleWhat to source
Fixed device, known array, raw-mic accessPrimary training, matched to geometryHeld-out evaluation from target roomsMeasured RIRs from similar rooms, real eval sessions
Device with closed vendor front end (AEC, AGC)Pre-training onlyFine-tuning and evaluation through the real chainRecordings captured on the shipping device
Meeting rooms with distributed devicesAugmentation for single-channel modelsTraining and evaluation for diarization and driftMulti-device sessions with close-talk references
Kiosks in public spacesBreadth across noise and RT60Evaluation per venue typeVenue-specific recordings and ambient noise beds
Wake wordPositive-sample augmentation at distanceFalse-accept testing on long real ambient audioHours of real household or venue background

A practical rule: never report far-field accuracy from a test set built with the same RIR family used in training. Hold out rooms, not just utterances, and keep at least one real-recorded test condition per target environment.

Metadata to require for real far-field recordings

Real far-field audio is only as useful as the geometry and room metadata that ships with it. Without distance, array layout and channel mapping, you cannot stratify evaluation, reproduce conditions in simulation or debug a beamformer. Treat these fields as acceptance criteria; the general field list is in speech dataset metadata fields.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "session_id": "room07_s012",
  "room": {"type": "conference_room", "dimensions_m": [6.1, 4.3, 2.7], "rt60_s_measured": 0.62, "surfaces": "glass wall, carpet"},
  "array": {"device_model": "ceiling_bar_v2", "num_channels": 8, "channel_order": [0,1,2,3,4,5,6,7],
            "mic_coordinates_m": "array_geometry.csv", "front_end_processing": "raw, pre-AEC"},
  "talkers": [{"speaker_id": "spk_031", "distance_m": 2.4, "azimuth_deg": 35, "moving": false,
               "close_talk_channel": "spk_031_lav.flac"}],
  "playback_active": true,
  "noise_sources": ["HVAC", "keyboard"],
  "sample_rate_hz": 16000, "bit_depth": 16, "codec": "FLAC",
  "clock_sync": "single ADC", "transcript_ref": "room07_s012.ctm"
}

Four failure modes recur when these fields are missing or wrong:

  1. Swapped or missing channel order, which silently breaks any spatial model.
  2. Processed audio labeled as raw, so AEC or noise suppression is applied twice in your pipeline and the training audio no longer matches production.
  3. Lossy compression or per-channel resampling, which corrupts inter-channel phase; see audio file specs.
  4. No close-talk or reference transcript alignment, which forces expensive re-transcription of reverberant audio.

For packaging multichannel sessions with segment timing, follow the conventions in speech dataset manifest packaging.

What to source for each side

Simulation and real capture need different inputs, and both carry license questions. Treat the inputs as separate line items in your data plan.

For simulation, source:

  • A large clean close-talk corpus whose license permits derivative augmented copies for commercial training.
  • RIRs, measured or synthetic, covering your target RT60 and distance range; measured collections like BUT ReverbDB include positions and room noises [4].
  • Noise beds recorded in environments like your deployment, with rights to mix and redistribute internally.
  • Your device's exact microphone coordinates and, ideally, its front-end processing chain for replay [1].

For real far-field audio, source:

  • Recordings captured on the target device or a matched array, raw and post-front-end where possible.
  • Parallel close-talk channels for transcription and enhancement targets [3][5].
  • Room diversity documented per session, so you can hold out rooms.
  • Speaker consent that covers recording at distance, bystanders in shared spaces and model training; review the scope against speech and voice recording license terms.

Real operational recordings from businesses, such as meeting-room or service-counter audio, often contain names, account numbers and other personal details spoken aloud. Ask how those were redacted and how redaction artifacts appear in the audio before you accept the data.

Sourcing real far-field speech through SourceX

SourceX sources operational datasets from US companies on request, including recordings of hands-on work, and manages licensing and ongoing purchases. Data is not held in stock, so a request does not guarantee a match. You describe the far-field data you need (rooms, distances, devices, channels, metadata), and SourceX looks for US businesses that hold it; every release is approved by the supplying company.

Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement. Describe your requirement on the SourceX buyer page.

Request real far-field speech data

If simulation has carried your far-field model as far as it can, the next step is real distant-microphone audio from rooms that match your deployment. SourceX follows a Find, Assess, Agree, Transact and Manage process, and nothing is contracted until a supplier agrees. Describe your far-field speech data requirement to SourceX.

Sources

  1. arXiv, "Spatial Attention for Far-field Speech Recognition with Deep Beamforming Neural Networks" (2019). https://arxiv.org/pdf/1911.02115
  2. arXiv, "Treble10: A high-quality dataset for far-field speech recognition, dereverberation, and enhancement" (2025). https://arxiv.org/pdf/2510.23141
  3. Sell et al., "DiPCo -- Distant Speech Recognition with Joint Learning of Acoustic Echo Cancellation, Dereverberation and Beamforming" (2019). https://arxiv.org/abs/1909.13447
  4. Brno University of Technology, Faculty of Information Technology (IEEE published version), "Building and Evaluation of a Real Room Impulse Response Dataset (BUT ReverbDB)" (2019). https://www.fit.vut.cz/research/publication-file/c159973/280107/IEEE_Final_Published_08717722-1.pdf
  5. TensorFlow Datasets, "AMI Meeting Corpus (community catalog entry)". https://tensorflow.org/datasets/community_catalog/huggingface/ami

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data