Speech and audio data
Overlapping Speech Data for Multi-Talker ASR and Meeting Transcription
Quick answer
An overlapping speech dataset is only useful if it can tell you who said which words while two or more people talked at once. That requires per-speaker references, ideally derived from close-talk or per-participant channels, word or segment timing, explicit overlap regions, and microphone metadata. Score those regions instead of masking them, and use speaker-attributed metrics such as cpWER and tcpWER rather than plain WER. Open meeting corpora are mostly small or dated, so licensed real meeting audio is often the path to scale.
By SourceX Editorial · Updated
This page covers the acoustic and labeling side of the problem. If you are licensing transcripts as text, see licensing meeting transcripts for AI training; for the broader category, start at the speech and audio data buyer's guide.
Why overlap breaks ordinary ASR training data
Overlap breaks ordinary data because a single-stream transcript cannot represent two simultaneous word sequences. Many single-channel recordings, such as room or speakerphone captures, were transcribed as one serialized text, so when two speakers collide, the transcriber either drops the quieter talker, merges words, or tags the span as [crosstalk]. A model trained on that output learns to suppress the secondary speaker, which is exactly the failure meeting transcription products are judged on.
The failure modes show up in three places. Backchannels ("yeah", "right", "mm-hm") vanish from references, so the model never learns to emit them. Interruptions get attributed to whoever held the floor. And diarization systems trained on single-label segments cannot assign two speakers to one frame, which is why overlap-aware formats matter (see diarization).
What a usable overlap reference contains
A usable reference gives each speaker their own time-stamped word stream, so that overlapped regions can be reconstructed rather than guessed. The cleanest way to get one is to record a close-talk or per-participant channel alongside the far-field mix. The DiPCo dinner-party corpus, for example, pairs close-talk microphones with several far-field arrays so transcribers can label each person from their own channel [2]. Challenge organizers have used close-talk recordings in the same way to build conversational diarization references [4].
Per-channel sources in business settings include multi-track exports from conferencing platforms (one track per participant), headset recordings in hybrid rooms, and dual-channel telephony, which separates the two sides of a call by design (see dual-channel call recordings for ASR and diarization). A single mixed MP3 from a room speakerphone is the weakest input, because overlap labels then depend entirely on human ear and cannot be verified.
Your reference package should include:
- Per-speaker transcripts with segment or word timestamps, not one interleaved text.
- RTTM (or equivalent) diarization files in which segments from different speakers may overlap in time. RTTM is a plain-text, one-turn-per-line format with fixed fields for start, duration and speaker ID [1].
- An explicit overlap flag or derived overlap table, so you can compute overlap ratio per session.
- A stated transcription convention for fillers, backchannels and partial words, aligned with your verbatim vs clean transcription standard.
Microphone setup and session metadata to require
Session metadata matters because the same conversation scores very differently from a headset, a laptop mic and a ceiling array. The AMI Meeting Corpus became a reference point partly because it recorded about 100 hours of meetings with synchronized close-talking and far-field microphones, room cameras, slide projector output and whiteboard capture on one timeline [5]. Commercial data rarely matches that, so ask what was captured and record it per session.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
session_id | mtg_000418 | Joins audio, RTTM, transcript and consent record |
platform | Conferencing app, cloud recording, per-participant tracks | Determines whether per-speaker channels exist |
capture_setup | 2 remote headsets + 1 room with 4-mic array | Near-field vs far-field split for eval slicing |
sample_rate_hz / codec | 16000 / Opus decoded to WAV | Codec artifacts differ from studio audio |
channels | 3 per-speaker tracks + 1 mixed | Reference quality and separation targets |
participant_count | 5 | Overlap density rises with talker count |
overlap_ratio | 0.14 (overlapped speech / total speech) | Lets you balance training mix and eval slices |
domain | Engineering standup | Vocabulary and turn-taking style |
consent_basis | All participants notified and consented | Required before any use |
redaction | Names replaced in transcript; audio bleeped at PII spans | Affects word timing and model behavior |
Ask suppliers to state whether timestamps are word-level or segment-level, how per-track drift was corrected, and whether redaction tones or silences were inserted. Redaction choices change what the model hears, as covered in how audio redaction affects speech model training. For file-level conventions, see packaging speech datasets: manifests, segment timing and diarization files.
Scoring overlap: cpWER, tcpWER and DER without a collar
Score overlapped regions directly, because the usual shortcuts hide exactly the errors multi-talker systems make. Plain WER on a concatenated transcript ignores speaker attribution, so a system that transcribes every word but assigns half of them to the wrong person can look strong. Diarization scoring with a forgiveness collar around boundaries also inflates results in dense overlap; the MISP 2022 challenge scored overlapping speech without applying a collar for that reason [1].
Speaker-attributed metrics close the gap:
- cpWER (concatenated minimum-permutation WER) concatenates each speaker's words, finds the best mapping between hypothesis and reference speakers, and computes WER under that mapping. It punishes speaker confusion but ignores timing.
- tcpWER (time-constrained cpWER) adds a temporal constraint, so a word only matches if it falls within a tolerance window of its reference time. This stops a system from earning credit for words placed in the wrong part of the meeting. Recent multi-speaker benchmarks report tcpWER for ASR alongside DER for diarization [3].
- DER with overlap included, and no collar or a small, documented one, for the diarization component [1].
Two practical rules keep these numbers comparable. Normalize text before scoring (casing, numerals, contractions, filler handling), since formatting differences otherwise count as errors; Whisper's authors built a normalizer for exactly that purpose [6]. And report results sliced by overlap ratio and by capture setup, because an average across headset and far-field sessions tells you little about either.
Open corpora versus licensed meeting audio
Open meeting corpora are excellent for evaluation but rarely enough for training a production system. They tend to be small, recorded years ago in lab rooms, and scripted or semi-scripted. DiPCo, for instance, is about 5 hours and is recommended for evaluation rather than training [2]. Before you train on any of them commercially, run a license audit of open speech corpora, since research-only terms are common.
Licensed business meeting audio covers what lab corpora miss: real agendas, domain vocabulary, remote participants on poor connections, and natural interruptions. The trade-offs are uneven channel quality, heavier privacy review, and the need to confirm that every participant consented, not only the meeting host. Some buyers combine a small open corpus as a fixed eval set with licensed audio for training, which also limits contamination between the two.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Requirement | Open lab corpus | Licensed business meeting audio | Synthetic overlap (mixed single-speaker clips) |
|---|---|---|---|
| Per-speaker reference | Often, via close-talk mics | Depends on platform tracks | Exact, by construction |
| Natural turn-taking and backchannels | Partly | Yes | No |
| Domain vocabulary | Limited | Strong | Inherits source clips |
| Scale | Small | Depends on supplier | Unlimited |
| Consent and rights review | Check license terms | All-participant review required | Depends on source clips |
| Best use | Fixed eval set | Training and domain eval | Pretraining and separation front ends |
Synthetic mixtures help separation models, but they lack the prosodic cues people use to grab or yield the floor. For more on that trade-off, see real vs synthetic speech data for ASR.
Consent and biometric review for multi-party recordings
Multi-party recordings need consent review for every participant, because one person's agreement does not cover the others. As of October 2026, California Penal Code section 632 prohibits recording a confidential communication without the consent of all parties [7], and meetings frequently include participants in different states. Texas lists a voiceprint as a biometric identifier and requires notice and consent before capturing one for a commercial purpose [8], which matters if your use includes speaker embeddings or voice identification.
For counsel and privacy reviewers, the useful questions are concrete. Was a recording notice shown or announced to all participants, and is that preserved per session? Did external guests (customers, candidates, vendors) join, and on what basis are they covered? Are speaker embeddings in scope or excluded by license? Health or HR meetings raise further issues covered in special category data in operational training datasets.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Request template for a multi-talker audio sourcing brief
A good brief specifies the overlap you need to see and the references you need to score it, not just hours. Use this structure when you describe the data to a supplier or broker.
Illustrative example: invented to show structure; it does not describe an available dataset.
Use: multi-talker ASR training + held-out eval (meeting transcription)
Sessions: internal meetings, 3-8 participants, US English, mixed remote/in-room
Audio: per-participant tracks where available + mixed track; 16 kHz or higher WAV/FLAC
Overlap: report overlap_ratio per session; target slice with ratio >= 0.10
References: per-speaker word timestamps; RTTM with overlapping segments allowed
Conventions: verbatim, backchannels kept, partial words tagged
Metadata: platform, capture_setup, participant_count, domain, redaction method
Privacy: names/emails/numbers removed or replaced; method documented; sample checked
Consent: all-participant notice/consent evidence per session
Eval split: speaker-disjoint and organization-disjoint from training
Metrics we will run: tcpWER, cpWER, DER (no collar)
How much to request depends on your baseline; how many hours of audio you need to fine-tune ASR works through that question. If you also need video and shared screens, see multimodal meeting recordings.
How SourceX handles meeting and call audio requests
SourceX sources operational datasets, including support and sales histories and new recordings of hands-on work, from US companies on request, and manages licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details such as names, emails and phone numbers are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. For context on demand, see do AI labs buy meeting recordings, or describe the audio you need.
Source multi-talker meeting audio with SourceX
If you need meeting or multi-party call audio with per-speaker references for multi-talker ASR, describe the data, not the businesses, and SourceX looks for US companies that hold it. Each dataset goes through rights review and is delivered under a license that defines records, uses, term and delivery. Start a buyer request at sourcex.si/buyers.
Sources
- arXiv, "The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition" (2023). https://arxiv.org/pdf/2303.06326
- VoxKitchen, "DiPCo dataset documentation". https://voxkitchen.readthedocs.io/en/latest/datasets/dipco/
- arXiv, "Benchmarking Speech Systems for Frontline Health Conversations: The DISPLACE-M Challenge" (2026). https://arxiv.org/pdf/2603.02813
- arXiv, "Summary of the DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments" (2023). https://arxiv.org/pdf/2311.12564
- TensorFlow Datasets, "AMI dataset (community catalog)". https://tensorflow.org/datasets/community_catalog/huggingface/ami
- arXiv (OpenAI), "Robust Speech Recognition via Large-Scale Weak Supervision" (2022). https://arxiv.org/pdf/2212.04356
- California Legislative Information, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.