Speech and audio data
Speaker Diarization Training Data: RTTM Labels, Overlap and Annotation Specs
Quick answer
A usable speaker diarization dataset is real multi-party audio paired with time-stamped speaker-turn labels, usually in RTTM, where every turn carries a start time, duration and a consistent speaker ID, and overlapping speech is labeled rather than erased. Buyers should specify the annotation rules (minimum pause, overlap, backchannels, non-speech), the channel setup that produced the reference, and how speaker IDs persist across files. Without those rules, two suppliers' "diarized" audio will disagree on the same conversation.
By SourceX Editorial · Updated
What RTTM labels must contain for diarization training
RTTM is a plain-text format with one speaker turn per line, and challenge organizers require start time, duration and speaker ID in fixed column positions [1]. The common line type is SPEAKER, followed by the recording ID, a channel number, onset in seconds, duration in seconds, two placeholder fields (<NA>), the speaker label, and two more placeholders. Because each line is an independent interval, overlap is expressed simply as two lines whose time spans intersect.
That simplicity hides several decisions a supplier makes silently. Recording IDs must match audio filenames exactly, times should be in seconds with at least millisecond precision, and the reference should cover a defined scoring region. Scoring tools such as dscore read RTTM references alongside system output and optional UEM files that declare which parts of each recording are scored [4]. Ask for UEM files whenever the start or end of a recording contains hold music, IVR prompts or silence that was never annotated.
Illustrative example: invented to show structure; it does not describe an available dataset.
SPEAKER call_0412 1 0.000 4.210 <NA> <NA> agent_017 <NA> <NA>
SPEAKER call_0412 1 4.050 6.880 <NA> <NA> cust_0412_a <NA> <NA>
SPEAKER call_0412 1 9.700 0.410 <NA> <NA> agent_017 <NA> <NA>
SPEAKER call_0412 1 10.930 3.260 <NA> <NA> agent_017 <NA> <NA>
In this sketch, the customer's turn starts 0.16 seconds before the agent finishes (an overlap), and the 0.41-second agent turn inside the customer's speech is a backchannel such as "mm-hm" that the spec chose to keep. Note the agent label agent_017: it is a pseudonymous ID that should stay the same in every file where that agent appears, while customer IDs are scoped to one call.
How channel setup determines label accuracy
Labels are most accurate when each speaker was recorded on a separate channel or close-talk microphone, because the annotator (or a forced aligner) can see exactly when each person is speaking. In the DISPLACE 2023 challenge, organizers built reference RTTMs by combining close-talk annotations into one file per conversation, and kept separate files for speaker and language diarization [2]. The same logic applies to commercial audio: dual-channel call recordings give near-free speaker separation, while a single far-field meeting microphone forces annotators to separate voices by ear.
The practical consequence is that you may want two things from one source. The per-channel or close-talk audio produces the reference labels, and a downmixed mono or far-field track is what you train and evaluate on, since that matches production. Ask suppliers which audio the labels were derived from, and whether the training mixture was created after labeling. Our guide to dual-channel call recordings for ASR and diarization covers stereo layouts, and audio file specs for speech datasets covers sample rate, codec and channel mapping.
Annotation rules to specify before a supplier labels anything
Every diarization label set encodes a guideline, and if you do not write it, the supplier's default becomes your ground truth. The rules that change DER the most are the ones that decide where a turn begins and ends. A written spec also lets you audit whether the labels follow it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Rule | What to decide | Example setting | Why it matters |
|---|---|---|---|
| Minimum pause to split | Gap length at which one speaker's speech becomes two turns | Split at pauses of 0.3 s or longer | Too low fragments turns; too high merges across short replies |
| Overlap | Whether simultaneous speech gets two intersecting segments | Always label both speakers | Erasing overlap trains models to ignore crosstalk |
| Backchannels | Whether "uh-huh", "right", "okay" count as turns | Label as turns, flag in a side file | Dropping them shrinks the minority speaker's speech |
| Laughter, coughs, breaths | Speech or non-speech | Non-speech unless overlapping words | Affects speech activity detection boundaries |
| Non-speech regions | Hold music, IVR, TTS prompts, ringing | Exclude via UEM; label IVR as system if kept | Synthetic voices are not human speakers |
| Unknown or extra speakers | Background talkers, transfers, conference joins | Unique ID per new voice, never reuse | Merged IDs corrupt speaker counts |
| Boundary precision | Annotator target accuracy | Within 50 ms of audible onset | Defines how much collar evaluation needs |
| Speaker ID scope | Per file or global | Global for staff, per file for callers | Enables cross-file speaker clustering |
The same table doubles as a request template: paste it into an RFP with your own settings, and ask suppliers to state their existing rule where it differs. Our page on full-duplex conversation data, overlaps and backchannels goes further on turn-taking labels for speech-to-speech models.
Why conversational diarization data is hard to buy off the shelf
Carefully annotated conversational diarization data is scarce: one research team reported that no such test set was available and chose to release a 20-hour set of its own [3]. Public corpora skew toward meetings, broadcast and read or role-played speech, with license terms that often bar commercial use. Commercial calls and meetings are where the target distribution actually lives, which is why buyers tend to look at operational recordings from businesses.
Operational audio brings its own failure modes. Contact-center recordings often contain transfers (a third speaker mid-call), conference bridges, IVR segments and agent-side whisper coaching, and meeting recordings contain people joining late, sharing one room microphone, or speaking through laptop echo. A dataset that excludes all of these will look clean and then fail in production. Ask for counts of recordings by number of speakers, by overlap ratio, and by presence of transfers or system prompts.
See spontaneous conversational speech data for why scripted speech underrepresents interruptions.
Metadata and file layout to request with the labels
Diarization labels are only reusable if each RTTM ties cleanly to its audio and to per-speaker metadata. At minimum, request one RTTM per recording (or one combined RTTM with consistent recording IDs), matching UEM files, a speaker table, and a manifest with duration and channel layout per file.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example | Notes |
|---|---|---|
recording_id | call_0412 | Must match RTTM column 2 and the audio filename |
channels | 2 (ch1 agent, ch2 customer) | Record whether labels came from separated channels |
sample_rate_hz | 8000 | Telephony vs wideband changes model choice |
num_speakers | 3 | Includes transferred-to agent |
overlap_ratio | 0.07 | Share of scored speech with two or more speakers |
speaker_id | agent_017 | Pseudonymous, stable across files |
speaker_role | agent / customer / system | Useful for role-aware diarization |
annotation_pass | double-annotated, adjudicated | Supports quality claims |
redaction_spans | 41.20-44.85 | Where audio was masked for personal data |
Redaction deserves its own line: if a supplier silenced or toned over spoken account numbers, those spans can look like non-speech gaps that split a turn. Ask for redaction spans as a separate file so you can exclude them from scoring. Our pages on audio redaction artifacts in speech model training and speech dataset metadata fields cover both topics in depth.
How to check label quality before acceptance
Treat any supplier's DER or agreement figure as a claim to verify, because vendor targets are self-reported; one 2026 annotation guide cites production targets below 10% DER without an independent benchmark [5]. Label errors also persist in well-known benchmarks, where an audit of widely used vision, text and audio test sets estimated an average error rate of at least 3.3% [6]. Your acceptance test should run on a sample of the delivered audio, with your own rules.
A practical acceptance check runs in three steps. First, have an internal annotator re-label a random sample of recordings blind, using your written spec. Second, score the supplier's RTTM against yours with a tool such as dscore, with zero collar and overlap included, and inspect the per-file breakdown [4]. Third, listen to the worst files to separate guideline disagreements (fixable by rule) from careless labeling (a quality failure).
The annotation quality audit guide gives a fuller sampling method, and evaluating diarization on your own audio covers collars and test-set design.
Rights and privacy issues specific to speaker-labeled audio
Speaker-labeled audio is more sensitive than ordinary transcripts, because consistent speaker IDs across files make it easier to link voices to people. Illinois' BIPA lists a voiceprint as a biometric identifier [7], so ask whether speaker embeddings or voiceprints were ever generated from the audio, and keep that question with counsel. Confirm that call recordings carried the consent notices required where they were made, and that the supplier has the right to license recordings of both parties.
Pseudonymous speaker IDs help but do not remove the voice itself. If speaker identity matters only within a recording, request per-file IDs for external parties and keep global IDs only for consenting staff such as agents. Spoken names, emails and account numbers inside the audio still need masking, which is where the redaction-span file above becomes essential.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Where SourceX fits for diarization data
SourceX sources operational datasets from US companies, including support and sales histories and new recordings of hands-on work, and manages the licensing process. Data is sourced on request rather than held in stock, so you describe the audio and label spec you need and SourceX looks for US businesses that hold it; a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows.
You can describe your speaker-labeled audio requirements on the buyers page, or start from call center audio datasets and meeting transcripts for AI training. For background on the term itself, see the diarization glossary entry and the speech and audio data buyer's guide.
Request speaker-labeled conversational audio
Describe the recordings, channel setup and RTTM annotation rules you need, and SourceX will look for US companies that hold matching data. Nothing is contracted until a supplier agrees, and any dataset is delivered under a license that defines records, uses, term and delivery. Start a diarization data request with SourceX.
Sources
- arXiv, "The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition" (2023). https://arxiv.org/pdf/2303.06326
- arXiv, "Summary of the DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments" (2023). https://arxiv.org/pdf/2311.12564
- arXiv, "The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines" (2022). https://www.arxiv.org/pdf/2208.08042
- GitHub (srvk), "dscore: diarization scoring tools". https://github.com/srvk/dscore
- AIxBlock, "Speaker Diarization Training Data: A 2026 Annotation Guide" (2026). https://aixblock.io/blogs/speaker-diarization-training-data
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.