Skip to content

Speech and audio data

Speaker Diarization Training Data: RTTM Labels, Overlap and Annotation Specs

Quick answer

A usable speaker diarization dataset is real multi-party audio paired with time-stamped speaker-turn labels, usually in RTTM, where every turn carries a start time, duration and a consistent speaker ID, and overlapping speech is labeled rather than erased. Buyers should specify the annotation rules (minimum pause, overlap, backchannels, non-speech), the channel setup that produced the reference, and how speaker IDs persist across files. Without those rules, two suppliers' "diarized" audio will disagree on the same conversation.

By SourceX Editorial · Updated

What RTTM labels must contain for diarization training

RTTM is a plain-text format with one speaker turn per line, and challenge organizers require start time, duration and speaker ID in fixed column positions [1]. The common line type is SPEAKER, followed by the recording ID, a channel number, onset in seconds, duration in seconds, two placeholder fields (<NA>), the speaker label, and two more placeholders. Because each line is an independent interval, overlap is expressed simply as two lines whose time spans intersect.

That simplicity hides several decisions a supplier makes silently. Recording IDs must match audio filenames exactly, times should be in seconds with at least millisecond precision, and the reference should cover a defined scoring region. Scoring tools such as dscore read RTTM references alongside system output and optional UEM files that declare which parts of each recording are scored [4]. Ask for UEM files whenever the start or end of a recording contains hold music, IVR prompts or silence that was never annotated.

Illustrative example: invented to show structure; it does not describe an available dataset.

SPEAKER call_0412 1  0.000  4.210 <NA> <NA> agent_017 <NA> <NA>
SPEAKER call_0412 1  4.050  6.880 <NA> <NA> cust_0412_a <NA> <NA>
SPEAKER call_0412 1  9.700  0.410 <NA> <NA> agent_017 <NA> <NA>
SPEAKER call_0412 1 10.930  3.260 <NA> <NA> agent_017 <NA> <NA>

In this sketch, the customer's turn starts 0.16 seconds before the agent finishes (an overlap), and the 0.41-second agent turn inside the customer's speech is a backchannel such as "mm-hm" that the spec chose to keep. Note the agent label agent_017: it is a pseudonymous ID that should stay the same in every file where that agent appears, while customer IDs are scoped to one call.

How channel setup determines label accuracy

Labels are most accurate when each speaker was recorded on a separate channel or close-talk microphone, because the annotator (or a forced aligner) can see exactly when each person is speaking. In the DISPLACE 2023 challenge, organizers built reference RTTMs by combining close-talk annotations into one file per conversation, and kept separate files for speaker and language diarization [2]. The same logic applies to commercial audio: dual-channel call recordings give near-free speaker separation, while a single far-field meeting microphone forces annotators to separate voices by ear.

The practical consequence is that you may want two things from one source. The per-channel or close-talk audio produces the reference labels, and a downmixed mono or far-field track is what you train and evaluate on, since that matches production. Ask suppliers which audio the labels were derived from, and whether the training mixture was created after labeling. Our guide to dual-channel call recordings for ASR and diarization covers stereo layouts, and audio file specs for speech datasets covers sample rate, codec and channel mapping.

Annotation rules to specify before a supplier labels anything

Every diarization label set encodes a guideline, and if you do not write it, the supplier's default becomes your ground truth. The rules that change DER the most are the ones that decide where a turn begins and ends. A written spec also lets you audit whether the labels follow it.

Illustrative example: invented to show structure; it does not describe an available dataset.

RuleWhat to decideExample settingWhy it matters
Minimum pause to splitGap length at which one speaker's speech becomes two turnsSplit at pauses of 0.3 s or longerToo low fragments turns; too high merges across short replies
OverlapWhether simultaneous speech gets two intersecting segmentsAlways label both speakersErasing overlap trains models to ignore crosstalk
BackchannelsWhether "uh-huh", "right", "okay" count as turnsLabel as turns, flag in a side fileDropping them shrinks the minority speaker's speech
Laughter, coughs, breathsSpeech or non-speechNon-speech unless overlapping wordsAffects speech activity detection boundaries
Non-speech regionsHold music, IVR, TTS prompts, ringingExclude via UEM; label IVR as system if keptSynthetic voices are not human speakers
Unknown or extra speakersBackground talkers, transfers, conference joinsUnique ID per new voice, never reuseMerged IDs corrupt speaker counts
Boundary precisionAnnotator target accuracyWithin 50 ms of audible onsetDefines how much collar evaluation needs
Speaker ID scopePer file or globalGlobal for staff, per file for callersEnables cross-file speaker clustering

The same table doubles as a request template: paste it into an RFP with your own settings, and ask suppliers to state their existing rule where it differs. Our page on full-duplex conversation data, overlaps and backchannels goes further on turn-taking labels for speech-to-speech models.

Why conversational diarization data is hard to buy off the shelf

Carefully annotated conversational diarization data is scarce: one research team reported that no such test set was available and chose to release a 20-hour set of its own [3]. Public corpora skew toward meetings, broadcast and read or role-played speech, with license terms that often bar commercial use. Commercial calls and meetings are where the target distribution actually lives, which is why buyers tend to look at operational recordings from businesses.

Operational audio brings its own failure modes. Contact-center recordings often contain transfers (a third speaker mid-call), conference bridges, IVR segments and agent-side whisper coaching, and meeting recordings contain people joining late, sharing one room microphone, or speaking through laptop echo. A dataset that excludes all of these will look clean and then fail in production. Ask for counts of recordings by number of speakers, by overlap ratio, and by presence of transfers or system prompts.

See spontaneous conversational speech data for why scripted speech underrepresents interruptions.

Metadata and file layout to request with the labels

Diarization labels are only reusable if each RTTM ties cleanly to its audio and to per-speaker metadata. At minimum, request one RTTM per recording (or one combined RTTM with consistent recording IDs), matching UEM files, a speaker table, and a manifest with duration and channel layout per file.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExampleNotes
recording_idcall_0412Must match RTTM column 2 and the audio filename
channels2 (ch1 agent, ch2 customer)Record whether labels came from separated channels
sample_rate_hz8000Telephony vs wideband changes model choice
num_speakers3Includes transferred-to agent
overlap_ratio0.07Share of scored speech with two or more speakers
speaker_idagent_017Pseudonymous, stable across files
speaker_roleagent / customer / systemUseful for role-aware diarization
annotation_passdouble-annotated, adjudicatedSupports quality claims
redaction_spans41.20-44.85Where audio was masked for personal data

Redaction deserves its own line: if a supplier silenced or toned over spoken account numbers, those spans can look like non-speech gaps that split a turn. Ask for redaction spans as a separate file so you can exclude them from scoring. Our pages on audio redaction artifacts in speech model training and speech dataset metadata fields cover both topics in depth.

How to check label quality before acceptance

Treat any supplier's DER or agreement figure as a claim to verify, because vendor targets are self-reported; one 2026 annotation guide cites production targets below 10% DER without an independent benchmark [5]. Label errors also persist in well-known benchmarks, where an audit of widely used vision, text and audio test sets estimated an average error rate of at least 3.3% [6]. Your acceptance test should run on a sample of the delivered audio, with your own rules.

A practical acceptance check runs in three steps. First, have an internal annotator re-label a random sample of recordings blind, using your written spec. Second, score the supplier's RTTM against yours with a tool such as dscore, with zero collar and overlap included, and inspect the per-file breakdown [4]. Third, listen to the worst files to separate guideline disagreements (fixable by rule) from careless labeling (a quality failure).

The annotation quality audit guide gives a fuller sampling method, and evaluating diarization on your own audio covers collars and test-set design.

Rights and privacy issues specific to speaker-labeled audio

Speaker-labeled audio is more sensitive than ordinary transcripts, because consistent speaker IDs across files make it easier to link voices to people. Illinois' BIPA lists a voiceprint as a biometric identifier [7], so ask whether speaker embeddings or voiceprints were ever generated from the audio, and keep that question with counsel. Confirm that call recordings carried the consent notices required where they were made, and that the supplier has the right to license recordings of both parties.

Pseudonymous speaker IDs help but do not remove the voice itself. If speaker identity matters only within a recording, request per-file IDs for external parties and keep global IDs only for consenting staff such as agents. Spoken names, emails and account numbers inside the audio still need masking, which is where the redaction-span file above becomes essential.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Where SourceX fits for diarization data

SourceX sources operational datasets from US companies, including support and sales histories and new recordings of hands-on work, and manages the licensing process. Data is sourced on request rather than held in stock, so you describe the audio and label spec you need and SourceX looks for US businesses that hold it; a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows.

You can describe your speaker-labeled audio requirements on the buyers page, or start from call center audio datasets and meeting transcripts for AI training. For background on the term itself, see the diarization glossary entry and the speech and audio data buyer's guide.

Request speaker-labeled conversational audio

Describe the recordings, channel setup and RTTM annotation rules you need, and SourceX will look for US companies that hold matching data. Nothing is contracted until a supplier agrees, and any dataset is delivered under a license that defines records, uses, term and delivery. Start a diarization data request with SourceX.

Sources

  1. arXiv, "The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition" (2023). https://arxiv.org/pdf/2303.06326
  2. arXiv, "Summary of the DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments" (2023). https://arxiv.org/pdf/2311.12564
  3. arXiv, "The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines" (2022). https://www.arxiv.org/pdf/2208.08042
  4. GitHub (srvk), "dscore: diarization scoring tools". https://github.com/srvk/dscore
  5. AIxBlock, "Speaker Diarization Training Data: A 2026 Annotation Guide" (2026). https://aixblock.io/blogs/speaker-diarization-training-data
  6. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data