Speech and audio data
Evaluating Diarization on Your Own Audio: DER, Collars and Test-Set Design
Quick answer
Diarization error rate (DER) is the sum of false-alarm speech, missed speech and speaker-confusion time, divided by total reference speech time. The number is only comparable when every system is scored with the same scorer, the same collar, the same overlap rule and the same scoring regions (UEM). To evaluate models or vendors for your product, freeze that protocol, score with a standard tool such as dscore on RTTM files, add turn-level checks that DER hides, and build the test set from audio that matches your production channels.
By SourceX Editorial · Updated
How to calculate DER, term by term
DER is a time-weighted error: DER = (false alarm + missed speech + speaker confusion) / total reference speech duration [3]. Each term is measured in seconds after the scorer finds the optimal one-to-one mapping between reference speaker labels and system speaker labels, because system labels like spk_0 are arbitrary [1].
- False alarm (FA): time the system labels as speech where the reference has none. Usually a voice activity detection (VAD) problem: hold music, IVR prompts, keyboard noise, a TV in the room.
- Missed speech (MISS): reference speech the system did not label. Includes overlapped speech when the system emits only one speaker for a region where the reference has two.
- Speaker confusion (CONF): speech assigned to the wrong mapped speaker. This is the clustering error most vendor comparisons actually care about.
Because the denominator is reference speech, not audio duration, DER can exceed 100% on a file where a system hallucinates a great deal of speech. Report the three components alongside the total; two systems with the same 12% DER can have very different failure profiles.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Quantity (one 15-minute two-party call) | Seconds |
|---|---|
| Total reference speech (overlap counted per speaker) | 600 |
| False alarm | 12 |
| Missed speech | 30 |
| Speaker confusion | 24 |
| DER = (12 + 30 + 24) / 600 | 11.0% |
Corpus-level DER sums each term across all files before dividing. Averaging per-file DERs weights a 40-second call the same as a 40-minute meeting, and the two numbers can rank systems differently, so state which one you report.
Why the 250 ms collar and overlap rule change the score
The collar is a forgiveness window around every reference segment boundary, and it can move DER by several points on conversational audio. The common 250 ms collar, inherited from NIST evaluation practice, excludes 250 ms on each side of each reference boundary from scoring, so small boundary misalignments are not counted [2]. On fast back-and-forth speech with many short turns, a large share of the audio sits inside collars, which flatters every system and compresses the gap between them.
Overlap handling is the second switch. Many older published results, for example on CALLHOME telephone speech, excluded overlapped regions from scoring; more recent challenges score everything. The MISP 2022 challenge, for example, used no collar so that overlapping speech was fully scored [3]. A system without overlap detection looks acceptable under the first protocol and much worse under the second, especially on meetings and family or multi-party calls [2].
The rule for comparisons: never put a vendor's published DER next to your measured DER unless both state collar, overlap treatment and UEM. Re-score every system yourself, from their raw RTTM output, under one protocol.
Scoring with dscore and RTTM files
Use a standard, open scorer rather than a vendor's internal tool. dscore takes a reference RTTM and a system RTTM, optionally a UEM file that defines which time ranges are scored, and reports DER and its components per file and overall [1]. It also reports Jaccard error rate (JER), discussed below.
RTTM is a plain-text format with one speaker turn per line, and scorers expect the file ID, channel, start time, duration and speaker label in fixed column positions [3]. A typical line looks like this:
SPEAKER call_0412 1 83.270 2.415 <NA> <NA> agent <NA> <NA>
Common scoring failures that silently corrupt results:
- File-ID mismatch. The second RTTM column must match exactly across reference, system and UEM files; a mismatch drops the file or scores it as all-miss.
- Missing UEM. Without a UEM, the scorer infers regions from the RTTMs, so leading silence or IVR menus may be in or out of scope depending on the file.
- Sample-rate or offset drift. If one pipeline trims a header or resamples with a delay, every boundary shifts and confusion inflates. Check audio specs first; see audio file specs for speech datasets.
- Zero-duration or negative turns from post-processing bugs, which some scorers reject and others ignore.
Pin the scorer version and its command-line flags (collar, overlap option, UEM path) in the test-set README so a rerun in six months produces the same number.
JER vs DER: which diarization metric to report
Report DER as the primary number and JER as a speaker-balanced secondary metric. DER is dominated by whoever talks most: an agent who speaks 70% of a call drives most of the score, and a customer who says little can be badly mislabeled with only a small DER penalty. JER computes an error per reference speaker (one minus intersection over union with the mapped system speaker) and averages across speakers, so each speaker counts equally [1].
JER exposes systems that merge a quiet participant into the dominant one, a common failure in three-way calls with a supervisor or interpreter. It does not replace DER for capacity planning, because downstream systems such as transcript attribution and analytics care about time-weighted accuracy.
What DER hides: short turns, backchannels and speaker counts
DER under-weights short conversational phrases because every error is weighted by its duration [4]. A system that attributes every "yeah," "mm-hmm" and "right" to the wrong speaker loses almost nothing in DER, yet those turns matter for sentiment, turn-taking analytics and conversational AI training data built from transcripts.
Add three checks next to DER:
- Turn-level attribution accuracy for reference turns under 1 second, scored as the share of short turns whose majority-overlapping system speaker maps correctly.
- Speaker-count error: absolute difference between reference and estimated number of speakers per file, reported as a distribution, not a mean.
- Speaker-change boundary precision and recall within a tolerance you set, which isolates segmentation from clustering.
These numbers are not standardized across papers, so they are for internal ranking only. Document their definitions in the protocol so vendors cannot dispute them later.
Designing a diarization test set from your own audio
A useful test set mirrors your production distribution of channels, speaker counts, overlap and acoustic conditions, not a public benchmark's. Conversational benchmarks such as DISPLACE 2023 cover multilingual, multi-speaker settings with their own protocol [5], but they rarely match a contact center's 8 kHz G.711 stereo calls or a clinic's far-field room audio. Public results tell you which systems to shortlist; your own audio tells you which one to buy.
Where possible, source dual-channel recordings, where each party is on its own channel, so per-channel voice activity gives a strong starting reference for who spoke when. Human annotators then correct boundaries, mark overlap and crosstalk, and label third parties. For why stereo matters, see dual-channel call recordings for ASR and diarization; for labeling conventions, see RTTM labels, overlap and annotation specs. Then score systems on the mixed mono signal they will see in production, not the separated channels.
Reference labels are themselves noisy. Northcutt and colleagues estimated an average label error rate of at least 3.3% across test sets of ten widely used datasets, including audio, and showed those errors can change which model ranks first [6]. Double-annotate a slice, adjudicate disagreements, and measure annotator-vs-annotator DER as your noise floor: a vendor difference smaller than that floor is not a real difference. Our guide to gold-label audits for delivered eval sets covers acceptance in more detail.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Stratum | Why it matters | Example target share | Reference method |
|---|---|---|---|
| Two-party stereo calls, 8 kHz | Production baseline | 40% | Per-channel VAD + human boundary correction |
| Three or more speakers (transfers, supervisors, interpreters) | Speaker-count and merge errors | 15% | Full manual annotation, overlap marked |
| High overlap (>10% of speech) | Overlap detection | 15% | Manual, overlap scored, no collar |
| Short-turn-heavy (backchannels) | DER blind spot | 10% | Manual, turns under 1 s tagged |
| Far-field or noisy rooms | VAD false alarms | 10% | Manual |
| Hold music, IVR, silence-heavy | False alarm stress | 10% | UEM excludes IVR or includes it, stated explicitly |
Stratify deliberately and report DER per stratum, not only overall; see stratified evaluation sets for rare and high-risk cases. Keep the set held out from every vendor and from your own training data, and rotate a portion as conditions change.
A frozen protocol card for vendor comparisons
Write the protocol before you see any results, and send it to every vendor unchanged. The card below is a template for what to fix in advance.
Illustrative example: invented to show structure; it does not describe an available dataset.
diarization_eval_protocol:
scorer: dscore # pin commit hash
metrics: [DER, JER, miss, false_alarm, confusion]
collar_seconds: 0.0 # also report 0.25 for comparison with legacy numbers
overlap: scored
uem: uem/test_v3.uem # IVR prompts excluded
aggregation: corpus_level # sum terms across files before dividing
oracle_vad: false # report an oracle-VAD run separately if useful
speaker_count_hint: none # do not pass known speaker counts
input_audio: mono_mix_8k # what production sees
extra_checks: [short_turn_accuracy_lt_1s, speaker_count_error]
noise_floor: inter_annotator_DER_on_double_labeled_slice
per_stratum_reporting: true
Two settings deserve explicit attention. Oracle VAD largely removes false alarm and missed speech from the comparison (missed overlap can remain) and isolates clustering, which is useful for diagnosis but misleading as a headline. Passing the true number of speakers to a system can lower confusion substantially, so never mix "known speaker count" and "estimated" results in one table.
Where licensed conversational audio fits
Many teams lack enough in-domain, multi-speaker audio to build strata like those above, especially for transfers, interpreted calls or field recordings. SourceX sources operational datasets, including support and sales histories and new recordings of hands-on work, from US companies on request; it holds no inventory, and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details such as names and phone numbers are removed or replaced before delivery with the method recorded and a sample checked (no method is perfect), and use is defined in a license. You can describe the evaluation audio you need by channel layout, speaker count and conditions rather than naming suppliers. For the wider cluster, start at the speech and audio data buyer's guide, and see the diarization glossary entry and evaluation datasets built from real business work. If you also measure transcription, pair this with fair WER measurement and normalization.
Request diarization evaluation audio from SourceX
Describe the conversational audio your diarization test set needs: channels, speaker counts, overlap and conditions. SourceX looks for US businesses that hold that data, assesses data and licensing permissions, and nothing is contracted until the supplier agrees. Start a buyer request.
Sources
- srvk (GitHub), "dscore: Diarization scoring tools". https://github.com/srvk/dscore
- Kili Technology, "Speaker diarization models: guide, benchmarks and failure modes" (2026). https://kili-technology.com/blog/speaker-diarization-models-guide-benchmarks-and-failure-modes-2026
- arXiv, "The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition" (2023). https://arxiv.org/pdf/2303.06326
- arXiv, "arXiv:2208.08042 (diarization evaluation of short conversational phrases)" (2022). https://www.arxiv.org/pdf/2208.08042
- arXiv, "Summary of the DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments" (2023). https://arxiv.org/pdf/2311.12564
- Northcutt, Athalye, Mueller (arXiv / NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.