Data quality, coverage and contamination
Resolving Annotator Disagreement: Adjudication, Majority Vote and Soft Labels
Quick answer
Annotation adjudication is the policy that turns several annotators' conflicting labels into what a dataset actually delivers. The main options are majority vote, expert adjudication, probabilistic aggregation such as Dawid-Skene, and soft labels that keep the vote distribution [1]. Choose based on whether disagreement is error or genuine ambiguity, and on whether the data feeds SFT, reward modeling or evaluation. Whatever the method, require the raw per-annotator labels alongside the resolved label so you can re-resolve later.
By SourceX Editorial · Updated
This page covers resolution policy only. Measuring agreement is covered in inter-annotator agreement metrics, and auditing a supplier's QA process in gold questions, honeypots and consensus. For the broader quality picture, start at the training data quality hub.
What each resolution method assumes about disagreement
Each method encodes a belief about why annotators disagree, and the wrong belief produces the wrong label. Majority vote assumes annotators are equally reliable and that errors are independent. Expert adjudication assumes one person (or a small panel) can see the right answer. Probabilistic aggregation assumes a single true label exists but that annotators have different, estimable error patterns. Soft labels assume the disagreement itself carries information about the item.
The research literature treats these as a spectrum rather than rivals. A general aggregation model can handle simple classification and more complex outputs such as spans or multi-object annotations [1]. Learning-from-disagreement research, surveyed by Uma and colleagues in 2021 [7], compares methods that abandon the single-gold assumption; even then, teams still need to agree up front on how models will be evaluated.
Majority vote: cheap, and wrong in predictable ways
Majority vote is acceptable when tasks are objective, annotators are calibrated and each item has at least three independent judgments. Its failure mode is that a confident-looking label can be wrong when the majority shares a misreading of the guideline; one industry guide warns that majority-vote labels can look clean while hiding errors [3]. Ties on two-annotator or even-numbered panels need an explicit rule, and "first annotator wins" is not a rule a buyer should accept.
Systematic errors survive voting. Northcutt and colleagues estimated an average label error rate of at least 3.3% in the test sets of 10 widely used benchmarks, including crowd-labeled sets such as ImageNet [4]. If your use is evaluation, that residual error caps the accuracy you can measure.
Consensus thresholds and the inconclusive bucket
A consensus threshold is the minimum agreement required before an item gets a final label, and items below it should be flagged rather than forced. The EPIC-KITCHENS VISOR benchmark offers a concrete pattern: a label is accepted only when 6 of up to 9 workers agree, and otherwise the item is marked inconclusive [2]. That design keeps low-agreement items out of the accepted set instead of hiding them behind a coin-flip majority.
For buyers, the useful questions are where the threshold sits, how many judgments each item received before resolution, and what happens to the inconclusive bucket. Common destinations are expert adjudication, removal, or delivery with a flag. Removal quietly shrinks coverage of hard cases, which matters for long-tail and edge-case coverage.
Expert adjudication: when a third reader decides
Expert adjudication is the right default when labels depend on domain judgment, such as clinical coding, legal clause classification or insurance claim outcomes. The adjudicator should see the guideline, the item and the competing labels, and should record a reason code, not just a final answer. Adjudication decisions should also feed back into the guideline so the same disagreement does not recur [3].
The main failure mode is a single adjudicator who becomes the de facto ground truth. Ask who the adjudicators are and how they are qualified (see verifying domain-expert annotators). Ask whether adjudicators were blind to annotator identity, and what share of items they overturned. A very low overturn rate can mean the adjudicator rubber-stamps the majority.
Probabilistic aggregation: Dawid-Skene and its descendants
Probabilistic aggregation estimates each annotator's reliability and weights their votes accordingly. The classic Dawid-Skene model estimates a confusion matrix per annotator without knowing the true answers, using expectation-maximization: it alternates between estimating each item's label probabilities and re-estimating each annotator's error pattern. Later models generalize the same idea to spans, boxes and multi-object outputs [1].
These models earn their keep when annotators vary widely in skill and each annotator labels enough items to estimate a confusion matrix. They struggle when most annotators label only a handful of items, when the label space is large, or when errors are correlated because everyone misreads the same instruction. The output is a posterior per item, so ask suppliers to deliver that posterior, not just its argmax.
Soft labels: keeping disagreement as signal
Soft labels deliver the distribution of annotator judgments (for example 0.6 / 0.3 / 0.1 across three classes) instead of a single class. They suit subjective or ambiguous tasks such as toxicity, sentiment, helpfulness or document-layout boundaries, where disagreement reflects real uncertainty rather than error. Training against the distribution penalizes a model less for predicting the minority answer on a genuinely contested item, and many teams combine a hard label with the distribution rather than choosing one.
Preference data is the clearest case. Llama 2's annotators picked the better of two outputs and also rated how strongly they preferred it, and the reward model used that rating as a margin [5]. A bare pairwise label hides how contested the comparison was, so keeping per-rater choices, or a margin or confidence field, lets you down-weight contested pairs; see preference data quality.
Choosing a policy by use: SFT, reward modeling, evaluation
The right resolution depends on what the label will do downstream. Training tolerates some noise and benefits from distributions, while evaluation needs labels you can defend item by item. The table below is a starting point, not a rule.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Use | Task character | Suggested resolution | Deliver alongside |
|---|---|---|---|
| SFT (classification or extraction targets) | Objective, guideline-driven | Majority vote with tie rule, or Dawid-Skene | Raw labels, per-item agreement |
| SFT (subjective attributes) | Ambiguous by nature | Soft labels or majority plus agreement weight | Vote distribution |
| Reward modeling | Pairwise preference | Keep per-rater choices; drop or down-weight low-agreement pairs | Rater IDs, margin or confidence |
| Evaluation / test set | Must be defensible | Expert adjudication of all non-unanimous items | Adjudicator reason codes, inconclusive list |
| Evaluation of subjective tasks | Ambiguous by nature | Report against the distribution, not a single gold | Full annotation matrix |
DocLayNet shows a related use of multiple annotations: a subset of pages was annotated two or three times to measure agreement, which then served as a reference point for model performance [6]. Evaluation buyers can ask for the same double-annotated slice even when the rest of a dataset is single-pass.
What to request in a multi-annotated label delivery
A resolved label without its history cannot be re-resolved, audited or re-weighted. Ask for the annotation matrix and the resolution metadata as part of the delivery specification, and describe this in your request so suppliers can say whether the history exists.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"item_id": "tkt-000412",
"guideline_version": "v3.2",
"raw_labels": [
{"annotator_id": "a17", "label": "billing_dispute", "seconds": 41},
{"annotator_id": "a03", "label": "refund_request", "seconds": 28},
{"annotator_id": "a22", "label": "billing_dispute", "seconds": 35}
],
"soft_label": {"billing_dispute": 0.67, "refund_request": 0.33},
"resolution_method": "expert_adjudication",
"resolved_label": "billing_dispute",
"adjudicator_id": "x02",
"adjudication_reason": "customer disputes charge; refund is secondary",
"status": "accepted"
}
Checklist for the specification:
- Number of independent judgments per item, and whether annotators saw each other's labels.
- Pseudonymous, stable annotator IDs so reliability can be estimated across items.
- Resolution method per item, plus the consensus threshold and tie rule.
- Inconclusive items delivered with a flag rather than silently dropped.
- Guideline version per label, since guideline changes shift agreement.
- Adjudicator reason codes and overturn rate for adjudicated items.
These fields also support human annotation provenance and make an annotation quality audit possible before acceptance. If you are sourcing operational records where labels came from business decisions rather than annotators, the question changes to whether outcome fields are reliable labels.
How SourceX fits multi-annotated label requests
SourceX sources operational datasets from US companies on request and manages the commercial process, including the license; it does not hold inventory, and a request does not guarantee a match. Describe the label history you need, such as raw per-annotator labels and resolution metadata, in your request through SourceX for AI data buyers. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. For the basics of the term, see the data annotation glossary entry.
Request multi-annotated data with its label history
SourceX sources datasets on request from US companies, with every release approved by the supplying company and terms agreed in a license before anything is transacted. Describe the data and the label history you need, and SourceX will assess whether a supplier holds it. Start a request at sourcex.si/buyers.
Frequently asked questions
Is majority vote ever enough for a test set?
It can be for objective tasks with several calibrated annotators per item, but label errors persist even in widely used benchmark test sets [4]. Adjudicating non-unanimous items is a safer default for evaluation.
Should low-agreement items be removed from training data?
Not automatically. Removal can strip out the hard and ambiguous cases a model most needs; flag them and decide per use [2].
Can I compute soft labels myself if a supplier delivers only final labels?
No. Soft labels and Dawid-Skene posteriors both require the raw per-annotator labels, which is why the annotation matrix belongs in the delivery specification [1].
Sources
- arXiv, "A General Model for Aggregating Annotations Across Simple, Complex, and Multi-Object Annotation Tasks" (2023). https://arxiv.org/pdf/2312.13437
- arXiv, "EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations" (2022). https://arxiv.org/pdf/2209.13064
- Koji, "Data annotation quality guide". https://www.koji.so/docs/data-annotation-quality-guide
- arXiv, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv, "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
- arXiv, "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
- Journal of Artificial Intelligence Research (Uma et al.), "Learning from Disagreement: A Survey" (2021). https://www.jair.org/index.php/jair/article/view/12712
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.