Skip to content

Data quality, coverage and contamination

Gold Questions, Honeypots and Consensus: Checking a Supplier's Annotation QA Process

Quick answer

Labeling suppliers control quality during production with three mechanisms: a qualification test before an annotator touches real work, honeypot (gold) questions with known answers hidden in the live queue, and consensus, where several annotators label the same item and their answers are aggregated [1]. To evaluate a supplier, ask for the thresholds behind each mechanism, the size and refresh rate of the gold pool, who authored the gold answers, and how disagreements are adjudicated rather than simply outvoted.

By SourceX Editorial · Updated

This page covers in-process controls. For checking a finished delivery, see how to audit annotation quality before acceptance; for the wider framework, start at the training data quality hub.

What the three mechanisms actually measure

Each mechanism measures something different, so a supplier that runs only one has a blind spot. Qualification measures whether an annotator understood the guideline at onboarding, honeypots measure whether they keep applying it under production pace, and consensus measures whether independent annotators converge on the same item [1].

  • Qualification task. A fixed test set with reference answers, scored before access to paid work. It screens out people who misread the ontology, but it says nothing about drift in week six.
  • Honeypot or gold question. An item with a known answer inserted into the normal queue so the annotator cannot tell it apart. Accuracy on these items gives a running per-annotator score, and aggregation models use the same signal to estimate each worker's reliability when weighting their labels [4].
  • Consensus task. The same item goes to N annotators; the label is accepted by vote, by a reliability-weighted model, or sent to review when agreement is too low [1][4].

The term golden dataset is sometimes used loosely for all of these. In a supplier review, keep them separate: a gold evaluation set used to score your model is not the same artifact as the gold control items used to score annotators, and the two should never share items.

What documented thresholds look like

Published thresholds exist, but they are project choices, not industry standards. The EPIC-KITCHENS VISOR paper documents one such configuration: workers had to reach 80% on a qualifier, maintain 90% accuracy on known samples during work, and a label was accepted only when 6 of up to 9 workers agreed [2]. That is a demanding regime suited to a benchmark, where label noise directly distorts reported results.

On honeypot density, Dataloop's documentation suggests mixing in roughly 3-5% honeypot items [1]. Treat that as one vendor's guidance. The right share depends on how fast annotators drift, how expensive each item is, and how many items each annotator completes; an annotator who labels 200 items a day at 3% sees about six gold items, which is too few to separate a 92% annotator from an 85% one with confidence on any single day.

ISO/IEC 5259-4 frames this as a process question: it sets out organizational approaches to data quality for training and evaluation data, explicitly including labelling for supervised ML, without prescribing numeric cut-offs [7]. The standard is useful as a reference when you ask a supplier to describe its process in writing, not as a source of thresholds.

How large and how fresh the gold pool must be

A gold pool that is too small stops measuring quality and starts measuring memory. CVAT's guidance warns that when the validation pool is small, annotators learn to recognize the gold frames [3]. Once that happens, gold accuracy stays high while production accuracy falls, which is the worst failure because the dashboard looks healthy.

Common symptoms that a pool has been learned:

  • Gold accuracy near 100% for long-tenured annotators while spot-check accuracy on random production items is materially lower.
  • Gold items answered markedly faster than comparable production items.
  • Gold items that are visually or textually distinctive (unusual resolution, a fixed set of 40 transcripts, a recurring customer name).
  • No change in gold accuracy after a guideline revision that should have changed some answers.

Ask how many gold items exist per label class, how often new items are added and old ones retired, and whether gold items are drawn from the same distribution as production. Gold items that over-represent easy cases inflate scores; gold items that are all edge cases punish annotators for the guideline's own ambiguity.

Who wrote the gold answers, and whether they are right

A gold question is only as good as its reference answer. Northcutt and colleagues found label errors in the test sets of widely used benchmarks [5], which is a reminder that "reference" labels written by humans carry their own error rate. If a gold answer is wrong, every annotator who labels correctly is penalized and may be removed.

For each gold set, a supplier should be able to say who authored the answers (a guideline owner, a domain expert, or the buyer), how many people reviewed each one, and how annotators can dispute a gold item. A healthy process logs gold disputes and retires items that are frequently and correctly disputed. For specialist work, pair this with verifying domain-expert annotator qualifications, because the gold author's expertise caps the whole program.

Why consensus can hide real disagreement

Majority vote produces a clean label and discards the information that annotators disagreed. On a genuinely ambiguous item, a 3-2 vote yields the same label as a 5-0 vote, and the training file records both identically. For classifier training this flattens uncertainty; for preference and evaluation data it can erase the very signal you are paying for, as discussed in measuring noise and agreement in preference data.

Reliability-weighted aggregation, where each annotator's vote counts according to an estimated accuracy derived partly from gold questions, generally handles uneven annotator quality better than a flat vote [4]. It does not resolve guideline ambiguity. Ask the supplier to deliver per-item vote distributions or agreement scores alongside the final label, and to report dataset-level agreement with a metric suited to the task; the choice between Cohen's kappa, Fleiss' kappa and Krippendorff's alpha depends on annotator count, missing labels and label type [6]. The mechanics of adjudication and soft labels are covered in resolving annotator disagreement.

What happens to annotators who fail

The consequence of failing a control matters as much as the threshold. A 90% gold bar means little if a failing annotator keeps working while their past items stay in the delivered set. Ask whether items labeled by an annotator during a failing window are re-queued, re-reviewed or kept, and whether removal is automatic or a team lead decision.

Also ask about the incentive structure. If annotators are paid per item and only lightly penalized for gold misses, speed wins; if gold failures trigger immediate removal with no retraining path, annotators may refuse ambiguous items or default to the most common label. Neither shows up in an aggregate accuracy figure, which is why the quality report a supplier hands over should include per-annotator distributions, not just a single mean.

Supplier QA questionnaire and illustrative QA record

Use the questionnaire below in a vendor review or attach it to a statement of work; the record shows the per-item fields worth requesting so you can re-run QA analysis yourself.

AreaQuestion to askAnswer that should prompt follow-up
QualificationWhat is the pass mark, how many items, and are qualifier items reused in production gold?No fixed pass mark; qualifier and gold sets overlap
Honeypot shareWhat share of each annotator's queue is gold, and is it constant or adaptive?A single global percentage with no per-annotator minimum count
Gold pool sizeHow many gold items per label class, and how often are they refreshed?One small pool unchanged since project start [3]
Gold authorshipWho wrote and reviewed gold answers, and can annotators dispute them?Gold written by one person with no dispute log
Ongoing thresholdWhat rolling accuracy triggers retraining or removal, over what window?Threshold checked monthly on a handful of items
RemediationAre items from a failing window re-reviewed or re-labeled?Removal only; past items kept as delivered
ConsensusHow many annotators per item, what acceptance rule, and what happens below it?Simple majority with ties broken arbitrarily
Disagreement outputWill per-item vote counts or agreement scores be delivered?Only the final label is retained
Agreement metricWhich metric is reported, at what level, and why that metric [6]?Raw percent agreement only
Gold vs. eval separationHow is overlap between gold control items and your evaluation set prevented?No check performed

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "tkt-000412",
  "final_label": "billing_dispute",
  "aggregation_method": "reliability_weighted_vote",
  "annotations": [
    {"annotator_id": "a-17", "label": "billing_dispute", "gold_accuracy_30d": 0.94},
    {"annotator_id": "a-22", "label": "billing_dispute", "gold_accuracy_30d": 0.91},
    {"annotator_id": "a-05", "label": "refund_request", "gold_accuracy_30d": 0.88}
  ],
  "vote_share": 0.67,
  "below_acceptance_threshold": false,
  "adjudicated": false,
  "is_gold_item": false,
  "guideline_version": "v3.2",
  "annotator_ids_pseudonymous": true
}

Fields such as guideline_version and pseudonymous annotator IDs also feed your human annotation provenance record, so the same export serves quality review and documentation.

How this fits into acquiring labeled operational data

In-process controls reduce error but do not replace acceptance testing. Combine the questionnaire above with a sample audit on delivery and with the broader checklist for evaluating data supplier quality. When the underlying records are business data rather than open content, rights and preparation matter as much as label accuracy.

SourceX sources operational datasets from US companies, such as support histories, engineering records and documents, and manages the commercial process, including licensing agreements. Every dataset is delivered under a license defining records, uses, term and delivery. If you are scoping expert annotations and labels, you can describe the data you need to SourceX.

Sourcing annotated data with a QA process you can inspect

SourceX sources data on request rather than holding stock, so describe the records, labels and QA documentation you expect, and note that a request does not guarantee a match. Each dataset is rights-reviewed and comes with diligence materials covering source, preparation and allowed use. Start a request at sourcex.si/buyers.

Sources

  1. Dataloop, "Creating Consensus, Honeypot, and Qualification Tasks". https://developers.dataloop.ai/tutorials/task_workflows/quality_control/chapter
  2. arXiv (Darkhalil et al.), "EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations" (2022). https://arxiv.org/pdf/2209.13064
  3. CVAT.ai, "Labeling quality control". https://www.cvat.ai/academy/labeling-quality-control
  4. arXiv, "A General Model for Aggregating Annotations Across Simple, Complex, and Multi-Object Annotation Tasks" (2023). https://arxiv.org/pdf/2312.13437
  5. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. arXiv, "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation" (2026). https://arxiv.org/pdf/2603.06865
  7. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data