Skip to content

Data quality, coverage and contamination

How to Audit Annotation Quality in a Labeled Dataset Before You Accept It

Quick answer

An annotation quality audit is an independent re-check of a delivered labeled dataset before you sign acceptance. Freeze the exact delivery, draw a stratified random sample, have qualified reviewers re-label it blind to the supplier's labels, adjudicate disagreements, classify every error, and report the error rate with a confidence interval per stratum. Accept only if the interval's upper bound clears the threshold written into the order, and score against your own gold labels rather than the supplier's QA report.

By SourceX Editorial · Updated

This page covers the end-to-end procedure. Agreement statistics, gold-question design and supplier QA process checks live on separate pages in the data quality hub.

Why buyers need their own audit, not the supplier's QA number

A supplier's reported accuracy measures their process against their own reviewers, so it cannot tell you whether the labels meet your definition of correct. Even benchmark test sets curated by research teams carry an estimated average label error rate of at least 3.3%, including at least 6% of the ImageNet validation set [1]. Buyer checklists from the annotation market recommend scoring a pilot against a ground-truth set with pass/fail criteria fixed before delivery, not trusting a demo [3]. Practitioners add that self-reported accuracy should be re-audited and that rework rate is a useful companion metric [4].

Three failure modes recur in delivered labeled data. Guideline drift happens when a later annotation batch silently applies a revised instruction. Class collapse happens when ambiguous items get pushed into a default class such as "other" or "no issue." Pipeline corruption happens when labels are misaligned to the wrong row after a join, shuffle or deduplication step, which no reviewer-level QA will catch.

Step 1: Freeze the delivery and write the acceptance bar first

Audit a fixed, hashed snapshot so the sample and the accepted dataset are provably the same files. Record the delivery manifest (file names, SHA-256 hashes, row counts per file), the guideline version each batch was labeled under, and the label schema with its allowed values. If the supplier redelivers, the audit restarts on the new snapshot.

Write the acceptance bar before you look at a single label. A usable bar names the metric (item-level error rate, or field-level error rate for extraction), the threshold, the confidence level, which error classes count, and what happens on failure: rework of a stratum, replacement, or rejection. ISO/IEC 5259-2 gives a vocabulary of measurable data quality characteristics you can borrow so the bar is unambiguous [8], and the training data quality metrics page explains how accuracy differs from completeness and consistency.

Step 2: Draw a stratified random sample sized for the decision

Sample size should follow from the error rate you must rule out, not from a round number. Stratify by the dimensions where problems concentrate: source system, time period, annotation batch or vendor team, label class, and guideline version. Within each stratum draw uniformly at random with a recorded seed, and oversample rare classes so a 2% class is not represented by four items.

Two rules of thumb help. To estimate an error rate near 5% within roughly plus or minus 2 percentage points at 95% confidence, plan on about 450 to 460 items per stratum you intend to judge separately. If an audit of n items finds zero errors, the 95% upper bound on the true error rate is roughly 3/n, so 300 clean items supports a claim of about 1% or less.

Use an interval method that behaves at small n and low error rates, such as Wilson or Clopper-Pearson. Normal-approximation (CLT-based) intervals tend to be too narrow on samples of fewer than a few hundred items [6]. For recurring deliveries, the acceptance-sampling logic in ANSI/ASQ Z1.4, with normal, tightened and reduced plans and switching rules, is a reasonable model for adjusting audit intensity batch by batch [7]; its tables were written for manufactured lots, so treat it as a pattern rather than a ready-made plan.

Step 3: Re-label blind with qualified reviewers

Reviewers must label sampled items without seeing the supplier's label, or anchoring will inflate agreement. Load the sample into your own annotation tool as a separate project, strip the supplier's label column, and keep a join key so labels can be compared later. Tools such as CVAT support ground-truth jobs and report quality results against them in QA analytics [5].

Reviewer qualification matters more than reviewer count. For expert work such as clinical coding, contract clause tagging or financial document extraction, use reviewers who pass a qualification test on your guideline; the expert annotator verification guide covers credential and test design. Use at least two independent reviewers on a subset so you can measure your own reviewers' agreement before trusting them as the reference.

Step 4: Adjudicate and classify every error

An error count is only actionable once each disagreement is adjudicated and assigned a cause. A senior adjudicator sees the item, the supplier's label, both reviewer labels and the guideline, then records a final label and a classification. Items the guideline cannot resolve are logged as ambiguous, not as supplier errors; the adjudication and disagreement resolution page covers majority vote, soft labels and escalation.

Illustrative example: invented to show structure; it does not describe an available dataset.

Error classDefinitionTypical causeCounts toward error rate
Wrong classLabel is a valid value but the wrong oneGuideline misread, fatigueYes
Missing labelRequired label or field left empty or defaultedClass collapse into "other"Yes
Boundary or span errorEntity span, bounding box or time range is off beyond toleranceTokenization mismatch, loose toleranceYes, if beyond agreed tolerance
Guideline violationLabel contradicts an explicit rule in the version in forceGuideline drift between batchesYes
MisalignmentLabel belongs to a different recordJoin, shuffle or dedup bugYes, and triggers full-file check
Ambiguous itemGuideline does not determine a single answerUnderspecified guidelineNo; fix the guideline

Misalignment deserves special handling because it is systematic. One confirmed misaligned row means checking the whole file's keys, not just counting one error.

Step 5: Compute, stratify and decide

Report error rates per stratum with intervals, because a dataset-level average can hide one bad batch. Here is a worked calculation for a single stratum.

Illustrative example: invented to show structure; it does not describe an available dataset.

audit_id: AQA-2026-10-example
snapshot_manifest_sha256: "<hash of manifest file>"
stratum: { source_system: "ticketing", period: "2025-Q3", label_field: "issue_category" }
sample_size: 409
scored_items: 400                 # ambiguous items removed from the denominator
adjudicated_errors: 12            # wrong class 7, missing 3, guideline violation 2
ambiguous_excluded: 9
observed_error_rate: 0.030
wilson_95_interval: [0.017, 0.052]
acceptance_bar: "upper bound <= 0.05 at 95% confidence"
decision: "fail - rework stratum and re-audit"

The observed 3.0% looks comfortably under a 5% bar, but the upper bound of about 5.2% does not clear it. The right response is usually to request rework of that stratum, or to audit more items to narrow the interval, rather than rejecting the whole delivery. Track rework rate across rounds as a second signal of supplier process health [4].

Special cases: model-assisted triage and work-generated labels

Model-based error detection is a good way to target the audit, but it is not a substitute for the random sample. Confident learning ranks likely label errors under a class-conditional noise assumption, and the open-source cleanlab package implements it [2]. Run it on the full delivery to find a high-yield review queue, and keep the random sample separate so your error estimate stays unbiased.

Labels generated by business work, such as ticket dispositions, QA scores, claim outcomes or invoice coding, need a different reference. Check them against later outcomes before treating them as ground truth: a ticket marked "resolved" that reopens within a week, or an invoice code later reversed by finance, is evidence the label was wrong at the time. Report the agreement between the original label and the later outcome per period, since process changes often show up as step changes. The buyer's page on licensing expert annotations and labels describes these label types in more detail.

For high-risk systems in the EU, Article 10 of the AI Act requires training, validation and testing data to meet quality criteria under data governance practices that include annotation and labeling [9]. As of October 2026 the high-risk start dates have reportedly moved to December 2027 and August 2028, but an audit record like the one above is useful evidence either way.

What the audit report should hand to the next team

The audit output should let a model team reuse the labels without re-running the audit. Include the snapshot hashes, sampling frame and seed, per-stratum intervals, the error taxonomy with counts, adjudicated corrections as a patch file keyed by record ID, the ambiguous items with proposed guideline edits, and the acceptance decision. The dataset quality report contents page lists what a supplier-side report should contain, which makes comparing the two straightforward. For supplier-level due diligence before purchase, see evaluating data supplier quality, and for the term itself, the data annotation glossary entry.

When you are sourcing labeled operational data, SourceX prepares diligence materials per dataset covering source, rights, preparation and allowed use, which gives your audit a documented starting point. You can describe the labeled data you need on the buyer page.

Getting audit-ready labeled data for annotation quality checks

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, and finance and legal workflows, and every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked. Tell us what labeled data you need at sourcex.si/buyers.

Frequently asked questions

Should I use inter-annotator agreement instead of an error rate?

They answer different questions. Agreement measures whether annotators are consistent with each other; an audit error rate measures whether labels match your adjudicated reference. Use agreement to qualify your own reviewers, and the error rate to make the acceptance decision. The gold questions and consensus page covers how suppliers use both during production.

What if the supplier's guideline differs from mine?

Resolve that before the audit, not during it. Map each supplier label value to your schema, list the rules that differ, and decide whether disagreements caused by a known guideline difference count as errors. Otherwise the audit measures the gap between two guidelines, not label quality.

Can an LLM act as one of the blind reviewers?

It can help triage, but validate it against human adjudication first and keep humans on the acceptance sample. The LLM judge scoring guide covers reliability and bias checks for that setup.

Sources

  1. Northcutt, Athalye, Mueller (NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  2. Northcutt, Jiang, Chuang (arXiv), "Confident Learning: Estimating Uncertainty in Dataset Labels" (2019). https://arxiv.org/pdf/1911.00068
  3. CloudPano, "How to Choose a Data Annotation Partner (Buyer's Checklist)". https://www.cloudpano.com/blog/how-to-choose-a-data-annotation-company
  4. prommer.net, "Training data vendor evaluation guide". https://prommer.net/en/tech/guides/training-data-vendor-eval/
  5. CVAT documentation, "QA analytics". https://docs.cvat.ai/docs/qa-analytics/
  6. ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  7. ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes" (2018). https://asq.org/quality-press/display-item?item=T1164
  8. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and ML, Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  9. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data