Skip to content

Evaluation and benchmarking datasets

Accepting a delivered eval set: gold-label audits and adjudication

Quick answer

Accept a delivered eval set only after an independent audit of its gold labels passes criteria you fixed in the contract. Write down a sample size and an error tolerance, have qualified reviewers re-label a blind random sample, and adjudicate every disagreement against the rubric. Then check label provenance, rubric agreement, schema and slice coverage. Release payment and first use only on a pass, with replacement, re-labeling or holdback as the agreed remedy for a fail.

By SourceX Editorial · Updated

Why gold-label errors matter more in eval data than in training data

A wrong gold label in an eval set directly corrupts the score you will use to choose models, so the tolerance has to be tighter than for training data. A noisy training example is diluted across millions of gradient steps; a noisy test item flips a pass to a fail for every model you run against it. Northcutt and colleagues estimated an average label error rate of at least 3.3% across the test sets of 10 widely used benchmarks, with at least 6% of the ImageNet validation set mislabeled, and showed that such errors can change which model ranks higher [1]. Curated, high-profile benchmarks were not exempt.

Benchmark maintainers have had to re-verify sets after release. SWE-bench Verified was built by having human annotators screen tasks after problems such as overly specific unit tests, underspecified issue descriptions and unreliable environment setup surfaced in the original set [2]. As of October 2026, OpenAI has also flagged SWE-bench Verified as contaminated, which is a separate failure mode: a correct label on a leaked item is still unusable. Treat acceptance as two gates, label correctness here and leakage checks covered in contamination-resistant evaluation design.

This page covers the eval-specific part of acceptance. Generic delivery checks (file counts, schema validity, duplicates, PII scans) belong in your acceptance criteria for licensed training data and apply here too.

Fix the acceptance criteria before the supplier starts labeling

Acceptance criteria only work if both parties signed them before the first label was written. If the threshold is set after you see the data, every borderline item becomes a negotiation. Put the criteria in the statement of work next to the evaluation dataset specification, so the rubric version, label schema and slice quotas are the same documents the auditors will use.

Illustrative example: invented to show structure; it does not describe an available dataset.

CriterionHow it is measuredExample thresholdIf it fails
Gold-label error rateBlind expert re-label of a random sample, errors confirmed by adjudicationUpper 95% bound below 3%Re-label the affected slice; re-audit a fresh sample
Critical-slice error rateSeparate sample drawn from each high-risk sliceZero confirmed errors in the slice sampleReplace items; re-audit that slice
Rubric reliabilityKrippendorff's alpha or Cohen's kappa between independent reviewers on the audit sampleAt least 0.7Revise rubric, re-label, re-test agreement
Label provenancePer-item label_source and reviewer IDs, checked against the sample100% of items carry provenance; no undisclosed model-only labelsTreat undeclared items as unlabeled
Reference answer sufficiencyAuditor confirms the reference is supported by the cited evidenceNo unsupported references in sampleRe-write references with evidence spans
Ambiguity flagsItems marked ambiguous by two or more auditorsBelow an agreed share, each with adjudication notesRemove or rewrite ambiguous items
Slice coverageCount per slice versus spec quotasWithin agreed tolerance of every quotaSupplier supplies missing items

The thresholds above are starting points, not norms. Set the gold-label tolerance from the smallest model difference you need to detect: if candidate models differ by two points, a 3% label error rate can erase the gap.

Size the audit sample from the error rate you can tolerate

The audit sample is set by the maximum error rate you will accept and the confidence you need, not by a percentage of the delivery. For a zero-defect plan, the number of items n you must inspect with no confirmed error to claim the true error rate is below p at 95% confidence is n = ln(0.05) / ln(1 − p).

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Worked example. Tolerance 3%: n = ln(0.05) / ln(0.97), about 99 items with zero confirmed errors. Tolerance 2%: about 149 items. Tolerance 5%: about 59 items.
  • Allowing a few errors. If the plan allows up to 2 confirmed errors, a Poisson approximation puts the sample near 210 items to support the same 3% claim at 95% confidence.
  • Slices. Run the calculation separately for each critical slice. A clean overall sample can hide a 15% error rate in a 40-item slice of rare refusals.

For recurring deliveries, attribute sampling standards such as ANSI/ASQ Z1.4 add switching rules that tighten inspection after failures and relax it after a run of passes [8]. The adaptation of AQL plans to label defects is covered in acceptance sampling for dataset deliveries. Remember that the eval set itself also needs enough items for the comparisons you plan; small evals make naive confidence intervals too narrow [9], which is the subject of sizing an eval set for statistical power.

Run the re-label blind, then adjudicate every disagreement

The core of a gold-label audit is an independent re-label by qualified reviewers who cannot see the supplier's answer. Showing reviewers the gold label first turns an audit into a rubber stamp. Market practice for golden datasets is validation by subject-matter experts [4], so match auditor qualifications to the domain: certified coders for medical codes, licensed practitioners for legal clause labels, senior engineers for code-repair tasks.

A workable protocol:

  1. Draw the random sample with a recorded seed, stratified by slice, and strip gold_label and reference_answer fields.
  2. Two auditors label each item independently using the exact rubric version in the spec.
  3. Compare auditor labels with each other and with the supplier label. Three-way agreement passes the item.
  4. Send every disagreement to an adjudicator, a third expert who sees all labels, the rubric and the evidence, and records a decision with a reason code.
  5. Classify each adjudicated item: supplier label wrong, auditor wrong, item ambiguous, or rubric gap.

Only "supplier label wrong" counts against the error tolerance. "Item ambiguous" and "rubric gap" outcomes are still findings: they mean the item cannot discriminate between models and should be fixed or removed, even if no one made a mistake.

Measure rubric agreement so you know the labels are learnable

Inter-annotator agreement tells you whether the rubric produces consistent labels at all, which sets the ceiling on how correct any gold label can be. Use Cohen's kappa for two raters on categorical labels and Krippendorff's alpha when you have more than two raters, missing ratings or ordinal scales [6]. Metric choice changes the number you report, so name the metric in the contract [7].

One practitioner guide to calibrating LLM judges treats roughly 0.7 human agreement as the floor: if humans cannot reach it on a rubric, the rubric is broken, not the raters [3]. The same logic applies to acceptance. Low agreement between your auditors is evidence the supplier's labels cannot be trusted either, whatever the raw error count says.

For generative tasks with free-text reference answers, agreement is measured on the judgment, not the string. Auditors score whether the reference is correct, complete and supported by the cited evidence; two acceptable phrasings are not a disagreement. Write that distinction into the rubric before the audit.

Check how each label was made, not only whether it looks right

Label provenance is an acceptance criterion in its own right, because a model-written label that looks plausible is not ground truth. When an LLM generates both the question and the "truth," you get a second model output, not a reference [5]. Scoring candidate models against it measures similarity to the generator model.

Require a per-item record of how the label was produced. Fields worth contracting for:

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "ev-000731",
  "slice": "refund_disputes/partial",
  "rubric_version": "2.3",
  "gold_label": "escalate",
  "reference_answer": "Escalate: partial refund exceeds agent authority under policy 4.2.",
  "evidence_spans": [{"doc_id": "pol-4", "start": 1180, "end": 1342}],
  "label_source": "human_expert",
  "model_assist": "draft_suggested_then_edited",
  "labelers": ["L-14", "L-22"],
  "adjudicated": true,
  "adjudication_reason": "rubric_gap_resolved_v2.3",
  "source_record_date": "2025-11-03"
}

Check the declared provenance against reality in the sample. Look for labels that match a known model's phrasing, timestamps too fast for human review, or model_assist values missing on items where drafts were clearly reused. Undeclared synthetic labels are a material defect; the trade-offs of generated sets are covered in synthetic evaluation data limits.

Tie remedies and payment to the audit result

An acceptance test without a remedy is just a report, so connect each failure to a specific action and to payment. Common contract mechanics buyers negotiate include:

  • Re-label at supplier cost for a failed slice, followed by a re-audit on a fresh random sample, not the same items.
  • Replacement items that meet the same slice and difficulty profile, audited before they are merged.
  • Payment holdback of a portion of the fee until the final audit passes, with a defined number of re-delivery cycles.
  • Rejection rights if the second audit fails, with return or deletion of delivered files as the license specifies.

Keep the audit sample out of the scored set, or re-audit with new items, because auditors have now seen it. Restrict who can view gold labels during the audit; access controls and canaries are covered in keeping a private eval set private.

If your system is high-risk under the EU AI Act, Article 10 requires that validation and testing data sets meet its quality and governance criteria [10]. As of October 2026, the amended application dates for high-risk systems reportedly fall in December 2027 (Annex III) and August 2028 (Annex I). Your audit records, adjudication logs and provenance fields are the evidence that work was done.

How sourced eval data from real operations fits this process

Eval sets built from real business records still need the same gold-label audit, because the original record shows what happened, not always what was correct. A closed support ticket records the agent's action; whether that action was the right one is a label someone has to assign and you have to verify.

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows. Every dataset is rights-reviewed for ownership and consents, delivered under a license that defines records, uses, term and delivery, and comes with diligence materials covering source, rights, preparation and allowed use. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Your acceptance audit then runs on top of that. To describe the operational data behind an eval set you plan to audit this way, start at SourceX for AI data buyers, and see evaluation sets built from real business work and the eval set glossary entry. For supplier selection before delivery, read evaluating data supplier quality and running a data pilot with a supplier; for leakage checks see contamination checks for licensed eval data. The broader map is in the evaluation datasets hub and the AI data hub.

Describe the eval set you need

SourceX sources operational datasets from US companies on request rather than from stock, and a request does not guarantee a match. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set in a license per deal. Describe the records and slices your eval set needs at sourcex.si/buyers.

Frequently asked questions

Should the supplier's own QA count toward acceptance?

No. Supplier QA reports (consensus rates, honeypot scores, reviewer pass rates) are useful diligence, but they measure the supplier's process against itself. Acceptance needs reviewers who answer to you and work blind to the supplier's labels.

What if my auditors and the supplier disagree on most items?

Stop counting errors and check the rubric. Widespread disagreement usually means the rubric is underspecified, so measure auditor-to-auditor agreement first; if that is also low, revise the rubric, re-label and restart the audit.

Can an LLM judge replace human auditors?

Not for acceptance of gold labels. A judge can pre-screen items for likely errors, but the judge itself needs calibration against human labels [3], and confirmed errors should come from human adjudication.

Sources

  1. Northcutt, Athalye, Mueller (arXiv; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  2. OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
  3. OneUptime, "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
  4. Sigma.ai, "Golden datasets: Evaluating fine-tuned large language models". https://sigma.ai/?p=26358
  5. OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
  6. Klaus Krippendorff, University of Pennsylvania Annenberg School for Communication, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  7. arXiv, "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation" (2026). https://arxiv.org/pdf/2603.06865
  8. ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes". https://asq.org/quality-press/display-item?item=T1164
  9. ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  10. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data