Skip to content

Evaluation and benchmarking datasets

Building a golden dataset for LLM evaluation from real business records

Quick answer

A golden dataset for LLM evaluation is a frozen, versioned set of real inputs, each paired with a reference answer that domain experts have independently checked, used to gate releases and compare models. To build one from business records, define what a correct output is before collecting anything, draw cases from records where the right answer was settled and reviewed, validate every reference with measured expert agreement, de-identify without removing what each case tests, and freeze the set before anyone tunes against it.

By SourceX Editorial · Updated

What separates a golden set from an ordinary labeled set

A labeled set becomes golden when three things hold: each reference has been verified rather than labeled once, the set is frozen under a version number, and its composition matches the decision it will drive. Vendor material describes golden datasets as examples crafted and validated to serve as ground truth, often by annotators with subject-matter expertise [1]; the golden dataset glossary entry gives the short definition.

Verification is the property teams most often assume rather than check. Northcutt, Athalye and Mueller estimated an average label error rate of at least 3.3% across the test sets of 10 widely used vision, language and audio datasets, and showed that such errors can change which model ranks higher [2]. If release decisions turn on differences of a few points, reference errors of that size can decide them.

Public benchmarks rarely substitute; one evaluation guide notes they often lack use-case specificity, may be contaminated and need quality and licensing audits [3]. A golden set also differs from a regression suite grown from your own failures; the LLM evaluation datasets hub compares all routes.

Define "correct" for each task before collecting cases

Write down, per task, what a correct output is and how it will be scored before you pull a single record, because the scoring method decides which records can serve as references. A record with a final, reviewed value supports exact-match scoring; a record with only a narrative outcome needs a rubric.

Reference typeWhat the source record must containExample taskScoring
Exact valueA final, reviewed field valueGL account on a posted invoice line; diagnosis code on an audited claimExact or normalized match
Required and forbidden claimsEvidence spans plus claims a correct answer must and must not makeAnswering a policy question from a knowledge baseClaim coverage and citation support
End stateThe system state after the work was doneTicket closed with refund issued and account hold clearedState comparison
Rubric judgmentA reviewed deliverable and the criteria it metA contract redline or customer replyExpert rubric scores, or a judge calibrated against them

The claims format follows practitioner guidance on RAG ground truth: find the authoritative evidence first, then write required and forbidden claims so differently worded answers still score [4]. The same guide warns that an LLM inventing a question and its "truth" without verification yields a second model output, not ground truth [4]. See rubric design with domain experts and LLM-as-a-judge calibration sets.

Where business records already hold a settled answer

The strongest golden cases come from records where the business already made a decision and someone reviewed it, so the reference reflects real practice. Practitioner guidance lists support tickets, human corrections, escalations, refunds and abandoned workflows as candidate sources, and stresses storing the full context, such as retrieved documents, tool results and prior turns [5].

Record sourceTypical systemsWhat supplies the referenceTrap to screen for
Resolved support ticketsHelpdesks such as Zendesk or ServiceNowFinal resolution, policy applied, QA scoreReopened tickets; policy changed later
Corrections of automated outputERP invoice capture, document intake queuesThe corrected field valuesOnly flagged documents get corrected
Escalations and refundsCRM case history, billing systemsOutcome and refund reasonGoodwill refunds that ignore policy
Adjudicated claimsClaims administration systemsPaid, denied or adjusted, with reason codesAppeals that later reversed the decision
Executed contracts versus draftsContract lifecycle management toolsClause positions in the signed versionCommercial trade-offs, not legal judgments
Merged code changesGit hosting with issue trackers and CIThe merged patch and the tests it passedTests that check the patch, not the behavior

SWE-bench showed the pattern for code: its 2,294 tasks came from real GitHub issues and the pull requests that resolved them, across 12 Python repositories, with resolution checked by tests [6]. The last column matters as much as the source, because an outcome field is a reference only if it was final and correct; see verifying outcome labels in operational records and outcome-labeled evaluation data.

When your own production data cannot supply the cases

Your own traffic is the natural first source, but it falls short before launch, in a new domain, when rare high-stakes slices are too thin, when privacy commitments do not cover reuse, and when comparing model vendors needs a set independent of your system's history. Records from companies that do the work can fill those gaps.

Check commitments before reusing customer logs. In January 2024, FTC technology staff warned that model-as-a-service companies may be liable if they break promises not to use customer data for undisclosed purposes such as training or updating models [7]; whether your terms cover building evaluation sets is a question for counsel. Another company's records also show how the task is done with no model involved, which suits measuring task competence rather than regressions.

SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing agreement. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match. Buyers describe the records and outcome fields they need, not the businesses, and every release is approved by the supplying company. See evaluation datasets built from real business work, or describe the records your golden set needs.

Composing the set: a production-matched core and an over-weighted tail

Split a golden set into a core sampled to mirror the production mix, so the aggregate score predicts real performance, and a tail that over-samples rare and high-risk cases, scored per slice. Do not blend the two into one headline number.

Research on data representativity makes the distinction precise: a dataset can be representative as a miniature of the target population, which suits average error, or as coverage of the input space, which is more robust to distribution shift and to accuracy gaps between groups [8].

  • Core. Sample by stratum (product line, document type, channel, month) in proportion to current traffic or, for a new domain, the supplier's case volumes.
  • Tail. Set a minimum count per high-risk slice, such as disputed refunds, multi-currency invoices or appealed claims; see stratified evaluation sets for rare and high-risk cases.
  • Size. Work back from the smallest difference a release decision must detect. An ICML 2025 position paper argues that central-limit-theorem confidence intervals come out too narrow below a few hundred items [9]; sizing an eval set for statistical power has the arithmetic.

Expert validation: turning a recorded outcome into a trusted reference

A recorded outcome becomes a golden reference only after qualified reviewers confirm it independently, their agreement is measured, and disagreements are settled under a written adjudication rule. Most of a golden set's cost and value sits here.

  1. Blind re-derivation on a sample. Two domain experts produce the reference without seeing the recorded outcome; their overrule rate shows how far that outcome field can be trusted.
  2. Agreement statistics. Measure agreement with Cohen's kappa for two raters or Krippendorff's alpha for three or more; one calibration guide stresses that if human reviewers cannot apply the rubric consistently, no amount of judge tuning will fix it [10]. In Krippendorff's definition, alpha of 1 indicates perfect reliability and 0 the absence of reliability [11].
  3. Adjudication. A senior reviewer resolves each disagreement with a written rationale; unsettled cases go to an ambiguous slice with acceptable variants, not a forced label.
  4. Error screening. Run a label-error detector such as confident learning and send flagged cases to experts as candidates: in Northcutt et al.'s audit, human reviewers confirmed about 51% of flagged candidates on average [2].
  5. Conflicts. Exclude reviewers whose own decisions are in the set; see sourcing domain-expert raters.

SWE-bench Verified shows the payoff and the limit. Human annotators screened the original benchmark for overly specific unit tests, underspecified problem descriptions and unreliable environment setup, producing a 500-item verified subset [12]; in a post dated 23 February 2026, OpenAI said it had stopped reporting it because it had become increasingly contaminated [13]. Validation fixes references, not exposure (contamination-resistant evaluation design). For a set built by a third party, the same checks become gold-label acceptance audits.

What a golden record should carry

A golden record should hold the input, the reference, the recorded outcome it came from, the validation evidence and the de-identification applied, so a reviewer can see why the reference is trusted. Store one record per line in JSON Lines [14] and freeze each version with a manifest of case IDs and content hashes.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "case_id": "ap-exc-00318",
  "golden_version": "v3.0",
  "split": "gate",
  "task": "invoice_exception_resolution",
  "input": {
    "invoice": "docs/ap-exc-00318/invoice.pdf",
    "purchase_order": "docs/ap-exc-00318/po.json",
    "goods_receipt": "docs/ap-exc-00318/receipt.json",
    "policy_ref": "ap-policy-2026-03#price-variance"
  },
  "reference": {
    "action": "hold_price_variance",
    "gl_account": "5010",
    "required_claims": ["unit price exceeds PO price beyond tolerance", "freight billed on a separate line"],
    "forbidden_claims": ["approve for payment"],
    "acceptable_variants": ["route_to_buyer_for_price_approval"]
  },
  "provenance": {
    "source_system": "erp_ap_workflow",
    "record_created_at": "2026-04-17",
    "recorded_outcome": "hold_price_variance",
    "outcome_final": true,
    "reopened": false,
    "policy_version_at_decision": "ap-policy-2026-03"
  },
  "validation": {
    "method": "blind_expert_rederivation",
    "reviewers": ["rev-ap-02", "rev-ap-05"],
    "agreed_with_record": true,
    "adjudicated": false,
    "rationale": "Invoice unit price above PO tolerance in the March 2026 policy."
  },
  "deidentification": {
    "method": "consistent_surrogates",
    "fields_replaced": ["vendor_name", "remit_to_address", "bank_account"],
    "property_under_test_preserved": ["price_variance", "po_line_match", "currency"],
    "rerun_after_processing": "passed"
  },
  "slices": ["price_variance", "multi_line_po", "tail"]
}

The record is pretty-printed here; in the file it occupies one line. recorded_outcome beside agreed_with_record shows whether the reference came from the record or from expert correction; rerun_after_processing shows de-identification was checked against the property under test. RAG tools use their own fields: the Ragas v0.1 documentation lists a question, retrieved contexts, the answer and a ground_truth reference [15], so check field names against the release you run (see question-answer-citation triples).

De-identifying cases without erasing what they test

Replace identifiers with consistent surrogates, then re-run each case to confirm the behavior it tests still occurs, because broken coreference or formatting quietly turns a hard case into an easy one. Practitioner guidance warns that replacing every entity with REDACTED can destroy coreference, formatting or retrieval behavior, and suggests keeping full traces in a restricted system with only a pointer in the evaluation repository [5].

Regulated records narrow the options. Under the HIPAA Privacy Rule, Safe Harbor requires removing 18 listed identifier types, including all elements of dates (except the year) directly related to an individual, while Expert Determination relies on a qualified expert finding the identification risk very small, with no numerical threshold set by HHS [16]. A case whose answer depends on a filing deadline or refund window may not survive Safe Harbor, so ask early whether an Expert Determination could preserve the intervals.

For health records, SourceX requires HIPAA de-identification by Safe Harbor or Expert Determination before anything is considered for a license. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded per dataset, and a sample is checked after processing; no method is perfect. Techniques are compared in de-identifying evaluation data without breaking the test.

Freezing, versioning and retiring cases

Freeze each version, change it only through a documented release, and compare scores only within a version. One practitioner puts it bluntly: sample the set, freeze it, version it like code and never let it grow organically [17].

  • Manifest. Case IDs, content hashes, slice counts, rubric version and a changelog for every release.
  • Splits. A development split for prompt and judge tuning, and a gate split nobody tunes against.
  • Retirement. Retire or re-reference a case when its policy changes, when an expert finds its reference wrong, or when it has been exposed outside your control (keeping a private eval set private).
  • Documentation. ISO/IEC 5259-4 gives a data quality process framework for training and evaluation data, from acquisition through labelling and evaluation [18]. NeurIPS 2026 requires Croissant-RAI-based metadata (limitations, potential biases, intended use) for its Evaluations and Datasets Track [19].

In the EU, Article 10 of the AI Act requires training, validation and testing data sets of high-risk systems to be relevant, sufficiently representative and, to the best extent possible, free of errors and complete in view of the intended purpose [20]. As of October 2026, Regulation (EU) 2026/1744, published 24 July 2026, has reportedly moved the Annex III high-risk date to 2 December 2027, also amending Article 10 [21]; see EU AI Act Article 10 for training, validation and test data.

License terms a golden set needs from a record supplier

A golden set built from another company's records needs rights a training license may not grant: sending cases to hosted models, outside expert relabeling, retaining frozen versions and publishing aggregate results. Settle these before scope is agreed.

  • Permitted use names evaluation and the endpoints cases may reach, including hosted model APIs and any judge model
  • Outside reviewers and adjudicators may access records under named confidentiality terms
  • Frozen versions may be retained for planned comparisons, and the license says what happens to them at term end
  • Aggregate scores may be published while cases and outputs stay private (publishing results on licensed evaluation data)
  • Each record carries source system, creation date, final and reopened flags, and the policy version at decision time
  • The supplier confirms the records never appeared in help centers, forums, court filings or case studies
  • Later time windows can be licensed for refreshed versions

Negotiating points are in evaluation-only data license terms. At SourceX, every dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and is delivered under a license defining the records included, permitted uses, term and delivery. Diligence materials on source, rights, preparation and allowed use are prepared per dataset for your review.

Need real business records for a golden evaluation set?

Describe the task, the outcome fields that must serve as references, the time window and slices you need, and the evaluation uses the license must cover. SourceX looks for US companies that hold such records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Describe your evaluation data needs.

Sources

  1. Sigma.ai, "Golden datasets: Evaluating fine-tuned large language models" (vendor page). https://sigma.ai/?p=26358
  2. Northcutt, Athalye and Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks," NeurIPS Datasets and Benchmarks (2021). https://arxiv.org/abs/2103.14749
  3. arXiv, "A Practical Guide for Evaluating LLMs and LLM-Reliant Systems" (2025). https://arxiv.org/pdf/2506.13023
  4. OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
  5. OneUptime, "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
  6. Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" ICLR (2024). https://arxiv.org/pdf/2310.06770
  7. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  8. arXiv, "Data Representativity for Machine Learning and AI Systems" (2022). https://arxiv.org/pdf/2203.04706
  9. ICML, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  10. OneUptime, "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
  11. Krippendorff, "Computing Krippendorff's Alpha-Reliability," University of Pennsylvania (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  12. OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
  13. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  14. jsonlines.org, "JSON Lines." https://jsonlines.org/
  15. Ragas, "Prepare your test dataset" (v0.1.21 documentation). https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
  16. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  17. Alpesh Nakrani, "Golden eval set" (practitioner blog). https://alpeshnakrani.com/blog/golden-eval-set/
  18. ISO/IEC, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  19. NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  20. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance." https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  21. Official Journal of the European Union, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data