Skip to content

Evaluation and benchmarking datasets

Private evaluation sets vs public benchmarks: when you need held-out data

Quick answer

Public benchmarks are enough to shortlist models and to check a vendor's published score, but not to choose between finalists or gate a launch. Their items are often in training data, their answer keys can contain errors, and their tasks are not yours. A private evaluation set, held out from every training pipeline and built from your own inputs, policies and outcomes, is the test that should decide a release. Run both: the public benchmark for the shortlist, the private set for the decision.

By SourceX Editorial · Updated

Definitions: held-out data and benchmark contamination. The LLM evaluation datasets hub maps the main sources of evaluation data.

What a public benchmark score proves, and where it misleads

A public benchmark score proves how a model performed on a fixed, published item set under one version and one harness. It does not prove that the model never saw those items, that the answer key is right, or that the task resembles the one you will ship.

Contamination. A 2024 study of benchmark inflation notes that many LLMs' training data contains benchmark test data and that a private holdout could verify published scores, yet most benchmarks have none; its authors build "retro-holdouts", sets constructed after release to match a benchmark, to reveal the gap [1]. The LiveBench authors call test-set contamination a well-documented obstacle that can quickly make a benchmark obsolete, citing Codeforces performance that drops after a model's training cutoff and a hand-crafted GSM8K variant showing that several models had overfit to GSM8K [2].

OpenAI said SWE-bench Verified had become increasingly contaminated, so score gains increasingly reflected training-time exposure [3]. It stopped reporting the benchmark and recommended SWE-bench Pro in the interim, noting that Pro is not perfect but empirically seems to suffer less from contamination [3].

Paraphrased leaks. N-gram overlap is the predominant decontamination check; the GPT-3 work treated a 13-gram match as a contamination signal [4]. LMSYS researchers showed that a 13B model trained on rephrased MMLU test items scored 85.9 on MMLU while n-gram overlap failed to flag the leak [5]. A "decontaminated" label means little until you know the method.

Answer-key errors. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets, and showed that such errors can change which model ranks higher [6]. A two-point gap between finalists can sit inside that noise.

Task mismatch. A 2025 practical guide notes that public benchmarks often lack use-case specificity and recommends evaluation data scoped to the system's tasks and kept distinct from training data [7]. A high score on multiple-choice reasoning says nothing about whether an agent applies your refund policy to a partial return or extracts the right total from your invoice layout.

Version and harness drift. Scores are tied to a release and an evaluation harness. SWE-bench Verified was a 500-item human-validated subset of SWE-bench that shipped with its own Docker-based harness [8], so numbers from SWE-bench, Verified and Pro are not comparable. Ask which version, harness and prompt settings produced any score a vendor quotes.

Who can see the items: public, held-out and private sets compared

The options differ in who can see the test items: anyone, only a benchmark maintainer, or only you and your supplier. Visibility sets both contamination exposure and how much an outsider can verify.

SWE-Bench Pro shows the tiers inside one benchmark. Arguing that permissively licensed public repositories are prime candidates for pre-training crawls, its authors built public and held-out sets from copyleft (GPL) repositories, kept the held-out set private for future overfitting checks, and added a commercial set from startup codebases whose results are published while the code stays private [9]. As market practice, the FinanceBench authors made a sample of 150 cases available open-source [10].

Public benchmarkHeld-out split kept by a maintainerIn-house golden set from your logsLicensed or commissioned private set
Who sees the itemsAnyone, including web crawlersThe maintainer; you see aggregate resultsYour teamYou and the supplier; ask who else
Contamination exposureHigh, and rising with the benchmark's ageLow until the split leaks or is reusedLow if kept out of training, prompt tuning and third-party APIsLow if never published and not resold as training data
Fit to your taskLow to mediumLow to mediumHigh for traffic you already serveHigh if well specified, including rare cases you do not yet see
Comparable with published resultsYesPartlyNoNo
Outside verificationFullPartialNone outside your teamOnly through the documentation supplied

Decision table: are public benchmark scores enough for this decision?

Public scores are enough when the decision is relative and coarse, such as shortlisting models or reproducing a vendor's claim. They are not enough when the decision is specific to your product: choosing between finalists, gating a release, or showing that a system fits a regulated use.

DecisionPublic benchmark enough?WhyWhat to add
Shortlist 10 or more candidate modelsUsuallyCheap and widely publishedMatch version and harness before comparing
Reproduce a vendor's claimed scoreYesReproduction needs the same items and harnessRun it yourself and record the version
Choose between two finalists for your productNoSmall gaps sit within contamination and label noiseA private set of your inputs, scored as a paired comparison (running a model bake-off)
Gate a release or catch regressionsNoMust reflect your policies and failure modes and stay unseen across releasesA frozen, versioned private set with slices for rare and high-risk cases
Evaluate your own fine-tuneNoPurchased training data can contain benchmark items or paraphrasesA held-out split carved before training and decontamination of the training set
Support a high-risk system under the EU AI ActNot on its ownTesting data must fit the intended purpose (see below)A documented test set matched to the intended purpose and population
Publish a research or marketing claimPublic plus a held-out checkReaders need reproducibility; held-out data guards against overfittingPublic results plus a disclosed private check, as SWE-Bench Pro does [9]

For high-risk AI systems in the EU, Article 10(3) of the AI Act requires training, validation and testing data sets to be relevant, sufficiently representative and, to the best extent possible, free of errors and complete in view of the intended purpose [11]. A general benchmark was not built for your intended purpose, so document why your test set fits it. As of October 2026 the Act has been amended by Regulation (EU) 2026/1744, and secondary reports say the high-risk application dates moved to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems [12]. See regulatory requirements for validation and test data.

Rating a benchmark's contamination exposure before you rely on it

Before a benchmark score enters a decision, rate its exposure on six signals. As a working rule (ours, not a published threshold), two or more high-exposure signals mean the score should screen candidates, not decide between them.

SignalLower exposureHigher exposure
Item visibilityItems gated or never published; answers scored server-sideItems and answers in plaintext on GitHub or Hugging Face
Release date vs training cutoffItems created after the model's cutoffBenchmark published long before the cutoff
Item originWritten new by experts or drawn from private recordsDerived from public web content such as GitHub issues, contest problems or encyclopedia text
RefreshQuestions rotated from recent sources, as LiveBench does [2]Static since release
Disclosed decontaminationMethod named, including paraphrase or embedding checksNone, or verbatim n-gram matching only
Score patternSimilar on held-out, rephrased or post-cutoff variantsDrops on held-out or post-cutoff variants

Two checks need no cooperation from the model developer. Where items carry dates, compare scores before and after the model's training cutoff, the pattern the LiveBench authors cite for Codeforces [2]. Or build a small retro-holdout matched to the public items and compare scores on the two [1]. A gap suggests exposure without proving it; see contamination checks for licensed evaluation data for detection methods and designing contamination-resistant evaluation sets for prevention.

What a private held-out set needs before it can gate a launch

A private set can stand in for a benchmark in a launch decision only if it samples your real inputs, carries ground truth you can defend, is sliced by the failures you care about, is frozen by version, and is documented well enough that reviewers can trust results they cannot reproduce.

  • Inputs from your distribution. Real requests, documents, tickets or code, including formats public sets skip: scanned PDFs, mixed-language threads, malformed CSV exports.
  • Defensible ground truth. Each label names its source: a final business outcome, an expert reference or an adjudicated human grade. Audits of public test sets show human labels carry errors [6], so plan gold-label audits and adjudication before acceptance.
  • Enough items for the gap you need to detect. A 2025 ICML position paper argues that CLT-based confidence intervals are too narrow for evals with fewer than a few hundred items [13]; sizing an eval set for statistical power shows the arithmetic per slice.
  • Creation dates. Timestamps let you keep items created after the models' cutoffs (post-cutoff evaluation data).
  • Documentation. A datasheet records a dataset's motivation, composition, collection process and recommended uses [14]. NeurIPS 2026 requires Croissant-RAI-based responsible-AI metadata for its Evaluations and Datasets Track [15], a sign that machine-readable documentation is becoming the expectation.

A 2025 paper on private data curators notes that open datasets let anyone inspect test data, model predictions and scoring, and that private evaluation gives this up while adding conflict-of-interest risk [16]. A set manifest like the one below recovers part of that transparency for reviewers.

Illustrative example: invented to show structure; it does not describe an available dataset.

eval_set_manifest:
  set_id: support-refunds-launch-gate
  version: 2026-09-r3                 # frozen; any edit creates a new version
  access_tier: private_held_out       # never published; excluded from training and prompt tuning
  decision: "ship only if candidate >= current model on every slice below"
  items_total: 1200
  slices:
    - {name: partial_refund_with_coupon, min_items: 180}
    - {name: chargeback_already_filed, min_items: 120}
    - {name: non_english_customer, min_items: 150}
  source:
    origin: licensed_business_records # alternatives: production_logs, commissioned_experts
    systems: [helpdesk_tickets, order_management, refund_ledger]
    created_between: [2025-12-01, 2026-06-30]   # after the cutoffs of the models under test
  labels:
    ground_truth: final_refund_disposition
    adjudication: two_reviewers_plus_lead_on_disagreement
    audit: random_sample_relabelled_before_acceptance
  decontamination:
    checked_against: [public_benchmarks_in_use, fine_tuning_sets_purchased]
    methods: [ngram_overlap, embedding_near_duplicates]
  exposure_log:
    third_party_apis: no_retention_terms_only
    runs: [{model: candidate_a, set_version: 2026-09-r3, date: 2026-09-12}]
  documentation: datasheet_v1         # motivation, composition, collection, recommended uses
  license: {scope: evaluation_only, publish: aggregate_scores_only}

Sourcing private evaluation data: your logs, commissioned experts or licensed records

Items no model has seen come from three places: your production logs, experts commissioned to write items against your rubric, and licensed business records whose recorded outcomes supply the labels. Launch-grade sets often combine them.

Vendors also sell ready-made private test sets; one vendor's blog argues that a set never published online avoids the risk that models were trained on it [17]. Treat that as a claim to verify, and ask whether the same items are licensed to developers of the models you will test. Independence checks for third-party evaluation vendors lists the questions.

SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing agreement. Datasets are sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset goes through rights review and is delivered under a license that defines which records are included and what they can be used for; names, emails and account numbers are removed or replaced before delivery, although no de-identification method is perfect. If records like these would fill gaps in your held-out set, describe the evaluation data you need, or see the record types on evaluation datasets built from real business work.

Whatever the route, settle before delivery whether use is limited to evaluation, whether items may be sent to third-party model APIs, and whether you may publish scores (evaluation-only data license terms, publishing results on licensed evaluation data).

Keeping a private set useful after the first run

A private set loses value through leakage, saturation and lost comparability, so plan access, refresh and a parallel public run before the first scoring job.

  • Leakage. Every hosted-API call sends test items outside your perimeter, so check retention and training terms and log which model saw which version; canary strings (unique markers embedded in the files) can help detect copies later. See keeping a private eval set private. On the supply side, SourceX delivers through private, access-controlled workflows, never email attachments.
  • Saturation. Once every candidate passes, the set no longer separates them; retire and replace items on a schedule (eval set refresh cadence).
  • Comparability. Keep one or two public benchmarks running in the same harness so your results can still be placed against published ones.

Red flags in benchmark claims and private-set offers

Each sign below means a score or set cannot carry the decision you want it to make.

  • A "contamination-free" claim. LiveBench's own authors describe their benchmark as contamination-limited [2].
  • A "private" set produced by rephrasing public benchmark items. Paraphrases evade n-gram checks [5], and the base items may already be in training data.
  • Labels with no recorded source or audit, and items with no creation dates.
  • A 50-item set used to decide a two-point difference between models.

Need held-out data your candidate models have never seen?

Describe the evaluation data you need: the task, the records or outcomes that should supply labels, the slices, the date window and the license scope. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Describe your evaluation data needs.

Sources

  1. arXiv:2410.09247, "Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts" (2024). https://arxiv.org/pdf/2410.09247.pdf
  2. White et al. (arXiv:2406.19314; ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
  3. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  4. arXiv:2406.04244, "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
  5. LMSYS Org, "LLM Decontaminator blog post (2023-11-14)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  6. Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. arXiv:2506.13023, "A Practical Guide for Evaluating LLMs and LLM-Reliant Systems" (2025). https://arxiv.org/html/2506.13023v1
  8. OpenAI (with the SWE-bench authors), "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
  9. arXiv:2509.16941, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/html/2509.16941v1
  10. Patronus AI (vendor documentation, market practice), "FinanceBench". https://arxiv.org/abs/2311.11944
  11. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  12. European Parliament and Council of the European Union (Official Journal, via EUR-Lex), "Regulation (EU) 2026/1744 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  13. ICML 2025 (poster), "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  14. Gebru et al. (arXiv:1803.09010; Communications of the ACM 2021), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
  15. NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  16. arXiv:2503.04756, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  17. Dataforce (vendor blog, market practice), "Disadvantages of standard LLM benchmarks". https://www.dataforce.ai/blog/disadvantages-standard-llm-benchmarks

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data