Skip to content

Evaluation and benchmarking datasets

Contamination-resistant evaluation: designing test sets that stay out of training data

Quick answer

Contamination-resistant evaluation means deciding where test items come from, how they are split and scored, and who may see them before any model is run, so that leakage into training data is unlikely rather than something you discover afterwards. The main levers are items drawn from never-published records or created after the training cutoffs of the models you test, a held-out tier whose items are never released, objective ground truth, scheduled refresh, and governance over curators and disclosure.

By SourceX Editorial · Updated

Why contamination has to be designed out, not filtered out

Filtering after the fact cannot make a public test set trustworthy, because standard overlap checks miss rewritten copies and nobody outside a lab can inspect a closed model's training corpus. N-gram overlap is still the predominant decontamination technique, and the match thresholds differ from lab to lab [1]. Yang et al. showed that paraphrased or translated test items pass such filters: a 13B model trained on rephrased MMLU items scored 85.9 on MMLU without being flagged by n-gram overlap [2].

The effect shows in widely used benchmarks. In a post dated 23 February 2026, OpenAI said SWE-bench Verified had become increasingly contaminated, so score gains increasingly reflected training-time exposure [3]. It stopped reporting the benchmark and recommended SWE-bench Pro in the interim, which it called not perfect but empirically less prone to contamination issues [3]. The Benchmark Inflation authors make the general point: a private holdout set could be used to verify public scores, but most benchmarks do not have one [4].

For a benchmark builder, an item's contamination status is mostly fixed when you choose its source and decide who sees it. Testing a delivered set for overlap is covered in contamination checks for licensed evaluation data, and the term itself in the benchmark contamination glossary entry.

Four routes from a test set into a training corpus

An item leaks in one of four ways: its text is published, its source material is published, it passes through evaluation traffic or vendor workflows, or a derivative of it is published. Each route needs its own control.

Leakage routeHow it happensDesign controlWhat remains
The item is publishedBenchmark files on GitHub or Hugging Face, examples in papers and model cards, prompts in leaderboard submissionsNever release the held-out tier; treat any released tier as burnedReleased-tier scores lose meaning over time
The item's source is publicQuestions written from Wikipedia, public filings or permissively licensed repositories, so the model has seen the source and often the answerSource from never-published records or post-cutoff materialPublic background knowledge can still answer some items
Evaluation traffic and vendorsItems sent to hosted model APIs or external LLM judges; annotators or vendors reusing items in training-data workRules on who runs the set, on which endpoints, under which termsDepends on enforcement and audit
DerivativesParaphrases, translations, solution write-ups, synthetic training examples generated from look-alike seedsKeep answers and rationales private; refresh items; measure the gap between public and private tiersParaphrased leakage evades n-gram checks [2]

Published text also replicates. Lee et al. found one sentence repeated more than 60,000 times in the C4 web corpus [5], so assume anything posted once can keep circulating in later crawls. SWE-bench shows the source route: it was built from real GitHub issues and pull requests in 12 popular Python repositories [6], and the SWE-Bench Pro authors argue that widely used public repositories, particularly permissively licensed ones, are prime candidates for web-crawled pre-training corpora [7].

Controls for the third route, such as canary strings and hosted-API exposure, are in keeping a private eval set private; generated look-alikes are covered in contamination through synthetic data.

Choosing sources that models are unlikely to have seen

Most of a set's resistance comes from its source: material that was never on the crawlable web, was created after the training cutoff of every model you will test, or sits under license terms that make it less likely to be in pre-training corpora. Each option has a weakness, so strong designs combine at least two.

Source typeWhy it resists contaminationWeaknessEvidence to keep per item
Never-published business records: support tickets, internal code repositories, claims files, finance and legal workflowsContent lived inside company systems, not on the public web"Never published" has to be checked; needs a license and often de-identificationSource system, supplier attestation, publication-check result, license scope
Post-cutoff public material: new competitions, papers, news, datasetsCreated after the model's training data was collectedAges into contamination as new models ship; closed models' cutoffs are self-reportedPublication date, the cutoff it was checked against, planned retirement date
Strong-copyleft public codeLess likely than permissive code to be in pre-training corpora, per the SWE-Bench Pro rationale [7]Reduces exposure rather than removing it; copyleft terms apply if you redistribute the codeLicense identifier and commit hash per repository
Items commissioned for the setThe text did not exist beforeAuthors may lean on public sources or LLM drafts; costly to scaleAuthor declaration, drafting tools used, reviewer sign-off
Templated or synthetic variants of public itemsCheap to generate and refreshInherits the public item's structure and often its answer; see synthetic evaluation data limitsSeed item ID, generator and prompt

SWE-Bench Pro combines sources: a public set from 11 strong-copyleft (GPL) repositories, a held-out set from 12 others, and a commercial set from codebases acquired from startups [7]. LiveBench takes the post-cutoff route, drawing questions from recently released math competitions, arXiv papers, news articles and datasets, and adding harder versions of older tasks from Big-Bench Hard, AMPS and IFEval [8]. Its authors cite evidence that model performance on Codeforces problems drops sharply for problems released after a model's training cutoff [8], which is the effect a post-cutoff design depends on.

"Never published" is a claim about the supplier's whole processing chain, which is why evaluation datasets built from real business work still need checking. A support ticket can surface in a help-center article, forum thread or public issue tracker; a contract can appear in a court filing; a SaaS processor's terms may have allowed training on customer content. Temporal designs also need the system creation time, not a last-modified or export time; see post-cutoff evaluation data and, for code, held-out coding agent evaluation sets from licensed private repositories.

Structuring the set: tiers, matched holdouts and objective answers

Split the set into tiers with different exposure rules, draw the private tier from the same distribution as anything you release, and score against answers that need no judge. Structure is what lets you measure the contamination you could not prevent.

  • Tiers. A common pattern is a public development tier (released and assumed burned), a private held-out tier (items and per-item outputs never published) and an unexposed reserve for the next refresh. SWE-Bench Pro keeps its held-out subset private for future overfitting checks and off public leaderboards, while releasing results, but not code, for its commercial subset [7].
  • Matched holdouts. Sample the private tier from the same sources, time window and difficulty mix as the public tier, so a score gap between them signals overfitting or leakage. Plan this before release: the Benchmark Inflation paper notes that building a statistically indistinguishable "retro-holdout" after a benchmark is public is non-trivial [4].
  • Objective ground truth. LiveBench scores answers automatically against objective ground-truth values and avoids LLM judges and crowdsourced scoring, which its authors say introduce bias and break down on hard questions [8]. Executable tests and exact values from recorded business outcomes also keep a private set scoreable without sending items to an external judge model.
  • Refresh and versioning. The LiveBench paper describes adding and updating questions monthly [8]. Refreshing limits exposure but breaks comparability, so version the set and never compare scores across versions; cadence trade-offs are in eval set refresh cadence.
  • Size for the gap you need to see. A small held-out tier can reveal only a large public-private gap. One ICML 2025 position paper argues that CLT-based confidence intervals come out too narrow on evals with fewer than a few hundred items and recommends other interval methods there [9].

Governance: access, disclosure and curator conflicts

Secrecy works only with written rules on who sees items, what gets published and who may curate, because a private set gives up the transparency that lets outsiders check a public one. The authors of "Peeking Behind Closed Doors" argue that private evaluation curators bring two specific risks: conflicts of interest from their business relationships with the model developers they evaluate, and annotator preferences that tilt results toward models trained on the same curator's data [10].

Decisions to record before the first run:

  • Item access. Name who can read held-out items (authors, adjudicators, the harness operator) and keep anyone producing training data for the same capability out of that group. Ask any curator whether it also sells training data in the same domain [10].
  • Where models run. Every hosted-API run sends items to a third party, so list allowed endpoints, check their retention and training-on-inputs terms, and log exposures per item.
  • Disclosure. Publish methodology, tier sizes, source types, time windows and aggregate scores with intervals; keep held-out items, per-item outputs and rationales private. Licenses can limit disclosure further; see publishing benchmark results on licensed evaluation data.
  • Label audits. Outsiders cannot report errors in items they cannot see. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets and showed such errors can change model rankings [11], so budget for gold-label audits.
  • Independence. If a third party builds or runs the set, apply the checks in third-party eval data vendors: independence and conflicts.

An item record that shows where each test case came from

Record provenance per item rather than per dataset, so that for any test case you can show where it originated, when it was created, which publication check it passed and how often it has been exposed. Item-level records also let you retire exposed items without discarding the whole set.

Illustrative example: invented to show structure; it does not describe an available dataset.

item_id: spb-2026q3-000412
benchmark_version: "2026-10"
tier: heldout_private                 # public_dev | heldout_private | reserve
source:
  origin_type: licensed_business_record   # licensed_business_record | post_cutoff_public | copyleft_public | commissioned
  source_system: helpdesk_ticketing
  record_created_at: "2026-06-14T09:32:00Z"
  timestamp_basis: system_created_at  # not last_modified or export time
  publication_check:
    help_center_and_forum_match: none
    exact_phrase_web_search: none
    supplier_attestation: never_published
    checked_on: "2026-09-30"
  license_scope: evaluation_only
  deidentification: names_and_account_numbers_replaced
item:
  task: refund_policy_decision
  input_ref: inputs/000412.json
  ground_truth:
    type: recorded_outcome            # recorded_outcome | executable_test | exact_value
    value: refund_denied_outside_window
    policy_clause: "returns-2.3"
    adjudication: two_reviewers_plus_tiebreak
authoring:
  rewritten_by: human_editor
  llm_drafting_used: false
exposure:
  allowed_endpoints: [self_hosted, hosted_no_training_on_inputs]
  external_runs: 6
  retire_after_external_runs: 20

Dataset-level documentation can then summarize these fields. Croissant-RAI adds machine-readable responsible-AI fields to the Croissant metadata vocabulary [12], and NeurIPS 2026 sets responsible-AI metadata requirements based on it for its Evaluations and Datasets Track [13]. In the EU, Article 10 of the AI Act applies data-governance practices, including design choices, data collection processes and the origin of data, to the testing data sets of high-risk AI systems as well as their training data [14]. As of October 2026, Regulation (EU) 2026/1744 has reportedly moved the high-risk application dates [16], and that regulation also amends Article 10 [16]; see regulatory requirements for AI validation and test data.

Worked example: tiering a support-policy benchmark

Illustrative example: invented to show structure; it does not describe an available dataset.

A team testing customer-support agents on refund and warranty policy licenses ticket histories created April to September 2026 by a US software company, under an evaluation-only license that permits publishing a small development tier:

  1. Publication check. Drop tickets whose text matches the company's help center, community forum or status page, or quotes public documentation verbatim.
  2. Ground truth. Use the recorded resolution and the policy clause applied, confirmed by two reviewers with a tie-break. Reopened or disputed tickets go to a separate slice.
  3. Tiers. 200 items form the public development tier, released with a canary string (a unique marker corpus builders can filter on). 1,500 form the private held-out tier; 300 stay in reserve.
  4. Matched sampling. Every tier is stratified by product line, policy type and month, so the development-versus-held-out gap is interpretable.
  5. Exposure rule. Held-out items run only on self-hosted models or endpoints whose terms exclude training on inputs; runs are logged per item, and items are retired after a set number of external runs.
  6. Reporting. Scores are reported per version with intervals. A model whose development score beats its held-out score by more than the interval is flagged for review, not ranked.
  7. Refresh. Each quarter a newer month joins the held-out tier and the oldest held-out month moves to the development tier.

De-identification can break items whose answers depend on names, dates or account history; see de-identifying evaluation data without breaking the test.

Questions to settle with a data supplier before licensing evaluation items

Licensed records resist contamination only if the supplier can answer these questions before you agree scope:

  • Has any of this content ever been public: help-center articles, forums, public repositories or issue trackers, court or regulatory filings, published case studies?
  • Has it been shared with or licensed to anyone who may train models on it, including other AI buyers and SaaS processors whose terms allow training on customer content?
  • Which timestamp is reliable for each record: system creation time, last-modified time or export time?
  • Can provenance be delivered per item: source system, creation time and the de-identification method applied?
  • Can the license limit use to evaluation, prohibit redistribution and state what you may publish about results? See evaluation-only data license terms.
  • Can later time windows be licensed for refresh?
  • Does the supplier or any intermediary also produce training data for the capability you are testing?

SourceX sources operational datasets from US companies on request; it does not hold them in stock, and a request does not guarantee a matching dataset. It does not source scraped public web content or train AI models. Buyers describe the data; SourceX looks for US companies that hold it, checks the data and the supplier's licensing permissions, and agrees allowed uses in a license that defines which records are included, their permitted uses, the license term and delivery. Diligence materials on source, rights, preparation and allowed use are prepared per dataset for your review, so bring the checklist above to that review, or describe the held-out records you need first.

What design cannot do, and where detection takes over

Design lowers the probability of contamination; it cannot prove its absence. Models can answer new items from public background knowledge, post-cutoff items age, private items can leak through evaluation traffic, and closed models' corpora cannot be inspected from outside. Even LiveBench's ICLR 2025 title calls it "contamination-limited" [15].

So run the overlap checks in contamination checks for licensed evaluation data on every delivered set, decontaminate your own training data against it, and track the public-private score gap over time. If you are still deciding whether you need a private set, start with private evaluation sets vs public benchmarks or the evaluation and benchmarking datasets hub.

Sourcing never-published records for a held-out benchmark?

Describe the records you need, the time window they must fall in, the ground truth they must carry, and the uses and publication rights your benchmark requires. SourceX looks for US companies that hold such records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Describe your evaluation data needs.

Sources

  1. arXiv, "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
  2. Yang et al. (LMSYS), "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples," arXiv (2023). https://arxiv.org/pdf/2311.04850v1
  3. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  4. arXiv, "Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts" (2024). https://arxiv.org/pdf/2410.09247.pdf
  5. Lee et al., "Deduplicating Training Data Makes Language Models Better," arXiv and ACL (2021-2022). https://arxiv.org/abs/2107.06499v1
  6. Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" ICLR (2024). https://arxiv.org/pdf/2310.06770
  7. arXiv, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/html/2509.16941v1
  8. White et al., "LiveBench: A Challenging, Contamination-Limited LLM Benchmark," arXiv and ICLR (2024-2025). https://www.arxiv.org/pdf/2406.19314
  9. ICML, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  10. arXiv, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  11. Northcutt, Athalye and Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks," NeurIPS Datasets and Benchmarks (2021). https://arxiv.org/abs/2103.14749
  12. Jain et al. (MLCommons), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI," arXiv (2024). https://arxiv.org/pdf/2407.16883
  13. NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  14. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance." https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  15. ICLR, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (poster page, 2025). https://iclr.cc/virtual/2025/poster/28134
  16. European Parliament and Council, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)," Official Journal of the EU (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data