Skip to content

Code and software engineering data

Code Benchmark Contamination: Detecting Leaks from Public Repos, Forks and Mirrors

Quick answer

Code benchmark contamination happens when benchmark problems, their reference solutions or their source repositories reach a model's training data, so scores measure recall instead of skill. Code leaks through more paths than prose: forks and mirrors of benchmark repos, vendored copies, solution write-ups and synthetic rephrasings. Catch it with layered checks: normalized token n-grams, AST or MinHash near-duplicate matching, embedding plus LLM review for rewrites, and a pre- versus post-cutoff performance comparison.

By SourceX Editorial · Updated

Why coding benchmarks are structurally exposed

Coding benchmarks built from public GitHub repositories are exposed by design, because the same repositories are prime candidates for the web-crawled corpora used in pretraining [4]. SWE-bench draws 2,294 problems from real issues and pull requests in 12 popular Python repositories, and the fix for each task is a merged commit that has sat in public history since before the benchmark existed [7]. Every crawl that ingests those repositories, their forks or their issue threads ingests the answers too.

The consequence is now visible in public reporting. As of October 2026, OpenAI states that SWE-bench Verified is increasingly contaminated, that score gains increasingly reflect training-time exposure, and that it has stopped reporting the benchmark [6]. Function-level sets such as HumanEval face the same problem in a different shape: short prompts and canonical solutions are copied into tutorials, course repositories and answer sites, where they are hard to trace back to the benchmark.

For a glossary-level definition see benchmark contamination. This page covers the code-specific leak paths and detectors.

Code-specific leak paths to model in your scan

Code leaks through repository mechanics that prose benchmarks rarely face, so a decontamination scan has to target each path explicitly. The paths below are a working taxonomy for scoping a scan, not findings about any specific benchmark.

  • Forks and mirrors. A GitHub fork, a GitLab or Gitee mirror, or a Software Heritage snapshot carries the full history, including the fix commit, under a different repository URL. Repository-name blocklists miss them.
  • Vendored copies. Projects copy a dependency's source into vendor/, third_party/ or site-packages/, so a benchmark repo's post-fix file can appear inside an unrelated project.
  • Issue and PR text. Issue bodies, review threads and commit messages describe the fix in natural language, which leaks the reasoning even when the diff is filtered.
  • Solution write-ups. Blog posts, notebooks and answer sites restate HumanEval-style problems with solutions, often with renamed functions and reformatted docstrings.
  • Synthetic rephrasings. Instruction data generated from benchmark prompts, or from a teacher model that saw them, produces paraphrased problems that no exact match will find [2].
  • Translations. A Python task reimplemented in JavaScript or Go is still the same task, and cross-language copies defeat token matching entirely [2].

Forks and vendored copies are also the core problem in deduplicating code training data, and the same index can serve both jobs.

N-gram overlap: thresholds, trade-offs and code normalization

N-gram overlap is the standard first-pass detector, but its thresholds were tuned for prose and need code-specific normalization before they mean anything [1]. The survey literature reports wide variation: GPT-3 work treated a 13-gram match as a contamination signal, GPT-4 reportedly used a 50-character overlap, and another study flagged an item when 70% of its 8-grams matched [1]. Short n-grams over-flag code because boilerplate (import numpy as np, def __init__(self):, license headers) repeats everywhere; long n-grams under-flag because renaming one variable breaks every window that contains it.

Normalize before you match. A workable recipe for code:

  1. Parse with a real grammar (for example tree-sitter) rather than regex, so strings and comments are handled correctly.
  2. Strip comments and docstrings into a separate text channel and check them against the problem statement on their own.
  3. Collapse whitespace and formatting, or run a canonical formatter such as Black or gofmt first.
  4. Alpha-rename local identifiers to positional placeholders (v0, v1, f0) while keeping API and library names, which carry the signal.
  5. Build token n-grams over the normalized stream, and drop n-grams that occur above a frequency ceiling across the whole corpus so boilerplate cannot trigger matches.

Then add structure-aware matching. Hashing normalized AST subtrees catches reordered statements, and MinHash with locality-sensitive hashing over normalized shingles finds near-duplicate files and functions at corpus scale, the same family of methods used to measure train-test overlap in deduplication work [9]. Report coverage per item, the share of a benchmark item's normalized n-grams found in the corpus, rather than a single yes or no flag.

Catching rewrites: embeddings, LLM judges and canaries

Paraphrased, translated or identifier-renamed benchmark items evade string and hash matching, so the last layer has to compare meaning. LMSYS showed that rephrased test samples slip past n-gram and embedding checks alike, and proposed a two-stage decontaminator: embedding similarity retrieves the top candidates for each test item, then a strong LLM judges whether each pair is the same problem [2]. For code, use a code-specific embedding model, embed both the problem statement and the reference solution, and send the top-k neighbors per item to the judge with the instruction to ignore naming and formatting.

Canary strings are a cheap complement. BIG-bench put a canary GUID in every task file so corpus builders could filter it, and a model that can complete the canary has likely seen the files [8]. Few coding benchmarks ship canaries, so if you build an internal eval, embed one in every file, README and issue template, and grep both your training mix and model outputs for it.

The post-cutoff test for models you did not train

When you cannot inspect the training data, compare a model's accuracy on problems published before its training cutoff with problems published after it. A sharp drop on post-cutoff problems of similar difficulty points to contamination rather than a capability gap. This is the logic behind refreshing benchmarks such as LiveBench, which draws questions from recent sources because static test data ends up in newer models' training sets [3].

Run the test carefully. Bucket problems by first-public date (issue creation, contest date or commit date, not benchmark release date), control for difficulty with a held-constant rubric or contest rating, and use enough items per bucket for the confidence intervals to separate. A flat curve does not prove cleanliness; it only fails to show a leak.

Decontamination run sheet

The artifact below is a run sheet an evaluation lead can adapt to scope a decontamination pass over a code training mix or to review a supplier's report.

Illustrative example: invented to show structure; it does not describe an available dataset.

Leak pathDetectorNormalizationFlag rule (tune locally)Action on hit
Fork or mirror of benchmark repoCommit SHA match plus MinHash on filesFormatter, comment stripShared root commit or fix-commit SHA presentDrop entire repository
Vendored copyPath heuristics (vendor/, third_party/) plus file MinHashFormatter, alpha-renameFile-level Jaccard above chosen threshold vs. benchmark post-fix fileDrop file
Solution write-upToken n-gram coverageAlpha-rename, boilerplate ceilingItem coverage above chosen thresholdDrop document
Issue or PR textText n-gram plus embeddingLowercase, strip markdownProblem statement overlapDrop thread
Synthetic rephrasingCode embedding top-k plus LLM judgeEmbed statement and solution separatelyJudge says same problemDrop sample, audit generator prompts
TranslationCross-language embedding plus LLM judgeNoneJudge says same problemDrop sample

A matching per-item record keeps the result auditable:

{
  "benchmark": "internal-swe-eval-v3",
  "item_id": "task-0417",
  "first_public_date": null,
  "ngram_n": 10,
  "ngram_coverage": 0.04,
  "max_file_jaccard": 0.12,
  "fix_commit_sha_found": false,
  "embedding_top1_cosine": 0.71,
  "llm_judge_verdict": "different_problem",
  "canary_found": false,
  "decision": "keep"
}

Log the corpus snapshot hash, tokenizer, normalizer version and thresholds with every run, so a later reviewer can reproduce the decision.

Contamination risk in licensed and held-out eval sets

Private, held-out and commercial evaluation sets exist because public-repo benchmarks cannot stay clean. SWE-Bench Pro has 1,865 problems from 41 repositories and splits them into public, held-out and commercial subsets, using copyleft code for the public and held-out portions to discourage inclusion in training corpora [4]. Scale reports 731 public and 858 held-out instances [5]. Copyleft is a deterrent, not a guarantee, and the copyleft contamination page covers the opposite risk of GPL code entering your own mix.

If you license an evaluation set built from private repositories, contamination becomes a diligence question. Ask the supplier to confirm:

  • the repositories were never public, including archived, temporarily public or open-sourced-then-closed periods;
  • no known public forks, mirrors, package-registry uploads or vendored copies exist, and how that was checked;
  • whether issue text, commit messages or code were ever pasted into public forums, public CI logs or third-party AI tools;
  • the first-commit and last-commit dates, so you can run your own post-cutoff analysis;
  • whether a canary string was embedded and where.

Then run your own scan of the delivered items against your training mix, using the methods above. The held-out coding agent evaluation sets page covers how such sets are built, and private eval sets vs public benchmarks covers when they are worth paying for. For the cross-data-type version of these checks, see contamination checks for licensed evaluation data.

Where SourceX fits in a clean code-eval pipeline

SourceX sources operational datasets, including engineering records, from US companies on request and manages the licensing process; it holds no stock, and a request does not guarantee a match. Every release is approved by the supplying company, every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. SourceX does not train models and does not source scraped web content. Buyers planning coding agent training data from private repositories can describe the repositories they need through the SourceX buyer request, and should run the contamination checks on this page on anything they receive.

For broader context, the code and software engineering data hub maps the cluster, and checking overlap between a new dataset and data you already own applies the same matching to non-benchmark overlap.

Request private code data for contamination-resistant evals

If your evaluations need code that has never sat in a public repository, describe the data rather than the businesses: languages, history depth, issue and test coverage, and intended evaluation use. SourceX looks for US businesses that hold it, assesses data and licensing permissions, and nothing is contracted until a supplier agrees. Describe your code evaluation data request.

Sources

  1. arXiv, "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
  2. LMSYS, "Catch me if you can! How to beat GPT-4 with a 13B model (LLM Decontaminator)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  3. arXiv (ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
  4. arXiv (Scale AI), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
  5. Scale AI, "SWE-Bench Pro" (2025). https://scale.com/blog/swe-bench-pro
  6. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  7. arXiv (ICLR 2024), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
  8. arXiv (BIG-bench authors), "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models" (2022). https://arxiv.org/pdf/2206.04615
  9. arXiv, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data