Skip to content

Data quality, coverage and contamination

Decontaminating a Licensed Training Set Against Public Benchmarks

Quick answer

To decontaminate training data, fix the list of benchmarks you will report on, normalize both the benchmark items and the training records the same way, match them with short n-grams plus a longest-match score, remove or quarantine every record that hits, and publish a per-benchmark decontamination report tied to the dataset version. Published thresholds range from 8-gram to 13-gram and 50-character overlap [1], but larger n misses more contamination [2][3], and no string method catches paraphrased test items [4][5].

By SourceX Editorial · Updated

This guide covers the training side: cleaning a purchased corpus or SFT set before it touches your model. Keeping a licensed evaluation set out of training is a separate problem, covered in contamination checks for licensed evaluation data. For the broader framework, start at the data quality, coverage and contamination hub.

Why licensed operational data still needs decontamination

Licensed private data can still contain benchmark items, because the people who produced it copied public text into their work. A support team pastes a Stack Overflow answer into a ticket; an engineering wiki quotes a HumanEval-style function; a training team builds internal quizzes from GSM8K or MMLU questions. Any of these can turn a benchmark item into a training example.

The risk is higher for SFT and instruction sets assembled by vendors, which sometimes seed prompts from public benchmarks or generate examples with a model that has memorized them. Generated examples that mirror test sets are covered in contamination through synthetic data. Testing whether licensed records already sit in Common Crawl or other pretraining corpora is a novelty question, handled in testing licensed data against public web crawls.

The cost of skipping the step is credibility. LiveBench's authors note that test data can end up in newer models' training sets and make a benchmark obsolete quickly [9]. As of October 2026, OpenAI has flagged SWE-bench Verified, a 500-sample human-verified subset of SWE-bench, as contaminated and stopped reporting it [8]. If your scores move on a benchmark you never decontaminated against, reviewers will discount them. The benchmark contamination glossary entry defines the terms used below.

What thresholds labs have reported

Published decontamination thresholds vary by an order of magnitude, so treat them as starting points rather than standards. The benchmark data contamination survey summarizes the main reported settings [1]:

Source (as summarized in the survey [1])UnitContamination ruleVerify against
GPT-3Word 13-gramsAny 13-gram shared with a test item flags itGPT-3 paper
GPT-4Characters50-character substring overlap (reported secondhand)GPT-4 technical report
Llama 2TokensToken-level n-gram matching (the survey reports 8-grams); check the paper's minimum match length and skipgram budgetLlama 2 paper [6]
Another study cited in the survey8-gramsItem contaminated if 70% of its 8-grams appear in training dataOriginal study

Read every row against the original report before you copy it into a policy. The survey is secondary, the GPT-4 figure is reported secondhand, and details such as tokenization, skip budgets and whether the rule flags the training record or the test item change the outcome.

The direction of the error matters more than any single number. A 2024 study (ConTAM) comparing four n-gram contamination metrics found longest-match the most effective and found n below 8 more accurate [2]. Larger n-grams and higher thresholds produce false negatives [3]: a 13-gram rule will not catch a math word problem whose numbers and names were edited, when the edits are frequent enough to break every 13-token window.

The decontamination workflow, step by step

A defensible decontamination run has five stages, and each produces something you keep. The stages below are a working practice, not a standard.

1. Freeze the benchmark list. List every benchmark you will report on, plus those your customers or reviewers commonly ask about, with exact split, version and hash. Include dev and validation splits if you use them for model selection. Record where each copy came from, because benchmark mirrors on the Hugging Face Hub often differ in whitespace, answer formatting and item count.

2. Normalize both sides identically. Apply Unicode NFKC, lowercase, collapse whitespace, strip punctuation and markup, and decide how to treat numbers and code. For code benchmarks, also normalize indentation and strip comments, or the same function will not match across a tab-indented repository and a space-indented benchmark. Run the same normalizer on training records and benchmark items; mismatched pipelines are the most common reason a run reports zero hits.

3. Match at two granularities. Build an index of benchmark n-grams (for example 8-grams and 13-grams over a fixed tokenizer) and scan the training set against it. For each training record, also compute the longest common substring with its nearest benchmark item, since longest-match catches partial copies that a single fixed n can miss [2]. Suffix arrays, the structure used for exact-substring deduplication of training corpora [7], scale this to large sets.

4. Remove or quarantine. Decide the unit of removal before you see the results: drop the whole record, drop the matched span, or move the record to a quarantine table for review. Whole-record removal is the safer default for SFT pairs, where a contaminated prompt with a clean answer still teaches the test item. For long documents, span removal keeps the rest of a support thread or contract usable.

5. Report and version. Write the decontamination report (artifact below), store it with the dataset version and reference it in the model's documentation. Re-run the whole pipeline whenever the benchmark list or the dataset version changes.

Catching paraphrased and translated test items

String matching cannot detect a benchmark question that was rewritten, translated or reformatted, so add a semantic pass for high-stakes benchmarks. The LMSYS team showed that rephrased test samples pass standard n-gram decontamination and can inflate a model's benchmark score substantially [4][5].

Their proposed "LLM decontaminator" works in two steps: retrieve the top-k most similar training records for each test item with embeddings, then ask a strong model whether each pair is the same question rephrased [4]. That design suits operational data, where a support agent may restate a public troubleshooting answer in their own words. Budget for it on the benchmarks you will headline, not every benchmark, because the judge step costs far more than n-gram matching.

Near-duplicate methods help with lightly edited copies. MinHash with locality-sensitive hashing over shingles, explained in exact and near-duplicate detection with MinHash and LSH, catches items that differ by a few words, and it can reuse the infrastructure you already run for deduplication [7].

Common failure modes in decontamination runs

Most decontamination failures come from process gaps, not from the choice of n. Check for these before you sign off:

  • Benchmark answers left out of the index. Indexing only questions misses training records that contain the answer key, such as a ticket that pastes a solved GSM8K solution.
  • Templates flagged as contamination. Multiple-choice scaffolding ("Which of the following...", "A) B) C) D)") and license headers create false positives; exclude boilerplate n-grams that appear in many unrelated items.
  • Short items silently skipped. A 13-gram rule cannot fire on a question shorter than 13 tokens; log how many benchmark items were too short to test.
  • Different tokenizers. Counting 8-grams with one tokenizer on the benchmark and another on the corpus makes thresholds meaningless.
  • Multilingual leakage. Translated benchmark items evade English n-grams entirely; only the semantic pass catches them [5].
  • Report without version. A report that does not name the dataset hash and benchmark versions cannot support a later score.

Artifact: decontamination report record

The report should let a reviewer reproduce the run and see hits per benchmark. Keep one record per benchmark per dataset version.

Illustrative example: invented to show structure; it does not describe an available dataset.

decontamination_report:
  dataset_id: support-sft-v3
  dataset_version_sha256: "9f2c...e41a"
  run_date: 2026-10-09
  normalizer: { unicode: NFKC, lowercase: true, strip_punct: true, code_comments: stripped }
  tokenizer: "model-tokenizer-v2"
  methods:
    - { type: ngram, n: 8, rule: "any shared 8-gram" }
    - { type: ngram, n: 13, rule: "any shared 13-gram" }
    - { type: longest_match, min_chars: 50 }
    - { type: semantic, retriever: "embedding top-5", judge: "LLM same-question check", scope: [gsm8k] }
  benchmarks:
    - name: gsm8k
      split: test
      version_hash: "b81d...07c3"
      items_indexed: 1319
      items_too_short_for_rule: 0
      training_records_flagged: 42
      flagged_by: { ngram_8: 31, ngram_13: 12, longest_match: 35, semantic: 9 }
      action: removed_whole_record
      quarantine_table: "quarantine.gsm8k_v3"
  false_positive_review: { sampled: 40, confirmed_contaminated: 33, template_only: 7 }
  reviewer: "eval-lead"

Fields such as items_too_short_for_rule and the false-positive sample make the limits of the run visible instead of implied. Dataset documentation formats give the report a home: Hugging Face recommends a dataset card for every dataset [11], and Croissant-RAI provides machine-readable fields for data life cycle documentation [10]. If you also prepare an EU training content summary, the same record supports it; see completing the EU training content summary for licensed datasets.

What to ask a data supplier before delivery

Decontamination is your job as the model builder, but supplier answers decide how much work it takes. Ask these questions while scoping a licensed corpus or SFT set:

  • Were any records written from, seeded with, or checked against public benchmarks or exam banks?
  • Were any records generated or rewritten by a model, and which one?
  • Can the supplier share record-level timestamps and source system fields (ticket IDs, repository paths) so you can trace a flagged hit?
  • Does the license allow you to remove records and keep a quarantine table for audit?

Pair the answers with a held-out evaluation strategy. Public benchmarks remain useful for comparison, but private evaluation sets and contamination-resistant evaluation design give you scores that do not depend on decontamination being perfect.

SourceX sources operational datasets, such as support and sales histories, engineering records and documents, from US companies on request, and prepares diligence materials covering source, rights, preparation and allowed use for each dataset. Buyers can describe the training data they need, knowing that a request does not guarantee a match.

Sourcing training data you can decontaminate and document

SourceX does not hold data in stock; it looks for US businesses that hold the data you describe, and each supplying company approves every release. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, so your decontamination report has a clear version to attach to. Describe your training data requirements to SourceX.

Sources

  1. arXiv, "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
  2. Maxim AI (Epochs), "ConTAM: Tackling Data Contamination in LLM Benchmarks". https://epochs.getmaxim.ai/contam-tackling-data-contamination-in-llm-benchmarks-ccceadf10d5f
  3. Maxim AI, "Evaluating data contamination in LLMs". https://getmaxim.ai/blog/llm-data-quality
  4. LMSYS Org, "Catch me if you can! How to beat GPT-4 with a 13B model (LLM Decontaminator)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  5. arXiv, "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples" (2023). https://arxiv.org/pdf/2311.04850v1
  6. arXiv (Meta), "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
  7. arXiv, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
  8. OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
  9. ICLR 2025, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2025). https://iclr.cc/virtual/2025/poster/28134
  10. arXiv (MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  11. Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data