Skip to content

Data quality, coverage and contamination

Contamination Through Synthetic Data: When Generated Training Examples Mirror Test Sets

Quick answer

Synthetic data contamination happens when a teacher model, prompted to generate training examples, reproduces benchmark items it absorbed during its own training, usually as paraphrases rather than copies. Exact-match and n-gram filters often miss these paraphrases, so scores on HumanEval, GSM8K or MMLU can rise without real capability gains [1][2]. Treat every generated set as untrusted: run semantic decontamination against your full evaluation suite, keep generation prompts and seeds, and validate on held-out data the teacher never saw.

By SourceX Editorial · Updated

How a teacher model leaks benchmark items into generated data

A teacher leaks benchmark content because it was trained on the open web, where popular test sets, their solutions and discussion threads circulate freely. When you ask it for "a Python function problem with tests" or "a grade-school word problem," the most probable outputs sit close to the canonical examples it memorized. Our working hypothesis, consistent with the evidence below, is that leakage is strongest where a benchmark has a narrow, templated format: code docstrings with unit tests, short arithmetic word problems, and four-option multiple-choice items.

Yang et al. showed the downstream effect directly: a 13B model fine-tuned on rephrased test items reached benchmark scores far above its real ability, and standard n-gram decontamination did not catch the rephrasings [1]. The same paper reports synthetic code instruction data, generated by a large model, containing items semantically close to HumanEval problems that overlap checks failed to flag [1]. That is the defining shape of synthetic leakage: same task, same test logic, new variable names and wording.

Three generation patterns raise the risk:

  • Few-shot seeding from public benchmarks. Using GSM8K or MBPP items as in-context examples invites the teacher to produce near-siblings of other items from the same set.
  • Rejection sampling against benchmark-style verifiers. Filtering generations by "passes these unit tests" selects for outputs that match the benchmark's solution space.
  • Self-instruct loops. Each round seeds the next, so one leaked item can multiply into many paraphrased variants.

Why n-gram decontamination misses synthetic paraphrases

N-gram filters fail on synthetic data because they test surface overlap, while a teacher's paraphrase preserves meaning and changes tokens. A 13-gram or 50-character overlap rule, the style used in many pretraining reports, catches copied passages from web crawls well. It does not catch "Write a function that returns the largest prime factor of n" rewritten as "Implement a routine computing the biggest prime dividing an integer," which tests the same thing [1].

Exact matching still earns its place as a cheap first pass. Suffix-array tools such as infini-gram make exact n-gram counts practical even across trillion-token corpora, so you can run it on every generated record without sampling [8]. The survey literature groups methods into matching-based checks, which need access to the training data, and behavior-based probes, which inspect the model, and notes that each misses cases the other catches [4]. For synthetic sets you control the training data, so lean on matching, but make it semantic.

A semantic decontamination workflow for generated SFT data

Semantic decontamination means comparing each generated record to each evaluation item by meaning, then adjudicating the closest pairs. LMSYS's LLM Decontaminator does this in two stages: embedding similarity retrieves the top-k nearest test items for each training example, and a strong LLM judges whether each pair is a rephrasing [2]. The judge step matters because cosine similarity alone flags many legitimate same-topic pairs.

Illustrative example: invented to show structure; it does not describe an available dataset.

StageInputMethodTypical setting to decideOutput
1. NormalizeGenerated records, eval itemsLowercase, strip whitespace and code comments, canonicalize variable names in codeWhich fields to compare (prompt only, or prompt plus target)Comparable text pairs
2. Exact and n-gram passNormalized textHashed n-gram overlap against all eval itemsn (for example 8 to 13 tokens); overlap ratio thresholdRecords with verbatim overlap removed
3. Near-duplicate passSurvivorsMinHash with LSH bandingJaccard threshold; shingle sizeTemplated near-copies removed
4. Embedding retrievalSurvivorsSentence or code embeddings; top-k nearest eval itemsk (for example 1 to 5); cosine floor for reviewCandidate pairs
5. LLM adjudicationCandidate pairsJudge prompt: "Is A a rephrasing of B or does it test the same answer?"Judge model must differ from the teacherRephrase verdicts
6. Human spot checkRandom sample of keeps and dropsReviewer labelsSample size; disagreement threshold that triggers re-tuningMeasured precision and recall of the filter
7. LogAll decisionsWrite verdicts with record IDsRetention period for logsAuditable decontamination report

Two settings decide most outcomes. First, compare targets as well as prompts: a generated math problem with fresh wording but the same numeric answer and solution path as a GSM8K item is still contaminated. Second, decontaminate against every benchmark you will report, plus your private evaluation sets, not only the headline ones. The Tulu 3 recipe is a useful public reference: it checks its post-training prompt sets, including synthetic ones, for overlap with its evaluation suite before training [5]. For the near-duplicate stage, see MinHash and LSH near-duplicate detection; for the general licensed-set procedure, see decontaminating a licensed training set against public benchmarks.

Probing the teacher before you generate at scale

Probing the teacher first tells you which benchmarks it has likely memorized, so you can decide where to tighten filters or switch generators. Deng et al. describe a guessing probe: mask a part of a benchmark item, such as one wrong option in a multiple-choice question, and ask the model to fill it in; exact recoveries the model could not infer from context suggest memorization [3]. Run this on a few hundred items per benchmark before committing compute to millions of generations.

Practical probes for a post-training team:

  • Masked-option recovery on multiple-choice sets, per the guessing approach [3].
  • Continuation tests: give the first half of a HumanEval docstring and measure how often the teacher completes it verbatim.
  • Canary checks: search generated output for benchmark canary strings or dataset-specific identifiers.
  • Answer-first prompts: ask the teacher to "write a problem whose answer is X" using benchmark answers, and check whether it returns the benchmark question.

A teacher that passes these probes can still leak, so probes reduce risk rather than certify a generator.

Generation records that make contamination traceable

Traceable generation records let you remove a whole contaminated family instead of one record at a time. Our recommendation, a hypothesis grounded in practice rather than a published standard, is to store the teacher model and version, the full prompt template, any in-context seed examples, sampling parameters and the random seed for every record. When adjudication flags one item, you can find every sibling produced from the same seed example or template and inspect them together. The provenance records page for synthetic training data covers the wider field set, including generator terms.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "syn-math-000481",
  "teacher_model": "teacher-llm-v3 (2026-06 snapshot)",
  "prompt_template_id": "math-wordproblem-v7",
  "seed_examples": ["internal-seed-0042", "internal-seed-0107"],
  "seed_source": "internal, not from any public benchmark",
  "sampling": {"temperature": 0.9, "top_p": 0.95, "seed": 771203},
  "verifier": "sympy-answer-check-v2",
  "decontam": {
    "ngram_13_max_overlap": 0.0,
    "minhash_max_jaccard": 0.21,
    "nearest_eval_item": "gsm8k-test-0913",
    "cosine": 0.83,
    "judge_verdict": "not_rephrase",
    "judge_model": "judge-llm-v2"
  }
}

Note seed_source: if seeds came from a public benchmark's train split, flag the whole family for closer review, because train and test splits of popular sets often share templates.

Measuring whether contamination already inflated your results

The reliable check for inflation is a gap between public benchmark scores and fresh or private evaluation data that no teacher could have seen. LiveBench was built around this problem: its authors note that test data can end up in newer models' training sets and make a benchmark obsolete, so they refresh questions from recent sources with objective answers [6]. OpenAI stopped reporting SWE-bench Verified in 2026 after concluding that training-time exposure increasingly drove score gains [7].

Signals that point to synthetic leakage in a trained checkpoint:

  • A large jump on one public benchmark with no matching gain on a held-out or freshly written set for the same skill.
  • Gains concentrated on the oldest, most widely discussed benchmark items.
  • An ablation where removing the synthetic slice erases most of the gain on that benchmark but little elsewhere.

Fresh evaluation items work best when they come from sources the teacher never saw, such as recent or non-public operational records. A licensed real-data holdout gives you that reference, and contamination-resistant evaluation design covers how to keep it clean over time. If you also use generated test sets, read where synthetic evaluation data misleads first.

Where licensed real data fits in a mixed post-training set

Licensed real records reduce, but do not remove, contamination risk in a mixed set. Operational data such as support tickets, engineering issue histories or finance workflows was not written to resemble public benchmarks, so it rarely paraphrases test items, and it can anchor evaluation sets that a teacher model has not memorized. It still needs the same decontamination pass, because real documents sometimes quote public material. Our guide to combining licensed and synthetic data covers mix ratios and roles, and the synthetic data glossary entry defines the terms used here.

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows; nothing is held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. If you need non-public real data for a holdout or a training mix, describe the data you need. For quality checks beyond contamination, start at the data quality hub or the synthetic data quality assessment guide.

Sourcing real data to validate synthetic training sets

SourceX looks for US businesses that hold the operational data you describe, and every release is approved by the supplying company. Personal details are removed or replaced before delivery, and delivery runs through private, access-controlled workflows only after an executed agreement. To describe a dataset for a holdout or training mix, go to SourceX for AI data buyers.

Sources

  1. arXiv (Yang et al.), "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples" (2023). https://arxiv.org/pdf/2311.04850v1
  2. LMSYS Org, "Catch me if you can! How to beat GPT-4 with a 13B model (LLM Decontaminator)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  3. arXiv (Deng et al.), "Investigating Data Contamination in Modern Benchmarks for Large Language Models" (2023). https://arxiv.org/html/2311.09783v2
  4. arXiv (Xu et al.), "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
  5. arXiv (Lambert et al., Allen Institute for AI), "Tulu 3: Pushing Frontiers in Open Language Model Post-Training" (2024). https://arxiv.org/pdf/2411.15124
  6. arXiv (White et al.); ICLR 2025, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
  7. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  8. arXiv (Liu et al.), "Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens" (2024). https://arxiv.org/html/2401.17377v4

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data