Code and software engineering data
Synthetic Code Data vs Licensed Real Code: Output Terms, Contamination and Coverage
Quick answer
Synthetic code data works for volume: instruction pairs, unit-test-verified exercises and RL tasks in popular languages where a compiler or test suite can check every sample. Licensed real code is needed where generators are weak or risky: legacy and internal stacks, long multi-file changes, real review judgment and held-out evaluation. Before training on model-generated code, check the generating model's output terms, decontaminate synthetic sets against your benchmarks, and record which records are synthetic.
By SourceX Editorial · Updated
This guide is for teams planning a post-training mix for code models. It sits under the code and software engineering data hub, and the general trade-offs across data types are covered in licensed vs synthetic vs scraped AI training data. Here the focus is what is specific to code: distillation terms, paraphrased benchmark leakage and the gap between generated tasks and enterprise repositories.
Where synthetic code data is enough
Synthetic code is enough when correctness is machine-checkable and the target distribution looks like public code. Execution feedback is the reason: a generated function paired with generated tests can be run in a sandbox, and failures dropped or turned into negative examples. That makes synthetic data strong for the following uses.
- SFT instruction pairs in mainstream languages. Python, TypeScript, Java and Go tasks with docstrings, signatures and passing tests.
- RL task generation with verifiable rewards. Problems where a hidden test suite, a type checker or a linter returns a pass or fail signal.
- Format and tool-use drills. Diff formats, patch application, JSON tool calls and terminal command sequences, where structure matters more than domain knowledge.
- Targeted augmentation. Variations of a known seed task (renamed identifiers, changed constraints, added edge cases) to widen coverage around a skill.
The weakness is the distribution itself. A model trained repeatedly on generated data loses the tails of the original distribution, a failure described as model collapse [7]. In code, the tails are exactly the unusual build systems, older idioms and house conventions that enterprise users bring to a coding assistant.
Where licensed real code is still needed
Licensed real code is needed wherever the generator has no good prior or the task depends on context it never saw. These gaps are working hypotheses from practice, so test them against your own evals rather than treating them as settled.
| Gap | Why synthetic falls short | What real data supplies |
|---|---|---|
| Legacy languages (COBOL, PL/I, RPG, older Fortran) | Few public examples, so generated code drifts toward textbook style | Copybooks, JCL, CICS and DB2 usage as written in production; see COBOL and mainframe code datasets |
| Internal frameworks and monorepos | The generator cannot invent APIs it has never seen | Real cross-file dependencies and build graphs; see repository-level code context |
| Long-horizon multi-file changes | Generated tasks tend to be short and self-contained | Commit sequences, migrations and refactors across many files |
| Review judgment | A generator rates its own output by its own taste | Accepted and rejected changes with reviewer rationale; see code review preference data |
| Held-out evaluation | A generator may reproduce public benchmark items | Code never published, so it cannot be in pretraining corpora |
Long tasks are where the gap shows most. METR measures model capability as the length of software task, in human working time, that a model completes with 50% success [8]. Tasks that take a skilled engineer hours rarely come out of a single generation prompt, and real repositories with history are the natural source of them.
Provider output terms on model-generated code
Before you train on code generated by another model, read that provider's output terms, because some restrict training competing models. As of October 2026, OpenAI's terms of use bar using Output to develop models that compete with OpenAI [2]; business and API terms may offer different permissions for specific training use cases. Anthropic's help center addresses whether customers can use outputs to train an AI model and ties the answer to its restrictions on competing products [1]. Open-weight models carry their own terms: archived Gemma Terms of Use define Model Derivatives to cover models trained to behave like Gemma through patterns transferred from its outputs, which reaches distillation and synthetic data methods [3].
How enforceable these restrictions are is debated. Lemley and Henderson argue that terms-of-use limits on AI outputs may rest on weak legal footing, which SpicyIP discusses [4]. Treat output terms as a contractual risk to review with counsel, not as a settled rule either way; the provenance angle is covered in more depth in checking provider output terms before training.
Three practical points follow for code teams:
- Terms attach to the account and version that generated the data. Record the provider, model ID, API or consumer product, and the terms version date for every generation run.
- Vendor synthetic datasets inherit the same question. If a vendor sells generated coding tasks, ask which models produced them and under which terms.
- "Open" datasets can hide distilled content. Dataset cards often omit or misstate licenses; one audit found license omission above 70% and error rates above 50% on popular hosting sites [9]. Verify before relying on a card.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Synthetic code and benchmark contamination
Synthetic code data can leak benchmark items in paraphrased form, so it must be decontaminated like scraped data. LMSYS showed that rephrased test samples, such as a HumanEval problem with renamed variables or translated to another language, pass n-gram overlap checks while still inflating scores; their LLM-based decontaminator found such overlap in widely used synthetic datasets [5]. A generator that memorized a benchmark will reproduce it with small variations when asked for "similar problems".
Public repositories create the same exposure for real code. The SWE-Bench Pro authors argue that permissively licensed public repositories are prime candidates for pretraining corpora, so benchmarks built from them are at risk; they drew on copyleft repositories and private commercial codebases instead, with public, held-out and commercial subsets across 41 repositories [6]. The lesson for buyers is that private, licensed code is the cleanest material for held-out evaluation. Detection methods for forks and mirrors are covered in code benchmark contamination, and limits of generated test sets in synthetic evaluation data limits.
A minimum decontamination pass for a synthetic code set:
- Exact and near-duplicate matching (MinHash or suffix-array) against every benchmark you report, plus their known forks.
- Embedding similarity to surface candidates, then an LLM judge or human check for semantic paraphrase.
- Identifier-normalized comparison (rename variables, strip comments) before matching.
- Cross-language check where your generator translates tasks between languages.
- Dedup against your training pool; see deduplicating code training data.
A hybrid mix: real seeds, synthetic volume
The most defensible pattern, offered here as a hypothesis to test, is licensed real code as seeds and held-out evaluation, with synthetic variation for volume. Real repositories provide the hard tasks and the eval; generators widen coverage around them; and every record carries a flag saying which is which. The general approach is described in combining licensed and synthetic data, and the record-keeping side in provenance records for synthetic training data.
Keep the eval split strictly real and never let seeds from the eval repositories feed the generator. Otherwise the synthetic set paraphrases your own held-out tasks, and you recreate the contamination problem in-house.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_id": "task-000417",
"origin": "synthetic",
"seed_record_id": "real-commit-88213",
"seed_split": "train",
"generator": {"provider": "example-provider", "model_id": "example-model-2026-05", "terms_version": "2026-03-01", "output_terms_reviewed": true},
"language": "java",
"task_type": "multi_file_refactor",
"files_touched": 4,
"verification": {"build": "maven", "tests_passed": 37, "tests_failed": 0},
"decontamination": {"minhash_max_jaccard": 0.21, "llm_judge_paraphrase_flag": false, "benchmarks_checked": ["HumanEval", "MBPP", "SWE-bench"]},
"license_basis": "derived from licensed seed; seed license permits derivative training data"
}
Illustrative example: invented to show structure; it does not describe an available dataset.
| Decision question | If yes | If no |
|---|---|---|
| Can a test suite or compiler verify every sample? | Synthetic is viable at scale | Prefer real data with human outcomes |
| Is the target language or framework well represented publicly? | Synthetic is viable | License real code |
| Do the generating model's output terms permit this training use? | Proceed, record the terms version | Switch generator or use real data |
| Will this data feed an eval you report? | Use private real code only | Synthetic acceptable after decontamination |
| Does the task span many files or long history? | License real repositories with history | Synthetic is viable |
Specifying the real-code side of the mix
Specify licensed real code by the gaps your evals expose, not by volume. Name the languages and versions, frameworks, build tools, whether full Git history and review threads are needed, and whether tests must run. The code dataset request specification lists fields to include, and open code datasets vs licensed private code explains what public corpora already cover.
SourceX sources operational datasets, including engineering records, from US companies on request; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Buyers describe the data they need on the SourceX buyer page, and SourceX looks for US businesses that hold it. For the broader comparison of generated and business data, see synthetic data vs real business data.
Request licensed real code for your code model
SourceX sources operational datasets from US companies, including engineering records, and manages the licensing process from Find and Assess through Agree, Transact and Manage. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Describe the real code your synthetic mix cannot cover at https://sourcex.si/buyers.
Frequently asked questions
Can we train on code generated by an open-weight model without restrictions?
Not automatically. Open-weight licenses can carry use restrictions and derivative definitions that reach models trained on their outputs, as the archived Gemma terms show [3]. Read the specific license version that applied when you generated the data.
Is n-gram decontamination enough for synthetic code?
No. Paraphrased and translated benchmark items pass n-gram checks, which is why LMSYS proposed an LLM-based decontaminator [5]. Combine lexical, embedding and judge-based checks.
How much real code does a hybrid mix need?
There is no fixed ratio. Size the real portion by the eval gaps it must close and by the held-out set you need, then measure whether synthetic additions improve scores on the real held-out set.
Sources
- Anthropic (Claude Help Center), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- OpenAI, "Terms of Use". https://openai.com/policies/row-terms-of-use/
- Google AI for Developers, "Gemma Terms of Use (archived version)" (2024). https://ai.google.dev/gemma/terms-archive/terms_04_01_24
- SpicyIP, "Discussing Lemley and Henderson's \"The Mirage of Artificial Intelligence Terms of Use Restrictions\"" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
- LMSYS Org, "Catch me if you can! How to beat GPT-4 with a 13B model (LLM Decontaminator)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
- Scale AI (arXiv:2509.16941), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/abs/2509.16941v1
- Nature (Shumailov et al., via RePEc), "AI models collapse when trained on recursively generated data" (2024). https://ideas.repec.org/a/nat/nature/v631y2024i8022d10.1038_s41586-024-07566-y.html
- METR (arXiv:2503.14499), "Measuring AI Ability to Complete Long Software Tasks" (2025). https://arxiv.org/html/2503.14499v3
- Longpre et al. (arXiv:2310.16787), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.