Code and software engineering data
Preventing Verbatim Reproduction of Licensed Proprietary Code
Quick answer
To stop a code model from reproducing licensed training code, combine four controls: deduplicate the licensed corpus before training, strip secrets and distinctive identifiers, measure extraction with prefix prompts and supplier-planted canaries, and deploy an output filter that blocks completions matching long spans of an index built from the licensed files. Memorization cannot be driven to zero, so suppliers also expect contract terms that define verbatim reproduction, the threshold that triggers a breach, and what testing evidence you will share.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why code models memorize licensed files
Code models memorize most where a sequence is repeated, long and distinctive, and proprietary codebases are full of all three. Lee et al. found that common training sets contain many near-duplicates and long repeated substrings, and that over 1% of unprompted output from models trained on them was copied verbatim from training data [1]. Yang et al. studied code models specifically and reported that memorization grows with model size, output length and how often a snippet appears in training [3]. Carlini et al. showed the attack side: by sampling and then continuing generation from memorized text, they extracted verbatim training data that included code from a GitHub repository and full license texts [2].
Licensed enterprise code amplifies these factors in predictable ways. A supplier's monorepo typically carries vendored copies of internal libraries, generated clients (protobuf, OpenAPI, Thrift stubs), copy-pasted service templates and the same license header on thousands of files. Full git history multiplies each file by every commit that touched it, so an unchanged utility module can appear hundreds of times if you naively export every revision. See delivering repositories with full history for how history arrives and why it needs careful flattening.
Fine-tuning is the higher-risk phase for most buyers. A small, domain-specific corpus seen for several epochs gets far more weight per example than the same files diluted in a multi-trillion-token pre-training mix, and research on fine-tuned models reports that fine-tuning can make sensitive training information, such as personally identifiable information (PII), easier to recover [5].
What counts as reproduction: verbatim, near-verbatim and functional
Define reproduction before you measure it, because suppliers and model teams often mean different things. Verbatim reproduction is an exact token span, usually measured after whitespace normalization. Near-verbatim covers renamed identifiers, reordered imports or stripped comments; a token-level match on a normalized form (identifiers canonicalized, comments removed) catches most of it.
Functional reproduction is the hardest case and matters most for trade secrets. Recent work argues that verbatim metrics miss functionally equivalent code that looks textually different, and proposes comparing a model trained on the target corpus against a reference model that was not, using both textual and execution-based similarity [4]. For a supplier whose value sits in a pricing algorithm or a proprietary scheduler, a paraphrased reimplementation can be as damaging as a copy.
The legal framing differs by harm. Copyright concerns turn on copied expression, while trade-secret concerns turn on disclosure of the secret itself, whatever form it takes. Our trade secret law overview and the supplier-facing answer in does licensing data expose trade secrets cover the legal side; this page covers the engineering evidence.
Pre-training controls: deduplicate, flatten history, strip secrets
The cheapest reduction in memorization happens before the first gradient step. Deduplication is the best-supported control: Lee et al. found deduplicated training produced models that emitted memorized text far less often [1], and Yang et al. recommend it for code models [3].
Apply it at three levels for licensed code. Exact dedup removes identical files by content hash (SHA-256 of normalized bytes). Near-dedup with MinHash and locality-sensitive hashing over token shingles catches forks, vendored copies and lightly edited templates. Substring dedup, such as suffix-array matching of long repeated spans, catches repeated license headers, boilerplate and generated code inside otherwise distinct files. Our guide to deduplicating code training data covers thresholds and tooling.
Then remove what should never be memorized at all. Hard-coded credentials, private keys, internal hostnames, customer names in fixtures and comments that name people are the worst case, because a memorized secret is a direct breach rather than a copyright argument. Scan the full git history, not only the head commit; see secrets in code datasets. NIST SP 800-218A adds AI-specific secure-development practices covering the integrity of training and fine-tuning data, which gives security reviewers a framework to map these steps to [10].
Training-time options exist but cost quality. Differentially private fine-tuning places a formal bound on how much any single example can influence the model [6], and it is worth evaluating for small, high-sensitivity corpora. Lowering epochs on licensed data, mixing it with larger corpora and holding back the most distinctive modules entirely are cruder but often sufficient.
How to measure code extraction from a fine-tuned model
Measure extraction with a fixed, repeatable protocol and report the numbers to the supplier as a rate, not an anecdote. The core test is prefix prompting: take files from the licensed corpus, feed the model the first N tokens, sample a continuation and measure the longest verbatim overlap with the true continuation. Carlini et al. and Yang et al. both rely on generation followed by matching against the training set [2][3].
Canaries make the result unambiguous. Ask the supplier to plant unique, meaningless strings or synthetic functions (for example, a function with a random 32-character name and a random constant) into a small number of files before delivery, and record which files hold them. If the model ever completes a canary from its prefix, memorization is proven without arguing about whether a common idiom was "copied."
Illustrative example: invented to show structure; it does not describe an available dataset.
| Test | Method | Metric | Example pass threshold to negotiate |
|---|---|---|---|
| Prefix extraction | 2,000 files sampled from the licensed corpus; 64- and 256-token prefixes; greedy and temperature 0.8 sampling | Share of continuations with a verbatim match of 50+ normalized tokens | Agreed with supplier, e.g. below a stated percentage |
| Canary recall | Prefix of each canary-bearing line | Count of canaries completed exactly | Zero |
| Unprompted sampling | 20,000 samples from empty or minimal context | Matches of 50+ tokens against licensed-corpus index | Agreed with supplier |
| Near-verbatim | Same samples, identifiers canonicalized and comments stripped | Match rate on normalized form | Reported, tracked across versions |
| Functional probe | Supplier names 10–20 high-value functions; prompt with docstring or signature only | Execution equivalence against supplier's tests | Reviewed case by case |
| Secret probe | Prefixes ending at api_key =, BEGIN RSA PRIVATE KEY, connection strings | Any real secret emitted | Zero |
Run the suite on every checkpoint you intend to ship, plus after each further fine-tune on top of the licensed model. Keep a held-out set of licensed files that were not trained on as a control, so you can distinguish memorization from the model simply writing plausible code in the supplier's style. Our page on evaluating code dataset samples covers what to check in the data itself before you train.
Deployment controls: output filters against a licensed-corpus index
An output filter is the backstop that catches what training controls miss. The pattern already exists in production: GitHub's Copilot filter compares each suggestion, together with about 150 characters of surrounding code, against an index of public code and can block matching suggestions [7]. GitHub reports matches in under 1% of suggestions, more often when the file is empty or nearly empty [7], and organizations and enterprises can set the block-or-allow policy centrally [8].
For licensed code, build the same structure over the licensed corpus instead of public code. Index normalized token n-grams (or rolling hashes of, say, 50-token windows) from every licensed file and every history revision you trained on. At inference, hash the candidate completion plus its context and block, truncate or regenerate when the overlap exceeds the agreed threshold.
Three failure modes recur. Filters that match only raw text miss reformatted output, so normalize whitespace, comments and identifier names before hashing. Filters tuned too tightly block shared idioms (a standard for loop, a common SQL join) and frustrate users, so exclude n-grams that also appear frequently in open-source corpora. And streaming completions need the check on the accumulated output, not each chunk.
Keep the index access-controlled. It is effectively a copy of the licensed corpus, so the license should state where it lives, who can query it and that it is destroyed on the same schedule as the training data.
Contract terms suppliers ask for
Suppliers of proprietary code usually want technical controls written into the license so they become obligations, not intentions. These are common requests in the market, not SourceX defaults; agree each one per deal with counsel. For broader license structure, see proprietary source code licenses for AI training.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Definition of reproduction: verbatim, near-verbatim or functional, with a token-length threshold and normalization rules.
- Required preprocessing: dedup levels, secret scanning on full history, removal of named modules.
- Extraction testing: the protocol above, canary planting by the supplier, cadence (each released checkpoint) and the report format.
- Output filtering: an index of the licensed corpus applied in production, with a stated threshold and logging of blocked events.
- Use restrictions: for example, evaluation-only use for the most sensitive modules, or no use in models exposed to the supplier's competitors.
- Incident handling: notification when a canary or secret appears in output, remediation steps and whether retraining is required.
- Downstream models: whether controls follow distillation, merging or further fine-tuning of the licensed model.
If you place a general-purpose model on the EU market, map these terms to your copyright policy. As of October 2026, the Copyright chapter of the GPAI Code of Practice asks signatories to adopt and implement such a policy, including measures that address the risk of infringing outputs [9]. Output filters and extraction reports are concrete evidence for that policy.
Fitting controls to the corpus and the model
Control intensity should scale with how distinctive and valuable the licensed code is. Commodity CRUD services and test code tolerate lighter controls; algorithmic cores, security code and anything resembling a trade secret warrant canaries, functional probes and possibly exclusion. Legacy estates such as COBOL and mainframe code are a special case: a small public corpus means licensed files may dominate what the model knows about the language, which raises memorization risk.
Distinctiveness also changes how you specify the request. State in your code dataset request specification which modules you need, whether full history is necessary and which memorization controls you will run, so suppliers can price and approve with the controls in view. The code data buyer's map links the rest of the cluster, and proprietary code datasets with full git history describes the category.
SourceX sources operational datasets, including engineering records, from US companies on request, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery; if you are scoping licensed code with these controls in mind, you can describe it to SourceX at https://sourcex.si/buyers.
Scope licensed code with memorization controls in mind
SourceX looks for US businesses that hold the code and engineering records you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license; nothing is contracted until a supplier agrees. Diligence materials on source, rights, preparation and allowed use are prepared per dataset, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the licensed code you need.
Sources
- Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini (arXiv / ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- Carlini et al. (arXiv / USENIX Security 2021), "Extracting Training Data from Large Language Models" (2020). https://arxiv.org/pdf/2012.07805
- Yang et al. (arXiv / ICSE 2024), "Unveiling Memorization in Code Models" (2023). https://arxiv.org/abs/2308.09932v2
- Meeus et al. (arXiv), "Detecting Functional Memorization in Code Language Models" (2026). https://www.emergentmind.com/papers/2606.12764
- arXiv, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
- Yu et al. (arXiv), "Differentially Private Fine-tuning of Language Models" (2021). https://arxiv.org/pdf/2110.06500
- GitHub Blog, "Introducing code referencing for GitHub Copilot" (2023). https://github.blog/news-insights/product-news/introducing-code-referencing-for-github-copilot
- GitHub Docs, "Finding public code that matches GitHub Copilot suggestions". https://docs.github.com/en/enterprise-cloud@latest/copilot/how-tos/get-code-suggestions/find-matching-code
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- National Institute of Standards and Technology, "Secure Software Development Practices for Generative AI and Dual-Use Foundation Models (NIST SP 800-218A)" (2024). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-218A.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.