Software companies
Memorization and regurgitation clauses for licensed source code
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
A memorization and regurgitation clause in a source code license limits the risk that a model trained on your code reproduces it verbatim. The three protections are output filters that block long matches, a buyer warranty against verbatim reproduction, and contamination checks that test the model for your code. They work best alongside careful scoping and secret removal before delivery.
Key takeaways
- Models can reproduce distinctive or duplicated training code verbatim, so a code license should address outputs, not just inputs.
- Output filters, a no-verbatim warranty and contamination checks protect different things and work best together.
- Buyers usually accept commercially reasonable obligations and reporting more readily than absolute guarantees.
- Excluding crown-jewel modules and removing secrets before delivery reduces exposure more than any clause can.
What do memorization and regurgitation mean for licensed code?#
Memorization means a model retains specific sequences from its training data closely enough to reproduce them; regurgitation is when it outputs those sequences, verbatim or nearly so, in response to a prompt. For licensed source code, the concern is that a model could hand your implementation to another company's developer.
Risk is not uniform across a codebase. Boilerplate and common patterns get reproduced because they appear everywhere, which matters little. Distinctive code that appears many times in a training set, such as an internal library vendored into many repositories, is more likely to be memorized than code seen once.
The question for a CTO is which parts of the repository would cause harm if reproduced: proprietary algorithms, security-sensitive code, customer-specific integrations and anything carrying embedded secrets. Those areas usually settle the scope before any clause is drafted.
The three contractual protections compared#
The three contractual protections act at different points in a model's life: output filters work when the model answers, the no-verbatim warranty sets the standard the buyer must meet, and contamination checks test whether the standard is being met. Comparing them side by side shows which to ask for first.
Most negotiated licenses combine all three in a softer form: a commercially reasonable filter obligation, a warranty tied to known reproductions, and periodic testing with a summary report to the licensor. The softer forms are easier to agree and still give the licensor a basis to act when a reproduction appears.
The order of priority depends on how the buyer will use the code. A model deployed as a public coding assistant warrants stronger filters and testing than one used only inside the buyer's research environment, and the clause can scale its obligations to the deployment.
| Protection | What the clause requires | What it can do | Limits |
|---|---|---|---|
| Output filters | Buyer runs filters that detect and block long verbatim matches to licensed code | Stops many exact reproductions at answer time | Near-copies with renamed variables can pass; filters need the delivered corpus as a reference |
| No-verbatim warranty | Buyer warrants that deployed models will not knowingly output substantial verbatim portions of licensed code | Creates a contractual standard and a remedy | Hard to define precisely; buyers resist absolute wording |
| Contamination checks | Buyer tests models with prompts drawn from licensed code and reports results | Produces evidence of memorization before and after release | Testing samples the model rather than proving a negative; needs an agreed method |
What should a memorization clause define?#
A memorization clause should define its terms precisely enough that both sides can tell when it has been breached. Vague wording such as a promise not to reproduce the licensed code invites a dispute the first time a partial match appears.
Thresholds deserve the most care. Set too low, the clause flags ordinary code that any developer would write; set too high, it allows meaningful copying. Engineers on both sides should agree the definition, not only counsel.
- Licensed code: the exact repositories, branches, date ranges and file types delivered.
- Verbatim output: a match threshold the parties agree, described by length or similarity, and whether near-copies count.
- Excluded material: boilerplate, common idioms and open source components that are not proprietary to you.
- Filter obligation: what the buyer must run, at which point in its systems, and how the delivered corpus serves as the reference.
- Testing: who designs the prompts, how often tests run and what the report contains.
- Notice and remediation: how either party reports a reproduction and what the buyer must then do, such as filtering, suppressing or retraining.
- Remedies: whether a breach triggers a cure period, termination, deletion or damages.
How do contamination checks work in practice?#
Contamination checks work by prompting the trained model with the opening portion of a licensed file or function and measuring how closely its completion matches the original. Repeating the test across a sample of files gives an estimate of how much licensed code the model has memorized.
The term contamination also covers a second risk: licensed code leaking into public benchmarks or evaluation sets the buyer publishes. The license can require that licensed code stay out of any published dataset, benchmark or model documentation examples.
Checks produce evidence, not certainty. Agree in advance which sample is tested, which result counts as a failure and what happens next, so a poor result leads to remediation rather than an argument.
Which steps before delivery matter more than the clause?#
The steps before delivery often matter more than the clause, because code that is never delivered cannot be reproduced. A CTO controls these steps directly, without negotiating with anyone.
Secret scanning deserves its own line in the plan. Gitleaks, an MIT-licensed tool for detecting passwords, API keys and tokens in git repositories, is widely used for this, though its README was updated in May 2026 to say the project is feature complete and future releases will be security patches only. Whatever tool you choose, scan the full commit history, not just the current branch.
| Pre-delivery step | Why it matters |
|---|---|
| Exclude crown-jewel modules | Core algorithms and pricing logic stay out entirely |
| Remove secrets and credentials from full history | Keys, tokens and passwords become a security risk if reproduced |
| Deduplicate vendored copies | Repeated code is more likely to be memorized |
| Separate third-party and open source code | You cannot license what you do not own, and copyleft terms may apply |
| Strip customer-specific integrations | Customer names, endpoints and configurations are not yours to license |
| Write a delivery manifest | Filters and checks need an exact reference set |
Illustrative: a routing software vendor negotiates output terms#
Illustrative: a fictional vendor of route planning software for regional freight carriers agrees to license its repository history, pull requests and code reviews for training coding assistants. The CTO's first concern is the optimization engine that sets the company apart from competitors.
The engine's modules are excluded from delivery, secrets are scanned out of the full history and vendored libraries are deduplicated. The license defines verbatim output by an agreed similarity threshold, requires the buyer to run an output filter against the delivered corpus and to test memorization on a sample after each major training run.
The buyer declines an absolute warranty but accepts one tied to known reproductions, with a duty to suppress them once reported. The final terms pair a clear testing method with a delivery manifest both sides can check against.
How SourceX handles output protections#
SourceX handles output protections as part of the Approval step of the SourceX five-step transaction, where the supplier reviews permitted use, restrictions and remedies before signing. Code packages are scoped and prepared before that step, so exclusions are settled early.
The SourceX Evidence Packet records provenance, licensing rights, permitted use, the privacy record and release authorization, including which repositories and modules were delivered. That record is the reference set that filters and contamination checks are measured against.
Frequently asked questions
Can we allow evaluation but prohibit training on our code?
Yes. Permitted use can be limited to evaluation, testing or reinforcement learning environments rather than training, with an express bar on later training. Evaluation-only use lowers memorization risk because the model is not trained on the code, though leakage into published benchmarks still needs its own restriction.
Will buyers accept a no-verbatim warranty?
Some will accept a qualified version. Absolute guarantees about model outputs are difficult for any developer to give, so warranties tied to knowledge, reasonable filtering and prompt remediation are more common. Expect negotiation over thresholds and remedies.
Does exclusivity reduce regurgitation risk?
Not directly. Exclusivity limits who receives the code, while regurgitation concerns what a model outputs to everyone who uses it. A non-exclusive license with strong output terms can protect you better than an exclusive one without them.
What about open source code inside our repositories?
Open source components are licensed to you under their own terms and are usually separated from the package or delivered with their notices intact. Copyleft licenses raise specific questions that are assessed with counsel, so many suppliers exclude those components entirely.
Should output obligations survive the end of the license?
Usually yes. A model trained during the term keeps whatever it memorized after the term ends, so filter, reporting and remediation duties often survive termination for any model trained on the code. Deleting the delivered files does not remove what a model has already learned, which is why survival language matters.
How do we know whether a model reproduces our code after release?
Through the reporting and testing the license requires, plus your own spot checks using prompts drawn from the delivered code. Keep the delivery manifest, because it defines the reference set for any comparison.
Sources
- Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories; on May 21, 2026 its README was updated to say it is feature complete and future releases will be security patches only. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.