Code and software engineering data
Copyleft Contamination in Licensed Code Datasets: GPL, AGPL and Snippet Risk
Quick answer
A licensed private codebase is rarely 100% first-party code. It usually contains vendored GPL or AGPL libraries, snippets pasted from public repositories and forums, output from licensed code generators, and modules that may belong to a client. Before you train on it, require the supplier to run file- and snippet-level license detection, deliver a per-component inventory with SPDX identifiers, and record a disposition for each finding: removed, kept with notice, or flagged. Then write your copyleft policy into the license.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Where third-party code hides inside a private repository
Most third-party code in a private repository sits in a handful of predictable places, and a supplier who checks only the top-level LICENSE file will miss nearly all of it. The repository's own license statement tells you what the owner intends. It says little about what engineers brought in over ten years of commits.
- Vendored directories.
vendor/,third_party/,external/,node_modules/committed by mistake, Govendor/trees, copied*.jarsources and Git submodules flattened into the tree. These often carry their originalCOPYINGorLICENSEfiles, which makes them the easiest to detect. - Copied snippets. Functions pasted from GitHub, Stack Overflow or library examples, usually with the header comment stripped. This is the hardest category, because nothing in the file announces a license.
- Generated code. Scaffolding, ORM models, protocol buffer stubs and parser tables generated by tools whose output terms vary, plus assistant-generated code whose terms depend on the tool's contract. Our synthetic vs licensed code comparison covers output terms in depth.
- Client-owned or contractor modules. Consultancies and agencies often hold code they wrote for clients, where ownership sits with the client under a services agreement. That is an ownership problem as much as a license problem; see code ownership due diligence.
- Git history. A GPL file deleted from
HEADin 2019 is still in every earlier commit. If you license full history, as described on the proprietary codebases page, the scan must cover history too.
Why copyleft matters for training data, and what is still unsettled
Copyleft licenses attach their main obligations to distribution, and how those obligations apply to model training, weights and outputs is legally unsettled as of October 2026. GPLv2 and GPLv3 condition distribution of copies and derivative works on releasing corresponding source under the same license. LGPL relaxes this for linking. AGPLv3 adds section 13, which extends source-offer duties to modified versions that users interact with over a network, and practitioner commentary generally reads it as not triggered by hosting an unmodified copy.
None of these licenses were written with training in mind, and no settled rule says whether a model trained on GPL code is a derivative work or whether a memorized output carries the license. Treat this as a policy choice your counsel signs off on, not a legal conclusion. The practical risks you can control are concrete: a model reproducing a recognizable copyleft function verbatim in a customer's proprietary product, an AGPL module surfacing in a hosted agent's output, and an audit trail that cannot show what was in the corpus.
Regulatory pressure points the same way. EU AI Act Article 53(1)(c) requires general-purpose AI model providers to put in place a policy to comply with Union copyright law [3], and the GPAI Code of Practice copyright chapter asks signatories to draw up, keep up to date and implement such a policy [4]. A component-level inventory of licensed code is the evidence that policy needs.
License is a design variable, not only a compliance box
License status also shapes evaluation design, because it predicts which code a model has already seen. The SWE-Bench Pro authors argue that permissively licensed public repositories are prime candidates for web-crawled pre-training corpora, and they draw part of their benchmark from copyleft and private commercial repositories to reduce that exposure [1]. The same logic applies to a buyer's own corpus.
If you exclude copyleft from training but keep it in an eval split, record that choice. If a supplier's private repository contains vendored GPL code that is also on public GitHub, those files are both a license question and a contamination question. See code benchmark contamination and held-out evaluation sets from private repositories.
Why declared license metadata is not enough
Declared licenses are often missing or wrong, so detection has to read the files themselves. The Data Provenance Initiative audit of popular datasets found license fields on aggregator platforms frequently unspecified or miscategorized compared with the original sources [2]. The same failure modes appear inside companies: a package.json says "license": "MIT" while a bundled file is GPL-2.0-only, or a header was removed when a function was copied.
Three detection layers cover most of the risk:
- License text and header detection. Full-text scanners such as ScanCode Toolkit compare each file against a database of known license texts and rules, and can emit SPDX or CycloneDX output. Licenses not on the SPDX License List should carry
LicenseRef-identifiers rather than being forced into the nearest match. - Package manifest and lockfile analysis. Resolve
package-lock.json,go.sum,Cargo.lock,pom.xmlandrequirements.txtto declared licenses, then reconcile declared against detected. - Snippet matching. Fingerprint functions and compare them against public code indexes to catch pasted code with no header. Expect false positives on boilerplate such as getters, test fixtures and generated stubs, and require a human review step for each match above a stated threshold.
Deduplication work overlaps with this, because a vendored library is also a near-duplicate of its upstream; see deduplicating code training data.
What to require from a code data supplier
Require a per-component inventory, a detection method statement and a recorded disposition for every third-party finding before the license is signed. The inventory should be machine-readable, keyed to file paths and commit ranges, and delivered alongside the dataset rather than as a PDF summary. Ask for these items in the request itself; our code dataset request specification shows where they fit.
- Tools and versions used, rule sets, snippet-match thresholds and whether full history was scanned.
- One row per component or matched snippet with SPDX expression, confidence and evidence location.
- A disposition per row: removed, kept with notice, or flagged for buyer decision.
NOTICEand attribution files for retained permissive code.- A list of unresolved items and
NOASSERTIONentries, not silently dropped rows. - A statement of known client-owned or contractor code and how it was handled.
Illustrative example: invented to show structure; it does not describe an available dataset.
| path | commits | detected_license (SPDX) | detection | confidence | disposition | notes |
|---|---|---|---|---|---|---|
third_party/libyaml/ | all | MIT | license file + headers | high | kept_with_notice | NOTICE entry added |
src/net/ws_frame.c | a41c..9e07 | GPL-2.0-or-later | snippet match | medium | removed | function replaced by placeholder; history rewritten in delivery copy |
services/report/ | all | AGPL-3.0-only | license file | high | removed | whole directory excluded |
ui/charts/vendor.min.js | all | NOASSERTION | none | low | flagged | minified, origin unknown |
src/billing/tax_rules.py | all | LicenseRef-client-owned | supplier attestation | n/a | removed | client module under services agreement |
Choosing a copyleft policy and writing it into the license
Most buyers pick one of three policies, and the right one depends on how the model will be used and distributed. The license should name the policy, the disposition values and who decides flagged items, so the choice is auditable later.
| Policy | What is excluded | Retained material | Fits when |
|---|---|---|---|
| Strict exclusion | All GPL, LGPL, AGPL, MPL and unknown | Permissive code with attribution files | Model or outputs ship into customer proprietary code |
| Strong-copyleft exclusion | GPL and AGPL; LGPL and MPL reviewed case by case | Permissive and weak copyleft with notices | Internal tools, research or eval-heavy use |
| Disclosure only | Nothing by default | All code, fully inventoried and tagged | Analysis, eval sets or experiments counsel has approved |
Pair the policy with operational controls: tag retained copyleft rows so they can be filtered at training time, keep the inventory with your dataset documentation, and decide in advance how removed files are handled in commit history. For broader scoring of a dataset before acquisition, see training data risk assessment and the AI training data due diligence checklist.
SourceX sources operational datasets, including engineering records, from US companies on request and manages the licensing process; every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Buyers can describe the code data and inventory they need on the SourceX buyers page.
Licensing private code with a clean third-party inventory
SourceX sources engineering records, including code, and other operational data from US companies on request, not from stock, and a request does not guarantee a match. Every release is approved by the supplying company and covered by a license that defines records, uses, term and delivery, with diligence materials on source, rights and preparation prepared per dataset. Browse the Code and software engineering data hub, then describe your request at https://sourcex.si/buyers.
Frequently asked questions
Is removing the LICENSE file enough to remove a copyleft component?
No. The code itself is what carries the license, so the files must be excluded or replaced, and earlier commits still contain them unless the delivery copy rewrites history. Deleting only the license text makes detection harder without changing the underlying obligation.
Should snippet matches below the threshold be ignored?
Record them rather than ignore them. Short matches are often boilerplate, but a pattern of many low-confidence matches in one module can signal a pasted file that was reformatted, which is worth a human look.
Does excluding copyleft code eliminate memorization risk?
It reduces one category of risk. Permissive code still carries attribution conditions, and a model can still reproduce any training code verbatim, so keep output filtering and attribution handling in your plan.
Sources
- arXiv (Scale AI authors), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- arXiv (Data Provenance Initiative), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/pdf/2310.16787
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.