Skip to content

Code and software engineering data

Copyleft Contamination in Licensed Code Datasets: GPL, AGPL and Snippet Risk

Quick answer

A licensed private codebase is rarely 100% first-party code. It usually contains vendored GPL or AGPL libraries, snippets pasted from public repositories and forums, output from licensed code generators, and modules that may belong to a client. Before you train on it, require the supplier to run file- and snippet-level license detection, deliver a per-component inventory with SPDX identifiers, and record a disposition for each finding: removed, kept with notice, or flagged. Then write your copyleft policy into the license.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Where third-party code hides inside a private repository

Most third-party code in a private repository sits in a handful of predictable places, and a supplier who checks only the top-level LICENSE file will miss nearly all of it. The repository's own license statement tells you what the owner intends. It says little about what engineers brought in over ten years of commits.

  • Vendored directories. vendor/, third_party/, external/, node_modules/ committed by mistake, Go vendor/ trees, copied *.jar sources and Git submodules flattened into the tree. These often carry their original COPYING or LICENSE files, which makes them the easiest to detect.
  • Copied snippets. Functions pasted from GitHub, Stack Overflow or library examples, usually with the header comment stripped. This is the hardest category, because nothing in the file announces a license.
  • Generated code. Scaffolding, ORM models, protocol buffer stubs and parser tables generated by tools whose output terms vary, plus assistant-generated code whose terms depend on the tool's contract. Our synthetic vs licensed code comparison covers output terms in depth.
  • Client-owned or contractor modules. Consultancies and agencies often hold code they wrote for clients, where ownership sits with the client under a services agreement. That is an ownership problem as much as a license problem; see code ownership due diligence.
  • Git history. A GPL file deleted from HEAD in 2019 is still in every earlier commit. If you license full history, as described on the proprietary codebases page, the scan must cover history too.

Why copyleft matters for training data, and what is still unsettled

Copyleft licenses attach their main obligations to distribution, and how those obligations apply to model training, weights and outputs is legally unsettled as of October 2026. GPLv2 and GPLv3 condition distribution of copies and derivative works on releasing corresponding source under the same license. LGPL relaxes this for linking. AGPLv3 adds section 13, which extends source-offer duties to modified versions that users interact with over a network, and practitioner commentary generally reads it as not triggered by hosting an unmodified copy.

None of these licenses were written with training in mind, and no settled rule says whether a model trained on GPL code is a derivative work or whether a memorized output carries the license. Treat this as a policy choice your counsel signs off on, not a legal conclusion. The practical risks you can control are concrete: a model reproducing a recognizable copyleft function verbatim in a customer's proprietary product, an AGPL module surfacing in a hosted agent's output, and an audit trail that cannot show what was in the corpus.

Regulatory pressure points the same way. EU AI Act Article 53(1)(c) requires general-purpose AI model providers to put in place a policy to comply with Union copyright law [3], and the GPAI Code of Practice copyright chapter asks signatories to draw up, keep up to date and implement such a policy [4]. A component-level inventory of licensed code is the evidence that policy needs.

License is a design variable, not only a compliance box

License status also shapes evaluation design, because it predicts which code a model has already seen. The SWE-Bench Pro authors argue that permissively licensed public repositories are prime candidates for web-crawled pre-training corpora, and they draw part of their benchmark from copyleft and private commercial repositories to reduce that exposure [1]. The same logic applies to a buyer's own corpus.

If you exclude copyleft from training but keep it in an eval split, record that choice. If a supplier's private repository contains vendored GPL code that is also on public GitHub, those files are both a license question and a contamination question. See code benchmark contamination and held-out evaluation sets from private repositories.

Why declared license metadata is not enough

Declared licenses are often missing or wrong, so detection has to read the files themselves. The Data Provenance Initiative audit of popular datasets found license fields on aggregator platforms frequently unspecified or miscategorized compared with the original sources [2]. The same failure modes appear inside companies: a package.json says "license": "MIT" while a bundled file is GPL-2.0-only, or a header was removed when a function was copied.

Three detection layers cover most of the risk:

  1. License text and header detection. Full-text scanners such as ScanCode Toolkit compare each file against a database of known license texts and rules, and can emit SPDX or CycloneDX output. Licenses not on the SPDX License List should carry LicenseRef- identifiers rather than being forced into the nearest match.
  2. Package manifest and lockfile analysis. Resolve package-lock.json, go.sum, Cargo.lock, pom.xml and requirements.txt to declared licenses, then reconcile declared against detected.
  3. Snippet matching. Fingerprint functions and compare them against public code indexes to catch pasted code with no header. Expect false positives on boilerplate such as getters, test fixtures and generated stubs, and require a human review step for each match above a stated threshold.

Deduplication work overlaps with this, because a vendored library is also a near-duplicate of its upstream; see deduplicating code training data.

What to require from a code data supplier

Require a per-component inventory, a detection method statement and a recorded disposition for every third-party finding before the license is signed. The inventory should be machine-readable, keyed to file paths and commit ranges, and delivered alongside the dataset rather than as a PDF summary. Ask for these items in the request itself; our code dataset request specification shows where they fit.

  • Tools and versions used, rule sets, snippet-match thresholds and whether full history was scanned.
  • One row per component or matched snippet with SPDX expression, confidence and evidence location.
  • A disposition per row: removed, kept with notice, or flagged for buyer decision.
  • NOTICE and attribution files for retained permissive code.
  • A list of unresolved items and NOASSERTION entries, not silently dropped rows.
  • A statement of known client-owned or contractor code and how it was handled.

Illustrative example: invented to show structure; it does not describe an available dataset.

pathcommitsdetected_license (SPDX)detectionconfidencedispositionnotes
third_party/libyaml/allMITlicense file + headershighkept_with_noticeNOTICE entry added
src/net/ws_frame.ca41c..9e07GPL-2.0-or-latersnippet matchmediumremovedfunction replaced by placeholder; history rewritten in delivery copy
services/report/allAGPL-3.0-onlylicense filehighremovedwhole directory excluded
ui/charts/vendor.min.jsallNOASSERTIONnonelowflaggedminified, origin unknown
src/billing/tax_rules.pyallLicenseRef-client-ownedsupplier attestationn/aremovedclient module under services agreement

Choosing a copyleft policy and writing it into the license

Most buyers pick one of three policies, and the right one depends on how the model will be used and distributed. The license should name the policy, the disposition values and who decides flagged items, so the choice is auditable later.

PolicyWhat is excludedRetained materialFits when
Strict exclusionAll GPL, LGPL, AGPL, MPL and unknownPermissive code with attribution filesModel or outputs ship into customer proprietary code
Strong-copyleft exclusionGPL and AGPL; LGPL and MPL reviewed case by casePermissive and weak copyleft with noticesInternal tools, research or eval-heavy use
Disclosure onlyNothing by defaultAll code, fully inventoried and taggedAnalysis, eval sets or experiments counsel has approved

Pair the policy with operational controls: tag retained copyleft rows so they can be filtered at training time, keep the inventory with your dataset documentation, and decide in advance how removed files are handled in commit history. For broader scoring of a dataset before acquisition, see training data risk assessment and the AI training data due diligence checklist.

SourceX sources operational datasets, including engineering records, from US companies on request and manages the licensing process; every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Buyers can describe the code data and inventory they need on the SourceX buyers page.

Licensing private code with a clean third-party inventory

SourceX sources engineering records, including code, and other operational data from US companies on request, not from stock, and a request does not guarantee a match. Every release is approved by the supplying company and covered by a license that defines records, uses, term and delivery, with diligence materials on source, rights and preparation prepared per dataset. Browse the Code and software engineering data hub, then describe your request at https://sourcex.si/buyers.

Frequently asked questions

Is removing the LICENSE file enough to remove a copyleft component?

No. The code itself is what carries the license, so the files must be excluded or replaced, and earlier commits still contain them unless the delivery copy rewrites history. Deleting only the license text makes detection harder without changing the underlying obligation.

Should snippet matches below the threshold be ignored?

Record them rather than ignore them. Short matches are often boilerplate, but a pattern of many low-confidence matches in one module can signal a pasted file that was reformatted, which is worth a human look.

Does excluding copyleft code eliminate memorization risk?

It reduces one category of risk. Permissive code still carries attribution conditions, and a model can still reproduce any training code verbatim, so keep output filtering and attribution handling in your plan.

Sources

  1. arXiv (Scale AI authors), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
  2. arXiv (Data Provenance Initiative), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/pdf/2310.16787
  3. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  4. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data