Software companies
GPL and AGPL code in your repo: can it be part of an AI training license?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
GPL and AGPL code your company did not write should usually be excluded from an AI training license, because you can pass it on only under its copyleft terms, which forbid the restrictions a data license adds. Code your company authored can stay, even in a repository that also holds copyleft files. Whether training itself triggers copyleft duties is unsettled.
Key takeaways
- Exclude GPL and AGPL files written by outsiders; you can redistribute them only under their own license.
- Typical data license terms, such as confidentiality and no onward sharing, may conflict with copyleft's ban on further restrictions.
- A company that holds all the copyright in its own GPL project can usually license that code on separate terms.
- Whether a model trained on copyleft code inherits copyleft obligations is unsettled, so suppliers avoid relying on either answer.
Can GPL or AGPL code go into an AI training license?#
GPL or AGPL code written by someone else generally should not go into an AI training license, because the company does not own it and can only pass it on under the original copyleft terms. A data license that adds confidentiality, use limits or a ban on redistribution may conflict with the copyleft rule that every recipient gets the same freedoms the distributor received.
The company's own code is a different matter. First-party code in the same repository remains licensable, so the practical task is separating files by origin rather than discarding whole repositories.
The wider question, whether training a model on copyleft code creates obligations for the model or its outputs, is not settled. Treat it as an open risk the buyer will weigh, not a question the supplier can answer on the buyer's behalf.
How copyleft code ends up in a company repository#
Copyleft code reaches company repositories through ordinary engineering habits, and teams often underestimate how much is there. A license scan usually finds it in places nobody remembers adding.
Acquired code deserves extra attention. An acquired company's repositories may hold copyleft components that were acceptable for internal use but were never reviewed for redistribution, and the engineers who added them have often left.
- Vendored dependencies: a GPL library copied into a vendor or third-party folder instead of installed through a package manager.
- Copied snippets: functions pasted from GPL projects or from answers that carried a copyleft license.
- Forks: an internal fork of an AGPL server or GPL tool, modified for the company's needs.
- Build and test tooling: GPL scripts, fixtures or sample files bundled for convenience.
- Acquired code: repositories inherited through an acquisition with no license inventory.
- Dual-licensed components: code offered under GPL or a commercial license, where nobody recorded which license the company actually took.
Why delivering copyleft files in a dataset still counts#
Delivering copyleft files in a dataset is still a distribution of those files, so the copyleft terms travel with them. The GPL lets anyone redistribute the source code if the license notice is kept and recipients receive the same rights, and both GPL version 2 and version 3 bar the distributor from imposing further restrictions on those rights.
Data licenses normally do add restrictions: permitted use limited to training and evaluation, confidentiality, deletion on termination and no onward sharing. Putting those terms on GPL files creates a conflict counsel would have to resolve, and the simplest resolution is exclusion. The supplier also cannot give meaningful warranties about code whose authors it never dealt with.
The AGPL adds a network clause: if you modify AGPL software and let users interact with it over a network, you must offer those users the corresponding source of your modified version. That matters for SaaS companies that run modified AGPL software in production, because their changes already sit under AGPL terms, and it makes AGPL forks worth isolating before any package is assembled.
Decision rule by code origin#
The decision rule is to keep what the company wrote, exclude copyleft code it did not, and document the call for everything in between. The table covers the origins that come up most often in software company repositories.
Files that mix first-party and copyleft code in the same function or module are the hard cases. The safer default is to exclude them rather than attempt a line-by-line split, and to note the exclusion in the inventory.
| Code origin | Default treatment | Reasoning |
|---|---|---|
| Written by employees as part of their job | Keep | The company usually owns it through employment |
| Written by contractors with a signed IP assignment | Keep | The assignment transfers ownership to the company |
| Company's own project released under GPL, with no outside contributors | Usually keep | As sole copyright holder, the company can license it on other terms |
| Company's GPL project with outside contributions and no contributor agreement | Exclude contributed files, or ask counsel | Contributors kept their own copyright |
| Vendored or copied GPL or AGPL code | Exclude | Licensable only under its own terms |
| Modified forks of GPL or AGPL projects | Exclude upstream code; review the company's changes | Changes may be derivative of the copyleft code |
| Permissive open source, such as MIT, Apache or BSD | Separate review | Notice duties travel with the files, and public code adds little a buyer does not already have |
How to find GPL and AGPL code before scoping a package#
Finding copyleft code takes a scan plus a conversation, because tools catch declared licenses and people remember undeclared copies. Run both before the package scope is shared with any buyer.
Keep the inventory with the package. The same record answers a buyer's diligence questions, supports a later removal request and shows a future acquirer of the company how third-party code was handled.
- Run a software composition analysis scan across every repository in scope, including archived and forked ones.
- Search file headers for license text and SPDX identifiers such as GPL-2.0-only, GPL-3.0-or-later and AGPL-3.0-only.
- Read package manifests and lock files to see which copyleft dependencies were vendored rather than fetched.
- Review git history for folders added in a single bulk commit, a common sign of vendored code.
- Ask the longest-serving engineers about forks and copied tools, since scans miss uncredited snippets.
- Record each result in a file-level inventory: path, detected license, origin and decision.
Illustrative: a DevOps tooling company cleans its monorepo#
Illustrative: a fictional DevOps tooling company wants to license its monorepo history, code reviews and linked issue tracker records. A license scan finds a vendored GPL compression library, an internal fork of an AGPL monitoring server, and a GPL command-line tool the company itself published years earlier.
The general counsel excludes the vendored library and the upstream code of the AGPL fork. The company's own changes to the fork are excluded too, because they are woven into the upstream files. The company-published tool stays in scope after its contributor history shows that every outside pull request came in under a contributor license agreement granting the company broad relicensing rights. Code reviews that quote excluded files are redacted, and the inventory goes into the transaction record.
How SourceX approaches copyleft code#
SourceX handles third-party code across the Rights and Preparation steps of the SourceX five-step transaction. The rights review separates first-party code from copyleft and other third-party code; preparation removes excluded files and redacts excerpts of them inside code reviews and issues.
The SourceX Evidence Packet carries the file inventory along with provenance, the licensing rights relied on and permitted use, so a buyer's counsel can see how each exclusion was decided. The supplier approves the final file list before Delivery.
Frequently asked questions
Is training a model on GPL code a form of distribution?
Nobody can say with confidence yet. Some argue a trained model is neither a copy nor a derivative of its training code; others argue models can reproduce code and should carry its terms. The question is unsettled in US law, so suppliers avoid depending on either view by excluding third-party copyleft files.
What about LGPL and MPL code?
Weaker copyleft licenses such as LGPL and MPL apply mainly to the licensed library or files rather than to everything linked with them. Those files must still stay under their own terms, so the same rule applies: exclude the third-party files and keep the company's surrounding code.
Can we keep code that merely calls a GPL library?
Usually the calling code is first-party and can stay, although whether linking creates a derivative work has long been debated for GPL software. A practical default is to keep code that imports a library, exclude the library's own files and note the dependency in the inventory.
Do buyers want copyleft code at all?
Public open source code is widely available to AI developers under its own licenses, so a supplier's distinctive contribution is private, first-party engineering history. Copyleft files add legal friction without adding much that is unique, which is another reason to leave them out.
What if a scan misses a copied GPL snippet?
Scans can miss short or uncredited snippets. Contracts usually handle this through a representation limited to the supplier's knowledge and a process for removing files found later. Record the scan method and date so the diligence effort is visible to the buyer.
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.