Software companies
Open source code in your repos: what to check before licensing it for AI
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Before licensing a codebase for AI training, separate the code your company wrote from open source and other third-party code, because you can only grant rights you hold. Permissive components usually need notices kept, copyleft components raise unsettled questions, and vendored or commercial third-party code is usually excluded. A clean package typically ships first-party code with a file-level inventory.
Key takeaways
- A license grant can only cover code the company owns or may sublicense; third-party code keeps its own terms.
- Permissive licenses generally require keeping copyright and license notices, so stripping headers creates new problems.
- Copyleft licenses generally bar adding further restrictions when the code is passed on, which can clash with the confidentiality terms in a data license.
- Code copied from Q&A sites and blogs is not license-free; it carries share-alike terms or no stated permission at all.
- First-party code still needs a chain of title check for contractors, acquisitions and customer-funded work.
Why does open source code matter in a data license?#
Open source code matters in a data license because the license is a promise about rights, and the supplier can only promise rights it holds. Data licensing agreements commonly include a warranty that the supplier owns or controls what it delivers, and some add an indemnity, so third-party code in the package becomes a contract risk as well as a copyright one.
Few software companies have a repository that is purely their own work. Third-party code arrives through vendored dependency folders, files copied in with their headers, forks of open source projects, snippets pasted from Q&A sites and blog posts, and commercial SDKs. Each has its own terms, and those terms travel with the code into any package you deliver.
How do the main license families compare?#
License families differ mainly in what they require when code is passed on to someone else. The table summarizes the usual pattern; exact obligations depend on the license version and on how the code is used, which counsel should confirm for anything you plan to include.
| License family | Common examples | Typical obligations when passed on | Usual handling in a training package |
|---|---|---|---|
| Permissive | MIT, BSD, Apache 2.0 | Keep copyright and license notices; Apache 2.0 also requires carrying NOTICE file contents and marking changed files, and includes a patent grant | Often excluded for simplicity, or included with notices intact and listed in the inventory |
| Weak copyleft | LGPL, MPL 2.0 | Share source for changes to the licensed library or files under the same license | Usually excluded or isolated pending counsel review |
| Strong copyleft | GPL family | Pass on the code and derivative works under the same license, with source, and without further restrictions | Usually excluded; counsel review if it is central |
| Network copyleft | AGPL | Adds a duty to offer source of modified versions to users who interact with them over a network | Usually excluded and flagged |
| Q&A site snippets | Answers copied from Stack Overflow | Published under a Creative Commons share-alike license, which requires attribution and passing adaptations on under the same terms | Usually excluded; if kept, traced to the answer and attributed |
| No clear license | Blog snippets, files with headers removed | Unknown; code without a license grant generally carries no permission | Excluded unless the origin can be traced and cleared |
| Proprietary third-party | Commercial SDKs, licensed libraries, customer-supplied code | Governed by the vendor or customer agreement, which often bars redistribution | Excluded unless the owner gives written permission |
Is training an AI model on copyleft code a form of distribution?#
Whether AI training on copyleft code counts as distribution is not settled, and reasonable lawyers disagree. The question is assessed deal by deal with counsel, and no supplier should treat either answer as safe by default.
A separate point is clearer and often decides the matter in practice. Delivering a dataset hands copies of the code itself to the licensee, which is a form of passing the code on. Copyleft licenses generally prohibit adding further restrictions when you do that, while a data license normally includes confidentiality, permitted-use and no-redistribution terms. That tension is a common reason suppliers exclude third-party copyleft files from the package rather than try to reconcile the two.
How to build a license inventory for a code package#
A license inventory for a code package is a file-by-file record of what is first-party, what is third-party and under which terms. Open source scanners such as ScanCode Toolkit and FOSSology detect license texts and headers, commercial software composition analysis tools add snippet matching, and SPDX license identifiers give a consistent way to record the results.
History needs its own pass when commits ship with the package. A copyleft library that was vendored years ago and later removed still sits in old commits, and so do files whose headers changed over time. If cleaning history of third-party code is impractical, a current-tree snapshot of first-party paths may be the cleaner scope, at the cost of losing the commit record.
- List the repositories in scope and the branches and history that will ship.
- Mark vendored and generated directories, such as third_party, vendor, node_modules or build output folders, for exclusion review.
- Run a license scanner across every file, including headers in older history if history ships.
- Use snippet matching where available to find code copied from open source projects without its header.
- Record an SPDX identifier or an unknown status for each third-party file.
- Flag commercial SDKs and customer-supplied code separately from open source.
- Have an engineer confirm high-impact findings and counsel decide on anything uncertain.
- Produce a manifest of included first-party paths and excluded third-party paths.
Does first-party code always belong to the company?#
First-party code does not automatically belong to the company, and chain of title is the second half of the review. Under US copyright law, work an employee creates within the scope of the job is a work made for hire owned by the employer. Code from an independent contractor usually falls outside the narrow commissioned-work categories, so the company generally needs a signed written assignment to own it.
Pull the agreements before the package is scoped. A missing IP assignment from an early contractor, or a development agreement that gave a customer ownership of a custom module, is far easier to handle as an exclusion now than as a warranty claim later.
| Source of code | Ownership question | Document to pull |
|---|---|---|
| Employees | Was the work within the job, and did the invention assignment agreement cover it? | Signed employee IP agreements |
| Contractors and agencies | Does the contract assign IP to the company, not just license it? | Contractor agreements and statements of work |
| Acquired companies | Did the purchase transfer the code and its history? | Purchase agreement and IP schedules |
| Customer-funded work | Who owns custom modules built under a customer contract? | Customer development agreements |
| Forks of open source projects | Which parts are upstream code and which are your changes? | Upstream license and fork history |
Illustrative: a construction software company cleans its package#
Illustrative: a fictional company that builds project management software for general contractors plans to license its repositories and pull request history. Its license scan finds a vendored mapping library under a copyleft license, a commercial PDF rendering SDK, a folder of MIT-licensed utilities and several headerless files that match a public project.
Counsel and the CTO exclude the copyleft library, the commercial SDK and the matched files, and leave out the MIT utilities too, since they add little value. The contract review then finds that a scheduling module was built under a development agreement that assigned ownership to a large customer, so that module is excluded as well. The final package is first-party code with a manifest listing every exclusion and the reason for it.
How SourceX approaches third-party code#
SourceX treats third-party code as part of the Rights step of the SourceX five-step transaction, before any code is prepared or delivered. The supplier's license inventory and chain-of-title review decide what is in scope, and SourceX does not make legal determinations on the supplier's behalf; counsel does.
Those decisions become the licensing rights section of the SourceX Evidence Packet. The licensee receives a package whose boundaries are written down: which paths are first-party, which were excluded and why.
Frequently asked questions
Do dependencies listed in package files count as third-party code?
Dependency manifests and lockfiles name packages and versions without containing their code, so they rarely raise the same issue. Dependency folders committed into the repository, or copies vendored into it, do contain third-party code and belong in the inventory.
Can we just remove license headers from third-party files?
No. Removing copyright and license notices generally breaches the conditions of even permissive licenses, and it hides the provenance a buyer will ask about. If a third-party file is not worth keeping with its notices intact, exclude it from the package instead.
What if our own product is open source?
If the company holds the copyright in its own code, it may be able to license that code on different terms, which is how dual licensing works. Contributions from outside developers are a separate question that depends on any contributor agreement, so counsel should review them before they are included.
Will a buyer ask for our license inventory?
Buyers commonly ask how a code package was sourced and what rights cover it, and a file-level inventory is the clearest answer. Treat the inventory and manifest as part of the delivery documentation rather than an internal working file.
Is code written with AI coding assistants a problem?
It may be. The US Copyright Office, which published its report on copyrightability and AI in January 2025, has taken the position that copyright requires human authorship, so heavily generated code may carry weaker ownership claims, and the assistant's terms of use also matter. Record where assistants were used heavily if you know, and raise it with counsel during the rights review.
Sources
- A work made for hire arises when an employee creates the work within the scope of employment, or when a work in certain statutory categories is specially commissioned under a signed written agreement; other contractor work needs an assignment. Source
- Part 2 (Copyrightability) of the US Copyright Office's Copyright and Artificial Intelligence report was published on January 29, 2025. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.