AI data market
Open-source code in your repositories: what it means for licensing to AI
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Open-source code in your repositories does not stop you licensing your own code to AI developers, but you cannot license rights you do not hold. Separate first-party code from vendored or copied open-source code, customer code and secrets, keep license notices with any third-party files you include, and disclose what remains so buyers can apply their own policies.
Key takeaways
- Your license can cover only code you own; open-source code stays under its own license terms.
- Most of a repository's licensing value sits in first-party code and its review and issue history, not in public libraries.
- Copyleft licenses such as GPL and AGPL raise open questions for AI training that counsel should assess deal by deal.
- Software composition analysis and an SBOM give buyers a component inventory they can check.
- Customer-owned code and committed secrets are separate problems and are usually excluded or removed.
Why open-source code complicates a repository license#
Open-source code complicates a repository license because a typical private repository mixes code the company owns with code it only uses under someone else's license. A company can grant an AI developer rights in its own code, but it cannot grant more rights in third-party code than the open-source license already gives everyone.
Open-source licenses fall into two broad families. Permissive licenses such as MIT, BSD and Apache 2.0 generally allow reuse and redistribution if copyright and license notices are kept. Copyleft licenses such as GPL, LGPL, AGPL and MPL attach conditions to distribution, such as making source available under the same terms. How either family applies to AI training, and whether delivering a code dataset counts as distribution, are questions without settled answers.
There is also a value point. Public open-source code is already available to buyers, so including it adds little. The licensable value of a private repository is usually the first-party code and its history: commits, pull request discussions, code reviews and the issues that drove each change.
Four categories to separate before any deal#
Separating a repository into categories turns a vague legal worry into a concrete scoping exercise. Engineering leads can usually do the first pass, with counsel reviewing the edge cases.
The hardest boundary runs between first-party code and code adapted from open source. A function copied from a public library and reworked over the years may now be mostly yours, or may still be a derivative of the original. Flag these cases rather than resolving them silently, and let counsel decide whether to exclude or disclose them.
| Category | Examples | Usual handling |
|---|---|---|
| First-party code | Services, business logic, tests and internal tools written by employees or contractors with assignments | Core of the license |
| Vendored or copied open source | Vendor and third-party directories, forked libraries, copied functions, snippets from public answers | Exclude, or include with notices and disclosure after counsel review |
| Customer code | Forks maintained for one customer, custom integrations a customer owns under contract, customer scripts attached to tickets | Exclude |
| Secrets and credentials | API keys, tokens, passwords, private keys and environment files anywhere in history | Find, revoke and remove before delivery |
| Generated and build files | Lockfiles, compiled output, minified bundles, generated clients | Usually exclude as low-value noise |
How to find open-source code that was copied in#
Copied open-source code is harder to find than declared dependencies, because it sits inside your own files with no manifest entry. A combination of tools and interviews usually finds most of it.
Record what each tool scanned and when. A scan result helps a buyer only if it is tied to the exact branches and commit range being delivered.
- Run a software composition analysis tool that scans both package manifests and file contents for known open-source code and snippets.
- Generate a software bill of materials in a standard format such as SPDX or CycloneDX, and keep it with the license file.
- Search for LICENSE, NOTICE and COPYING files, and for license headers inside source files.
- Compare vendored directories against package manifests to find libraries copied in rather than installed.
- Review git history for large single-commit imports, which often mark copied projects.
- Ask long-tenured engineers about forks and patched libraries that never made it upstream.
Copyleft licenses: what to check with counsel#
Copyleft licenses deserve a specific review because their conditions are triggered by actions such as distribution or, for AGPL, network use. Counsel will want to know which copyleft components exist, where they sit, whether they were modified, and whether delivering them inside a dataset could be treated as distribution.
A conservative approach is to exclude copyleft files from the package and reference them only through the dependency list. Where a buyer wants them included, keep every license notice intact, disclose the components, and let the buyer apply its own policy. None of this settles whether training a model on copyleft code creates obligations for the model; that question is unresolved and is assessed deal by deal.
If your company has published some of its own code under an open-source license, that code is already available to anyone on those terms. Its private history, such as internal reviews and design discussions, may still be licensable.
What buyers will expect you to disclose#
Buyers expect a plain description of what is in the package and what was removed. Licensing provenance has become a visible concern across AI training data: the Data Provenance Initiative audited more than 1,800 fine-tuning datasets and documented their sources, licenses and creators, and the Data & Trust Alliance's Data Provenance Standards include metadata elements for license to use and copyright status.
For a code package, that translates into a short disclosure set: the component inventory with licenses, the directories excluded and why, how copied snippets were handled, how customer code was identified and removed, the secrets scanning method and results, and any known gaps. A buyer's own counsel may run a separate scan, so consistency between your inventory and their findings matters more than perfection.
Illustrative: a logistics software company with a monorepo#
Illustrative: a fictional company sells transportation management software to freight brokers. Its monorepo holds the core TMS services, a vendor directory of copied libraries, a forked GPL-licensed PDF library, customer-specific EDI mapping scripts and years of pull requests linked to Jira issues.
The CTO ran a software composition analysis scan and produced an SBOM. The vendor directory and the GPL fork were excluded and listed in the disclosure. Customer EDI mappings were excluded because several customer contracts assigned ownership of custom work to the customer. Gitleaks, an MIT-licensed secrets scanner, found carrier API keys in early commits; the keys were revoked and the affected files rewritten in the delivered copy.
The package that remained was first-party TMS code with its review and issue history, delivered with the component inventory. Counsel signed off on the exclusions, and the company kept its own repository unchanged.
How SourceX handles mixed repositories#
SourceX handles mixed repositories in the Rights and Preparation steps of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. The fit check asks for metadata such as hosting platform, languages, years of history and known third-party or customer code, not the code itself.
The SourceX Evidence Packet records provenance, licensing rights including the component inventory and exclusions, permitted use, the privacy record covering secrets and de-identification, and release authorization. As with any large dataset, the code remains in the company's own storage or travels on encrypted drives rather than being hosted by SourceX.
Frequently asked questions
Can we license code that depends on open-source libraries without including them?
Yes. Dependencies declared in package manifests can be left out of the package and referenced through the manifest, since buyers can obtain public libraries themselves. The license then covers only your first-party code and history, which is where the value usually sits.
Is permissively licensed code safe to include?
Permissive licenses generally allow redistribution with notices kept, so inclusion is often possible. It rarely adds value, though, because the code is public. Excluding it anyway keeps the package focused and the disclosure short, and counsel can confirm which approach suits the deal.
Does AI-generated code in our repositories count as third-party code?
Not in the same way, but it raises its own questions about authorship and ownership that are not settled. Record when code assistants were adopted by each team and disclose it. Review discussions and issue histories remain human records regardless.
Will buyers run their own open-source scan?
Expect it. Buyers' legal and engineering teams often scan delivered code for licenses and secrets. Providing your own inventory, scan method and exclusions up front shortens that review and reduces the chance of a surprise finding late in the deal.
What if a customer's code was merged into our main product?
That needs a contract review. If the customer's agreement gave it ownership of custom work, or restricted reuse of its code, the merged portions may have to be excluded even though they now sit in your product. The original pull requests and commit history usually show where the code came from.
Sources
- The Data Provenance Initiative released a first audit covering 44 data collections spanning more than 1,800 fine-tuning text-to-text datasets, documenting their sources, licenses, creators and other metadata. Source
- The Use group of the Data & Trust Alliance Data Provenance Standards includes elements for license to use, intended data use, and copyright, patent and trademark status. Source
- Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.