Software companies
Open-source code inside your repositories: what must be excluded
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Open-source code inside proprietary repositories should be excluded from an AI data license, or handled only on its own license terms, because a company can license only rights it holds. Exclusion has to cover vendored libraries, copied snippets, generated code and forks, and it has to reach commit history and code reviews, not just today's files.
Key takeaways
- A company can license only code it holds rights in; third-party code is excluded or handled on its own terms.
- Vendored libraries, committed dependencies, copied snippets, forks, generated code and bundled assets all need screening.
- Deleting files from the current branch does not remove them from commit history or old review diffs.
- An exclusion manifest records what was removed, why, how it was found and who approved it.
- Copyleft code carries the strongest conditions, and code with no license gives no permission to redistribute at all.
- Whether permissively licensed code may stay with its notices is a counsel decision made deal by deal.
Why does third-party code have to come out of a data license?#
Third-party code has to come out of a data license because the company can grant only rights it holds. Open-source licenses permit use and redistribution on conditions. Permissive licenses such as MIT, BSD and Apache 2.0 mainly require their notices to travel with the code. Copyleft licenses such as the GPL family generally require derivative works to be passed on under the same terms without added restrictions, which sits badly with a confidential, use-limited data license.
Whether training an AI model on open-source code triggers those conditions is unsettled, and the answer may differ by license family. That uncertainty is a reason for caution, not a reason to include. The cautious default is to ship first-party code only and treat third-party code as excluded unless counsel approves a specific approach.
The general counsel's job is to set the rule; engineering's job is to find everything the rule applies to. The second part is harder than it sounds, because open-source code rarely sits in one tidy folder.
The screen list: where open-source code hides#
Open-source code hides in predictable places, and a screen list keeps the search systematic. Work through each category for every repository in scope, including archived ones.
Commercial third-party code, such as licensed SDKs or code received from a customer or partner, follows the same path even though it is not open source. Its license or contract rarely allows redistribution in a dataset.
| Category | Where it hides | How it is found | Default treatment |
|---|---|---|---|
| Vendored libraries | vendor, third_party, lib or external folders | Path rules and license files | Exclude |
| Committed dependencies | Package folders committed by mistake | Path rules and package manifests | Exclude |
| Copied snippets | Inside first-party files, often near a comment or URL | Snippet matching and comment search | Exclude the snippet or the file |
| Forks and patched upstream projects | Internal copies of open-source projects with company changes | Repository origin and history | Exclude, including company patches, unless counsel agrees otherwise |
| Generated code | Clients, stubs and scaffolding produced by tools | Generator headers and build configuration | Exclude unless generator terms are reviewed |
| Bundled assets | Minified scripts, CSS frameworks, fonts, icons | File type rules and headers | Exclude |
| Test fixtures and sample data | Files copied from open-source projects or public datasets | Manual review and matching | Exclude |
How do license families change the treatment?#
License families change how much is at stake if third-party code slips through, but they rarely change the default of exclusion. Copyleft code carries the strongest conditions, and code with no license at all gives no permission to redistribute it.
Record the detected license for every finding, even when the treatment is exclude. The license family tells reviewers how carefully the filtered history must be checked, and it gives counsel a clear basis for any exception. The table shows usual starting points; counsel sets the actual rule.
| License family | Examples | Main condition | Usual starting point |
|---|---|---|---|
| Permissive | MIT, BSD, Apache 2.0 | Keep copyright and license notices; Apache 2.0 adds patent and NOTICE terms | Exclude; counsel may allow with notices |
| Weak copyleft | LGPL, MPL 2.0, EPL | Changes to the covered files or library stay under the same license | Exclude |
| Strong copyleft | GPL-2.0, GPL-3.0 | Derivative works passed on under the same terms, without added restrictions | Exclude |
| Network copyleft | AGPL-3.0 | Source must also be offered to users who interact with modified software over a network | Exclude |
| Source-available | Licenses that limit commercial or competing use | Use limits set by the licensor | Exclude |
| No license found | Code from blogs, gists or public repositories with no license | Default copyright applies, so there is no permission to redistribute | Exclude |
Generated code is the category teams forget#
Generated code is the category teams forget because it looks like the company's own code and often lives beside it. API clients, database models, protocol buffer stubs and project scaffolding are produced by tools whose templates may carry their own license terms.
Code suggested by AI coding assistants raises a related question: output may occasionally reproduce public code, and the assistant's terms may address ownership. Treat generator output as excluded until someone has checked the tool and template terms, and flag AI-assisted code for counsel's view rather than assuming it is first-party.
Exclusion has to reach history, reviews and tickets#
Exclusion has to reach every layer of a code package, not just the current snapshot. A library vendored years ago and later deleted still sits in commit history, appears in old pull request diffs and may be attached to issues as a patch.
Filtering history changes commit identifiers, so links between issues, commits and reviews have to be remapped in the delivered copy. Plan that step early; a clean package with broken links loses much of what made the history useful.
| Layer | What to exclude | How |
|---|---|---|
| Current snapshot | Excluded paths and files | Path filters applied to the export |
| Commit history | Every version of excluded files | Filtered history built on a separate copy, never the live repository |
| Pull request diffs | Diff hunks that touch excluded paths | Drop or trim the affected hunks |
| Review comments | Comments quoting excluded code | Search for and remove quoted blocks |
| Issues and attachments | Patches and files copied from upstream projects | Scan attachments with the same rules |
How to scan repositories before scoping#
Scanning works best as a layered pass that combines automated tools with human review of everything unresolved. No single tool finds every category on the screen list.
Secret scanners belong in the same pass, because credentials must come out too. Gitleaks, an MIT-licensed tool, detects passwords, API keys and tokens in Git repositories and files. TruffleHog, licensed under AGPL-3.0, can try to log in with a secret it finds to confirm whether it is still live, so agree with security before enabling that check. Running these tools internally does not put their code into the package.
- Read package manifests and lockfiles to list declared dependencies.
- Apply path rules for vendor, third-party and dependency folders across every branch in scope.
- Detect license files and license headers inside source files.
- Run snippet matching through a software composition analysis tool to find copied code.
- Generate a software bill of materials in a standard format such as SPDX or CycloneDX.
- Log every finding in the exclusion manifest with its path, category, detected license and the tool that found it.
- Send every unknown or ambiguous finding to a named reviewer.
Who decides what stays and what goes?#
Counsel decides the rules and engineering applies them. The general counsel or outside counsel sets the treatment for each license family, decides whether permissively licensed code can ever stay with its notices, and rules on contractor code, AI-assisted code and anything received from customers or partners.
Engineering owns detection: the path rules, the scanning tools, the filtered history and the remapped links. Findings that do not fit a rule, such as a file with no license header that resembles a public project, go back to counsel rather than being decided in a pull request. Keeping that boundary clear is what makes the exclusion manifest credible to a buyer's reviewers.
Illustrative: a routing software company builds its exclusion manifest#
Illustrative: a fictional delivery routing software company is preparing to describe its main monorepo and several service repositories in a fit check. The general counsel sets the rule: first-party code only, with third-party code excluded unless approved.
Engineering finds a vendored geospatial library, a fork of an open-source job queue with years of company patches, generated API clients, a committed dependency folder that was later deleted and copied snippets in a few utility files. The fork and its patches are excluded entirely. The deleted folder is removed from history in a filtered copy, and links between Jira issues and commits are remapped.
The result is an exclusion manifest listing each path, category, license detected, decision, method and reviewer. The general counsel approves it before the package scope is shared with anyone outside the company.
How SourceX handles third-party code#
SourceX handles third-party code in the Rights and Preparation steps of the SourceX five-step transaction. The supplier's counsel sets the inclusion rule, the exclusion manifest shows how it was applied, and the supplier approves the result before Delivery.
That record carries into the SourceX Evidence Packet alongside provenance, permitted use, the privacy record and release authorization, so a buyer can see what was left out and why without ever receiving the excluded code.
Frequently asked questions
Can permissively licensed code stay in the package if we keep its notices?
Sometimes, but it is a counsel decision. Keeping notices may satisfy some permissive licenses for redistribution, yet the data license's restrictions and the buyer's intended use still need review. A cautious package excludes permissive code anyway, because its value to the buyer is low and the review effort is not.
What about code our engineers contributed to open-source projects?
Code already contributed upstream is public under that project's license, so the company cannot license it exclusively. Internal copies of those projects mix company changes with upstream code, which is why forks are usually excluded as a whole.
Does excluding third-party code make the package less useful?
Usually not. Buyers want the company's own engineering work: its code, reviews, issues and fixes. Dependency names in manifests can often remain as context, so the buyer still sees which libraries the code relied on without receiving them.
What if a scanner flags a file we believe our team wrote?
Investigate before deciding. Check the file's commit history, its original author and whether similar code exists in a public project from an earlier date. If the company's authorship is clear, record the evidence in the manifest. If it remains uncertain, exclusion is usually the safer choice.
Do we have to tell the buyer what was excluded?
The license agreement will usually include statements about rights in the delivered code, and the exclusion manifest supports them. Sharing a summary of the categories excluded, without the excluded code itself, also helps the buyer understand gaps in the history.
Is code written by contractors treated as third-party code?
Not automatically. Contractor code is first-party when the contract assigns the work to the company. Where assignment is missing or unclear, counsel decides whether to exclude it, which is a separate check from the open-source screen.
Sources
Related resources
- QuestionWho owns enterprise data?
- InsightCan engineering firms sell their data to AI companies?
- InsightAI disruption and software valuations in 2026: does proprietary data help?
- InsightProprietary algorithms in your code: exclude or include?
- SolutionData partnerships between businesses and AI developers
- SolutionTurn the data your company already creates into a licensing asset
See if your company qualifies
A short company assessment. No data uploads are needed.