Software companies
Third-party and open source code in repositories you want to license
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Open source and other third-party code in a repository you want to license must be sorted by origin before scoping, because a company can grant only rights it holds. First-party code is usually in scope; vendored SDKs, client-funded work and copyleft components usually are not. Exclude by path and by history, then deliver a file-level origin manifest.
Key takeaways
- Sort every file by origin: employee-written, contractor, client-funded, open source, vendored commercial or generated.
- Client-funded custom code is the category CTOs most often forget, and its ownership sits in the statement of work.
- Copyleft components raise unsettled questions for AI training, so many packages exclude them by default.
- Exclusion must reach commit history, code review diffs and code pasted into tickets, not just the current tree.
- A file-level origin manifest lets the buyer verify scope and lets counsel tie each exclusion to a reason.
Which code in your repositories can you actually license?#
The code you can license is the code your company owns outright, plus anything whose license terms clearly permit the use you are granting. Most software company repositories mix origins: code employees wrote, code contractors wrote, open source pulled in as dependencies or copied as snippets, commercial SDKs checked in for convenience and sometimes code a customer paid for and owns.
A data license for AI training is a grant of rights by the supplier. If a file belongs to someone else, the supplier cannot grant rights in it simply because it sits in the company's GitHub organization. Origin, not location, decides scope.
Which license terms apply to your code, and how they apply to AI training, is assessed with counsel for each deal. What follows is the rights map and the exclusion step that make that review possible.
A rights map by code origin#
A rights map sorts every directory, and eventually every file, into one origin. The positions below are typical starting points, not conclusions about your code.
| Code origin | Typical rights position | Default treatment | What to check |
|---|---|---|---|
| Employee-written first-party code | Owned through employment and invention assignment terms | In scope | Signed invention assignments, especially for founders and early staff |
| Contractor or agency code | Owned only if the contract assigns IP to the company | In scope if assigned | IP assignment clauses in each contractor agreement |
| Client-funded custom development | Often owned by, or restricted to, the client | Exclude unless the contract says otherwise | Statement of work and development agreement IP terms |
| Permissive open source (MIT, BSD, Apache) | Owned by its authors; broad license with notice conditions | Exclude or mark; include only if counsel agrees | License files, notice requirements, modified copies |
| Copyleft open source (GPL, LGPL, AGPL) | Owned by its authors; conditions attach to distribution | Exclude by default | Whether delivering the dataset counts as distribution |
| Vendored commercial SDKs and libraries | Licensed to you for use, rarely for redistribution | Exclude | The vendor license and any redistribution limits |
| Code from acquired companies | Depends on the acquisition documents and the target's own contracts | Review separately | Purchase agreement IP schedule and legacy contractor terms |
| Copied snippets and assistant-generated code | Mixed and often unclear | Flag and review high-risk directories | Comments, commit messages and tool adoption dates |
Why copyleft code raises open questions#
Copyleft code raises open questions because its obligations generally attach when the code is distributed (and, for AGPL, when modified code is offered to users over a network), and it is not settled whether training a model on code, or delivering a code dataset to a buyer, counts as distribution in the relevant sense. Delivering files to another company does put copies in that company's hands, which some readings treat as enough to trigger notice or source obligations.
Rather than resolve that question deal by deal, most suppliers and buyers keep GPL, AGPL and similar components out of the package. Permissive licenses are less fraught but still carry conditions, such as keeping copyright and license notices with copies, which a de-identification pass that strips comments can accidentally break.
Client code is the category that surprises CTOs#
Client code surprises CTOs because it often lives in ordinary-looking repositories. Vertical software companies frequently build integrations, custom reports or whole modules under a statement of work, and the development agreement may assign that code to the customer or restrict its use to that customer.
Look for repositories named after customers, branches created for a single implementation, directories of customer-specific scripts and forks of a customer's own codebase that your team maintained. Implementation teams also keep migration scripts that embed a customer's schema and business rules, which can be confidential information even with no personal data present.
Contractor code has the mirror problem. If an early contractor never signed an IP assignment, the company may hold only an implied license to use that code, which is weaker than ownership. Counsel decides whether a confirmatory assignment is needed before the code is included.
The exclusion step, in order#
The exclusion step runs on a separate export copy, never on production repositories. Rewriting history matters because a file deleted today still sits in every older commit, so a package that ships full history without that rewrite ships the excluded code too.
- Run software composition analysis on each repository to list open source components, versions and licenses.
- Mark vendored directories, committed dependency folders and copied third-party files by path.
- Map repositories and branches to customer contracts and flag client-funded work.
- Confirm IP assignment for every contractor who committed code, using commit author history as the list.
- Remove excluded paths from the working tree and from history in the export copy.
- Strip the same code from code review diffs, issue comments and support tickets where engineers pasted it.
- Scan the export for secrets with a tool such as gitleaks or TruffleHog, and rotate anything live you find. Gitleaks is now feature complete and receives security patches only, so confirm the scanner you choose is maintained.
- Produce a file-level origin manifest and have engineering and counsel sign it.
Illustrative: a fleet maintenance software vendor sorts its monorepo#
Illustrative: a fictional fleet maintenance software company has a GitHub monorepo with first-party services, a vendored mapping SDK licensed from a commercial provider, a GPL-licensed PDF library copied into a utilities folder, and two separate repositories built under development agreements for a large trucking customer.
Software composition analysis finds the PDF library and several permissive packages. The CTO excludes the mapping SDK, the GPL library and both customer repositories, removes them from history in an export copy and strips matching diffs from pull request reviews. One early contractor's commits are held back until counsel confirms an assignment.
The resulting package covers first-party services and their code review history, with a manifest listing every file by origin. Counsel signs the manifest, and the CEO approves the scope.
What to represent about third-party code#
Representations about third-party code should match the manifest and go no further. A warranty that the package contains no third-party code at all promises more than most scans can prove; a warranty that listed origins were excluded using a described process is easier to stand behind.
Buyers often run their own scans, so differences between your manifest and their results will surface. Agree in advance how a discovered component is handled and how indemnity applies; the allocation of open source risk is a common point of negotiation with counsel on both sides.
| Topic | Harder to stand behind | Easier to stand behind |
|---|---|---|
| Third-party code | The package contains no third-party code | Listed origins were excluded using the described process |
| Ownership | The company owns all code in the package | The company owns or holds an assignment for each file in the manifest |
| Open source | No open source license applies to the package | Known components were identified by scan and removed |
| Remedy | Open-ended indemnity for any match | Removal, corrected delivery, then agreed indemnity terms |
How SourceX handles third-party code#
SourceX treats code origin as part of the Rights step in the SourceX five-step transaction. Before any code is prepared, the supplier and its counsel agree which origins are in scope, and the Preparation step applies exclusions by path and by history.
The SourceX Evidence Packet records provenance and licensing rights for the package, including the origin manifest and the exclusion method, so permitted use is tied to code the supplier actually holds. The supplier approves the final scope before Delivery.
Frequently asked questions
Can we include permissive open source if we keep the notices?
Possibly, if counsel agrees and the buyer wants it, but many packages leave it out anyway. Buyers can usually obtain public open source directly, so it adds little that is unique while adding notice obligations and scan noise. Your own code that calls those libraries stays in scope.
Does deleting a vendored folder remove it from the dataset?
Not if the package includes history. Every earlier commit still contains the folder, and code review diffs may show it too. Remove the path from history in a separate export copy and confirm the result with a scan of the exported repository before delivery.
What about code our engineers wrote with AI coding assistants?
Treat it as first-party for scoping, but flag directories where assistant use was heavy if you can identify them, and ask counsel about it, because copyright in largely machine-generated code is unsettled. Buyers increasingly ask whether code was human-written, because human-generated signal is part of what they license. Note the dates assistant tools were adopted.
What if we cannot tell where some files came from?
Exclude them or hold them back until their origin is confirmed. Files of unknown origin are common in older repositories and in acquired codebases. A smaller package with a clean manifest is easier to approve than a larger one with unexplained files.
Do buyers run their own open source scans?
Many do, often before accepting delivery. Share your manifest and scan method up front so differences can be resolved quickly, and agree in the license how a discovered third-party component is handled rather than leaving it to a dispute after delivery.
Sources
- Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin. Source
- The gitleaks README states that Gitleaks is feature complete, that future releases will be security patches only, and that the maintainer is shifting focus to Betterleaks. Source
- TruffleHog is an AGPL-3.0 open-source secret scanner that scans sources including Git, chats, wikis, logs, object stores and filesystems. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.