Software companies
Open-source license audit before licensing your codebase
By SourceX Editorial · Updated
Short answer
An open-source license audit before licensing your codebase finds every file your engineers did not write, sorts it as permissive, copyleft, commercial or unknown, and removes or documents it before delivery. The working rule: code you cannot trace to your own employees or assigned contractors stays out of the package until it is identified and cleared.
Key takeaways
- Scan the delivery copy and its full git history, not only the current dependency manifest.
- Classify every finding as permissive, weak copyleft, strong or network copyleft, commercial, or unknown.
- Treat unknown provenance as a restriction until someone resolves it.
- Exclusion has to reach history too, because vendored code deleted years ago still sits in old commits.
- The audit record belongs in the SourceX Evidence Packet so a buyer can see what was removed and why.
What an open-source license audit answers before a data license#
An open-source license audit before a data license answers one question: which files in the package your company can actually license. A codebase that ships a working product can still contain vendored libraries, copied snippets, generated files and code from acquired companies, and the company may hold no rights to grant in any of them.
This audit differs from the compliance review your team may run before a release. Release reviews ask what ships to customers and whether notices are included. A data license hands over the repository itself, often with years of history, for a use many open-source licenses never contemplated. Whether open-source terms attach to AI training is unsettled, so the safer starting point is to separate first-party code cleanly from everything else.
Step 1: scan the package and its history#
The scan step inventories every third-party component in the repositories you plan to license, including components that were deleted long ago. Point every scan at a dedicated delivery copy so the results match exactly what a buyer would receive.
Software composition analysis tools automate much of this and can export a software bill of materials in a standard format such as SPDX or CycloneDX. Snippet matching and history scanning are the easiest parts to skip, and they are where surprises tend to come from.
- Dependency manifests and lockfiles, to list declared packages and versions.
- Vendored directories, such as vendor, third_party or copied lib folders checked into the repository.
- License files and headers anywhere in the tree, including nested folders.
- Snippet matching against public code, to catch copied functions that carry no header.
- Generated and minified files, such as bundled JavaScript, protocol buffer output and client SDKs.
- Git history, with the same checks run on old commits where removed vendored code still exists.
- Code from acquisitions or contractors, flagged by author, date or repository origin.
Step 2: classify each finding#
Classification sorts each finding into a category that sets its default treatment. Counsel makes the final call on edge cases; the table gives engineering a consistent first pass so every repository is handled the same way.
Why exclude permissive code at all? Permissive licenses grant broad rights, but those rights run from the original authors to everyone, not from you to the buyer. Dropping unmodified dependencies loses little, because public packages are available to anyone, and it keeps the license focused on what your engineers wrote.
Public code with no license at all is not free to reuse. By default its authors keep their rights, so a function copied from a public repository without a license file belongs in the unknown class, not the permissive one.
| Class | Examples | Default action | Record for the audit |
|---|---|---|---|
| Permissive | MIT, BSD, Apache | Exclude unmodified copies; send modified ones to counsel | License text, copyright notices, location |
| Weak copyleft | LGPL, MPL, EPL | Exclude by default, including modified files | Affected files and whether they were changed |
| Strong or network copyleft | GPL, AGPL | Exclude, and check whether first-party code was mixed into the same files | Component, paths and any entangled first-party files |
| Commercial or proprietary | Licensed SDKs, purchased components, partner code | Exclude unless the vendor agreement clearly allows the use | Contract reference and paths |
| Unknown | Files with no header, no match and no clear author; code copied from public repositories that carry no license | Exclude until resolved | Path, author history and why it is unresolved |
Step 3: exclude, then prove the exclusion#
Exclusion means producing a delivery copy in which classified files and their history are gone, then proving it with a second scan. Deleting a vendor folder in a new commit does not count, because the files remain in every earlier commit.
Expect the delivery copy to be noticeably smaller than the original repository. That is a sign the audit worked, not a loss of value: buyers want your engineers' work, not another copy of popular libraries.
- Write the exclusion list as paths and patterns and keep it under version control.
- Rebuild history in the delivery copy with a history-rewriting tool, or export snapshots without history if history is out of scope.
- Remove excluded paths from pull request diffs as well, so review records do not reintroduce them.
- Rerun the full scan, snippet matching included, on the finished copy.
- Compare file lists and hashes between the audit register and the delivery copy.
- Have a second engineer spot-check a sample of first-party files for missed headers.
Step 4: document what was scanned, found and removed#
Documenting the audit turns an internal cleanup into evidence a buyer and its counsel can rely on. The record should let someone who was not involved reproduce what was scanned, what was found and what was removed.
| Record | What it contains |
|---|---|
| Scope statement | Repositories, branches and date range included in the package |
| Tool record | Scanners used, their versions and the date each scan ran |
| Findings register | Every component found, with its class and location |
| Exclusion list | Paths and patterns removed, with the reason for each |
| Open items | Unknown files still under review and who owns each one |
| Sign-off | Engineering reviewer and counsel approval, with dates |
Mistakes that weaken an audit#
The mistakes that weaken an audit usually come from scanning what is convenient rather than what will be delivered. Each one leaves a gap a buyer's counsel is likely to find.
Scanning only the main branch misses long-lived release branches that carry their own vendored fixes. Trusting the dependency manifest misses copied code that was never declared. Stripping license headers to tidy files destroys the evidence of where code came from. And treating code from an acquired company as first-party, without checking the purchase agreement and its contributors' assignments, assumes rights the company may not have.
Illustrative: an insurance agency software vendor audits its repositories#
Illustrative: a fictional vendor of agency management software for commercial insurance brokers plans to license its main application repositories, their pull requests and the linked Jira issues. The CTO runs a composition scan on a delivery copy, then adds history and snippet scans.
The scans find a PDF rendering library vendored years ago and later removed, a rating engine module acquired from a smaller company with no assignment on file, generated API clients, and a handful of utility functions that match public code. The vendored library and generated clients are excluded along with their history. The acquired module goes to counsel and stays out pending the original purchase agreement. The matching utility files are excluded, and the audit record lists each decision with its reason.
How SourceX uses the audit#
SourceX treats an open-source license audit as part of the Rights step of the SourceX five-step transaction. The findings register and exclusion list feed the licensing rights section of the SourceX Evidence Packet, and the scan records support its provenance section. SourceX does not decide license questions for a supplier; it asks for a clear record of what was excluded and who approved the remaining scope.
Frequently asked questions
Do we need a commercial scanning tool?
Not necessarily. Open-source scanners can inventory dependencies and license headers, and many teams pair them with manual review. Commercial composition analysis tools add larger match databases and snippet detection, which matters most for older codebases with copied code. Choose by the size and age of the code rather than by price tier.
What about code written with AI coding assistants?
Code your engineers produced with an assistant is usually treated as first-party work, but note it where you can, because assistants can suggest code that resembles public sources. Snippet matching helps catch close matches. Your assistant vendor's terms and your internal policy on assistant use belong on the open items list for counsel.
If counsel lets us keep modified permissive files, what changes?
Keep the original license texts and copyright notices with those files in the package, and record the modification in the findings register. Removing notices is a common mistake that turns a minor issue into a larger one. Excluding unmodified third-party code entirely, and sending only modified files for review, keeps the number of these cases small.
How long does an audit stay valid?
An audit describes one delivery copy at one point in time. If the package is refreshed with newer commits or additional repositories, rerun the scans on the new material and update the findings register before that release is approved.
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.