Software companies
Which repositories to include in a code history license
By SourceX Editorial · Updated
Short answer
A code history license should include repositories your company wrote and owns outright, especially those whose history links commits to issues, reviews and tests: core product, internal services, infrastructure code and test suites. Carve out forks of open-source projects, vendored libraries, code built for specific clients and anything holding secrets or customer data. Decide ownership first, then value.
Key takeaways
- Ownership decides inclusion before value does: a repository the company cannot show it owns stays out.
- Forks, vendored dependencies and generated code add volume but rarely add licensable value.
- Customer-specific code often falls under client contracts and needs its own review before inclusion.
- A monorepo is scoped path by path, with a rule for commits that touch both included and excluded code.
- Every included repository is scanned for secrets across its full history, not just the default branch.
Start from a repository inventory, not memory#
Scoping a code history license starts from a full repository inventory, because most engineering organizations hold more repositories than anyone can list from memory. Export the list from your GitHub organization, GitLab groups or Bitbucket workspace, including private and archived repositories.
For each repository, record its name, purpose, creation date, last activity, main languages, archived status, the issue tracker it links to, and who wrote it: employees, contractors, an acquired team or an outside project. That last column drives most of the include and exclude decisions that follow.
Archived repositories deserve a second look rather than a quick cut. A retired product or a replaced service often carries years of reviewed changes and closed issues, and age alone is no reason to leave it out.
The include and exclude checklist#
The include and exclude checklist sorts repositories by who owns the code and what its history adds. Use the default column as a starting position, then let the rights review confirm or override it repository by repository.
| Repository type | Default | Why | Check before deciding |
|---|---|---|---|
| Core product application | Include | Longest linked history of issues, reviews and releases | Paths that hold vendored or client-specific code |
| Internal services and APIs | Include | Shows how real systems are designed and changed | Hardcoded endpoints and credentials |
| Infrastructure and deployment code | Include with care | Useful for operations and DevOps agents | Hostnames, account IDs and secrets in history |
| Test suites and fixtures | Include | Tests make changes verifiable for evaluation | Fixtures copied from real customer data |
| Internal tools and scripts | Include if linked to tracked work | Shows everyday engineering tasks | Scripts with embedded production access |
| Forks of open-source projects | Exclude | The upstream code is not yours to license | Whether your own patches are worth separating |
| Vendored third-party libraries | Exclude | Third-party code carries its own license terms | Folders that copy dependencies into your tree |
| Customer-specific code | Hold for review | Client contracts may give the customer ownership or confidentiality rights | The master agreement and statement of work for each client |
| Code from acquired companies | Hold for review | Ownership depends on the purchase agreement | Assignment of IP and any carve-outs |
| Generated code and build output | Exclude | Adds volume without showing human decisions | Committed bundles and generated API clients |
| Personal and experiment repositories | Usually exclude | Unclear ownership and little linked history | Whether the work was done for the company |
Ownership questions to settle for every repository#
Ownership is the first gate because a license can only grant rights the company actually holds. A repository the company cannot show it owns stays out of scope, however useful its history looks.
These questions are answered with counsel, and the answers can differ from client to client and from one acquisition to the next.
- Employee IP assignment agreements: do they cover everyone who committed, including founders who wrote code before the company was formed?
- Contractor and agency agreements: do they assign the code to the company, or only grant a license to use it?
- Client contracts and statements of work: does any client own its custom work or treat that code as its confidential information?
- Acquisition documents: did the purchase agreement assign the acquired company's code and history, and were any repositories carved out?
- Open-source licenses: which components came from outside projects, and under what terms?
- Joint development or partner agreements: was any code written together with another company?
How to scope a monorepo#
A monorepo is scoped by path, not as one unit, because a single repository can hold the core product, vendored dependencies, client-specific modules and infrastructure side by side. Draw the scope as a list of included and excluded directories.
Ownership files such as CODEOWNERS help map paths to teams, and dependency folders are usually easy to spot by name. The harder part is history. A commit that touches both an included and an excluded path has to be filtered so only the included changes are delivered, or dropped altogether. Agree that rule before preparation starts, because it changes how the delivery copy is built.
Keep the working repository untouched. Scoping and filtering happen on a separate copy, so day-to-day engineering never depends on licensing decisions.
What makes one included repository worth more than another#
An included repository is worth more when its history connects code changes to the reasons for them. Commits that reference issue keys, pull requests with review comments, continuous integration results and tagged releases let an AI developer follow a change from request to outcome.
Long-lived repositories with many contributors usually rank above short projects, and repositories with meaningful tests rank above those without, because tests let a buyer check whether a proposed change actually works. A small service with clean issue links can outrank a large repository whose history was squashed into a single commit during a past migration.
Secrets and personal details inside included repositories#
Every included repository is scanned for secrets across its full history, because a credential deleted from the current branch still sits in old commits, in other branches and in tags. Revoke and rotate anything found first; cleaning it out of the delivery copy does not make a leaked key safe.
Open-source scanners such as Gitleaks, an MIT-licensed tool that looks for passwords, API keys and tokens in git repositories and files, are a common first pass. Check the status of any scanner you rely on: the Gitleaks README was updated in May 2026 to say the tool is feature complete, that future releases will be security patches only and that its maintainer is shifting focus to Betterleaks. Default rules also miss formats unique to your company, so add patterns for internal token prefixes and connection strings.
Personal details hide in code history too. Commit author names and email addresses, customer records copied into test fixtures, and support emails pasted into code comments are removed or replaced with consistent pseudonyms before delivery.
- Scan every branch and tag, not only the default branch.
- Include configuration files, notebooks and committed environment files in the scan.
- Rotate each live credential before the delivery copy is cleaned.
- Record what was found and removed for the privacy record.
- Re-scan the final delivery copy, not just the working repository.
Illustrative: a logistics software company scopes its GitLab estate#
Illustrative: a fictional transportation management software company keeps its code in GitLab. Its inventory lists the core TMS application, carrier integration services, Terraform modules, a fork of an open-source routing library, EDI adapters built for individual shippers, a driver app from an acquired company and a repository of data science notebooks.
The CTO includes the core application, the integration services, the test suites and, after a full-history secrets scan, the Terraform modules. The routing fork and vendored folders are excluded. The EDI adapters are held: counsel finds that some client agreements give the shipper ownership of custom work and others do not, so only adapters built under the company's standard terms move forward. The driver app is included once the purchase agreement confirms its code was assigned. The notebooks are excluded because they hold extracts of customer shipment data.
The result is a written scope, a list of repositories and paths with an ownership note for each, ready for the rights review.
How SourceX approaches repository scope#
SourceX handles repository scope within the Rights and Preparation steps of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. At the fit check SourceX sees a repository list and descriptive metadata, never code, and nothing leaves company storage until the company approves delivery.
For each package, the SourceX Evidence Packet records provenance, the licensing rights behind each repository, permitted use, the privacy record of what was removed and the release authorization. SourceX's rights in a deidentified dataset are set out in the signed supplier agreement, and the company licenses its history rather than selling its code.
Frequently asked questions
Should archived repositories be included?
Often yes. Archived repositories from retired products or replaced services can hold long, well-linked histories. They go through the same ownership and secrets checks as active ones. The extra step is confirming whether the build still runs, or documenting that it does not, so the buyer knows what can be executed.
Can we license code history without including the current code?
Yes. Scope can be cut by date or by release, for example history up to an older version with recent work kept out, which limits competitive exposure. The license should state which commits and which time range are covered, so both sides know exactly what was delivered.
Do we need consent from former employees who wrote the code?
Usually not for ownership, which generally rests on employment and IP assignment agreements rather than individual consent. Former employees' names and emails in commit metadata are still personal details and are normally replaced during preparation. Contractors and founders who wrote code before incorporation differ, because their agreements may not assign the code, so review both with counsel.
What about repositories that contain copyleft open-source code?
Copyleft components, such as code under the GPL, are usually carved out, or the whole path is excluded, because their license terms travel with the code. Find them with a software composition analysis tool or your dependency manifests, and record what was excluded and why.
Does licensing code history expose our product to competitors?
The license terms manage that risk. Typical protections include permitted use limited to training or evaluation, no redistribution, confidentiality and deletion at the end of the term. Some companies also exclude modules that hold pricing logic or other trade secrets. Settle these limits before the scope is final.
Sources
- Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin. On May 21, 2026, its README was updated to state that Gitleaks is feature complete, that future releases will be security patches only and that the maintainer is shifting focus to Betterleaks. Source
Related resources
- InsightSelling an MEP engineering firm: what buyers value in 2026
- IndustryConstruction data
- QuestionDo AI labs buy code?
- InsightCan you license CAD and engineering drawings to AI companies?
- InsightCan you license code reviews and pull requests to AI companies?
- SolutionEnterprise data: the records of how organizations actually work
See if your company qualifies
A short company assessment. No data uploads are needed.