Code and software engineering data
Who Owns the Code? Ownership Due Diligence for Licensed Codebases
Quick answer
Source code ownership due diligence means proving, module by module, that the company licensing a repository actually holds copyright in what it ships to you. In a typical private codebase, ownership is split across employees, contractors, agency clients, acquired entities and inbound open-source or vendor code. Request the assignment documents, contract excerpts and component inventory that already exist, map ownership per directory, and exclude any part whose chain of title you cannot document rather than walking away from the whole repository.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why code ownership is harder to verify than it looks
Code ownership is hard to verify because one repository can hold work from dozens of authors under different legal relationships, and git history records who committed, not who owns. The supplier's logo on the repository proves possession, not title. AI teams do acquire private commercial codebases, for example to build held-out software engineering benchmarks that public training data has not seen [1], so this review is now a routine step in code data procurement rather than an edge case.
The core US rule is simple: copyright vests in the author, and an employer or commissioning party is treated as the author only for a work made for hire [2]. Everything else moves by assignment, and an assignment of copyright is valid only with a written instrument signed by the owner or an authorized agent [3]. A code license for training or evaluation therefore depends on a paper trail, and the gaps cluster in predictable places. The general chain-of-title framework is covered in Chain of Title for AI Training Data; this page covers the traps specific to software.
The five ownership layers in a private codebase
Most private codebases contain five ownership layers, and each needs different evidence. Treat the list below as a hypothesis to test against the actual repository, with counsel confirming the rules for each jurisdiction where authors worked.
- Employee code. Code written by employees within the scope of employment is generally a work made for hire, owned by the employer [2]. Whether a given author was an employee and whether the work fell within the scope of employment are factual questions [4]. Founders who wrote code before incorporation, interns, and engineers employed by a foreign affiliate are the usual exceptions.
- Contractor code. Independent contractors own what they write unless they sign an assignment. Labeling contractor software as "work made for hire" often does not work for commissioned software and can create employment law side effects in some states, so well-drafted contractor agreements rely on a present assignment instead [5]. A contract that grants the company only a license, not an assignment, means the company cannot sublicense that code for AI use without checking the license scope.
- Client-owned code in agency repositories. Development agencies and systems integrators commonly write code that their client agreements assign to the client, while the agency keeps pre-existing tools and libraries. That client code may sit in the agency's GitHub organization, but the agency usually needs the client's authorization to license it. See Who owns data in a client project? and the software development agencies buyer page.
- Acquired code. Code that came with an acquisition belongs to the buyer only if the deal transferred it. Asset purchases list assigned IP on schedules, and stock purchases leave title with the acquired entity, which may still exist as a subsidiary. Earlier exclusive licenses granted by the target, or code that the target itself never secured from its contractors, travel with the code.
- Inbound third-party code. Open-source dependencies, vendored copies, copied snippets, commercial SDKs and platform customizations (for example Salesforce Apex or SAP ABAP extensions written against vendor frameworks) are owned by third parties under their own terms. Open-source licenses usually permit copying, but conditions such as copyleft travel with the files; see Copyleft Contamination in Licensed Code Datasets.
Evidence to request for each layer
The fastest diligence asks for documents the supplier already keeps, not new legal opinions. Most software companies already have an employee invention assignment template, a contractor master services agreement, client statements of work, M&A disclosure schedules and some form of dependency manifest. Ask for templates plus signature evidence, and accept redacted contract excerpts limited to the IP clauses.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Layer | Typical gap | Evidence to request | If the evidence is missing |
|---|---|---|---|
| Employees | Pre-incorporation founder code; unsigned PIIA for early hires | Invention assignment template, list of authors with signed copies, founder IP assignment to the company | Exclude modules dominated by unassigned authors |
| Contractors | License-only terms; offshore subcontractors not bound | MSA and SOW IP clauses, subcontractor flow-down, signed assignment dates vs commit dates | Exclude contractor-authored paths or obtain a confirmatory assignment |
| Agency clients | Client owns deliverables under the SOW | Client contract IP clause and written client authorization for AI licensing | Exclude the client repository entirely |
| Acquisitions | Stock deal left title in subsidiary; target's own gaps | Asset schedule listing the repositories, IP assignment agreement, entity chart | License from the entity that holds title, or exclude |
| Inbound third-party | Vendored SDKs with no-redistribution terms; copied snippets | SBOM (SPDX or CycloneDX), license scan report, list of commercial SDK licenses | Strip vendor directories and flagged files before delivery |
| AI-assisted code | Unclear authorship of generated lines | Coding assistant policy and tool terms in force during the period | Note in the risk register; counsel decides |
Reading git history as ownership evidence
Git history is the best map of who wrote what, but it is evidence to cross-check, not proof of title. Run git shortlog -sne --all and group author emails by domain: company domains point to employees, agency or contractor domains point to the contractor layer, and personal addresses need a lookup against the HR and vendor rosters. Compare each author's first and last commit dates against the start date of their assignment agreement, because an assignment signed after the work may not reach earlier commits unless it assigns past work expressly.
Also inspect .gitmodules, vendor/, third_party/, node_modules/ committed by mistake, and Co-authored-by: trailers that hide a second author. CODEOWNERS files show who maintains a path today, not who owns it. For repositories delivered with full history, the delivery format matters too; see Delivering Code Repositories with Full History.
Map ownership per directory, then exclude what you cannot clear
Ownership should be recorded per directory or module, because blocking a whole repository over one unclear folder discards most of the value. Build the map from the git authorship analysis plus the evidence table above, assign each path a status, and agree that unclear paths are removed before delivery.
Illustrative example: invented to show structure; it does not describe an available dataset.
repository: billing-platform
history_range: 2017-03-01..2026-06-30
paths:
- path: services/invoicing/
owner_layer: employee
evidence: [piia_template_v3, signed_piia_roster]
status: cleared
- path: integrations/acme-erp/
owner_layer: agency_client
evidence: [sow_2021_ip_clause]
status: excluded # client owns deliverables; no client authorization
- path: mobile/
owner_layer: contractor
evidence: [msa_2019_assignment_clause]
status: cleared_from_2019-09-14 # commits before signature excluded
- path: vendor/payments-sdk/
owner_layer: inbound_commercial
evidence: [sdk_license_no_redistribution]
status: excluded
- path: libs/pdfgen/
owner_layer: acquired
evidence: [asset_schedule_b_item_7]
status: cleared
Two practical rules keep the map honest. Exclusions must remove the files from every commit in the delivered history, not only from the final tree, so ask how the supplier rewrites history (for example with git filter-repo) and verify with a path search across all commits. Excluded code should also be excluded from derived artifacts such as issue-to-fix pairs and review threads that quote it; see Issue-to-Fix Pairs from Private Repositories.
Confidentiality and use limits that survive ownership
Owning the copyright does not end the review, because contracts and confidentiality duties can still restrict AI use. A company can own its code and still be bound by a customer contract that treats customer-specific configuration, schemas or embedded data as confidential. Code can also embed secrets, customer names and personal data in fixtures and comments; scanning for these is covered in Secrets in Code Datasets.
The legal theory for training on copyrighted material is still contested. The US Copyright Office's Part 3 report on generative AI training remains a pre-publication version as of October 2026 [6], so a documented license from the actual owner is the safer foundation than relying on a fair use argument. Non-US buyers should also check export rules on technical code; see Export Controls on Source Code for AI Training.
Questions to put in the data provider questionnaire
A short set of code-specific questions surfaces most ownership gaps before you see a sample. Add these to your data provider due diligence questionnaire:
- Which entity holds copyright in each repository, and has any repository moved between entities through an acquisition, reorganization or asset sale?
- What share of commits, by path, came from non-employees, and do signed assignments cover their full commit date ranges?
- Did any repository contain work delivered to clients, and do you hold written client authorization for this use?
- Have you granted exclusive licenses, source escrow or code ownership to any customer or partner covering this code?
- Can you produce an SBOM and a license scan, and list commercial SDKs whose terms restrict redistribution?
- Which paths will you remove, and how will removal be applied to full history?
Pair these answers with the general AI training data due diligence checklist and the sample review steps in Evaluating a Code Dataset Sample Before You License It.
How SourceX handles ownership review for code
SourceX sources operational datasets, including engineering records, from US companies and manages the commercial process through licensing and ongoing purchases. Every dataset is rights-reviewed for ownership and consents, diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and every release is approved by the supplying company. Data is sourced on request rather than held in stock, so a request does not guarantee a match; you can describe the code data you need on the buyer page. For the categories involved, see proprietary code datasets with full git history, licensing source code for AI training and the code and software engineering data hub.
Request code data with documented ownership
Describe the repositories, languages and history depth you need, and SourceX looks for US businesses that hold that data, reviews ownership and consents, and puts allowed uses, records, term and delivery into a license. Nothing is contracted until a supplier agrees. Start at sourcex.si/buyers.
Sources
- arXiv (Scale AI authors), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- U.S. Government Publishing Office (govinfo), "17 U.S.C. 201 - Ownership of copyright" (2024). https://www.govinfo.gov/content/pkg/USCODE-2024-title17/html/USCODE-2024-title17-chap2-sec201.htm
- Legal Information Institute, Cornell Law School, "17 U.S. Code 204 - Execution of transfers of copyright ownership". https://law.cornell.edu/uscode/text/17/204
- Ninth Circuit Model Civil Jury Instructions, "17.11 Copyright Interests - Work Made for Hire by Employee". https://www.ce9.uscourts.gov/jury-instructions/civil/chapter-17/17-11-copyright-interests-work-made-for-hire-by-employee/
- Farella Braun + Martel, "Employment Law Issues to Consider Before Including Work Made for Hire Clauses in Contractor Agreements". https://www.fbm.com/publications/employment-law-issues-to-consider-before-including-work-made-for-hire-clauses-in-contractor-agreements/
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.