Software companies
AI-generated code in your repo: can you license it for AI training?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
You can usually include AI-assisted code in an AI training license, but your rights are strongest in human-authored portions, because US copyright generally requires human authorship. Split the repository by date range or evidence of assistant use, disclose AI-assisted periods to the buyer, and check each coding tool's terms for limits on using its output.
Key takeaways
- Copyright in code generated with little human input is uncertain, so ownership warranties should be narrower for those periods.
- The terms of each coding assistant the team used can restrict how outputs are used; read them before scoping.
- Tool adoption dates, account logs and commit metadata let you divide a repository into periods with different treatment.
- Code reviews and issue discussions stay human-written even in heavily assisted teams, so they often carry the cleanest rights.
Can AI-generated code be licensed for AI training?#
AI-generated code can usually be included in an AI training license, but the company's rights in it may be weaker than in code its engineers wrote themselves. In the US, copyright generally protects human authorship, so code produced by an assistant with little human creative contribution may receive thin protection or none.
That does not make the code unusable. A company can still deliver it under a contract, and confidentiality and trade secret protection do not depend on authorship. What changes is what the company can promise: it should not warrant full copyright ownership of periods dominated by generated code.
How much human selection, arrangement and editing is enough remains unsettled. Treat the issue as a disclosure question rather than a yes or no, and let the buyer decide what it wants with accurate information.
How AI-assisted code differs from human-written code in a license#
AI-assisted code differs from human-written code on ownership, warranties and third-party terms, even when the two look identical in the repository. The table sets out the differences counsel and the CTO should agree on before scoping.
| Question | Human-written code | Heavily AI-generated code |
|---|---|---|
| Copyright ownership | The company owns it through employment or assignment | Protection uncertain where human input was minimal |
| Warranty the company can give | Ownership and non-infringement, often to its knowledge | Narrower: a right to deliver, without an ownership warranty |
| Third-party terms | Employment and contractor agreements | Coding assistant terms of service also apply |
| Risk of reproduced third-party code | Depends on engineering practices | Assistants can reproduce public code, so license scans matter |
| Value as training data | Shows how people reason, write and revise | May be less distinctive, depending on the buyer's goals |
Check the terms of every coding assistant your team used#
Coding assistant terms of service decide who holds rights in outputs and whether outputs may be used for certain purposes. Some vendors state that outputs belong to the user; some AI service terms restrict using outputs to develop or train competing models.
List every assistant used, the plan or account type, and the period of use. Enterprise and individual plans often carry different terms, and engineers sometimes used personal accounts before the company bought licenses. Counsel should read each set of terms and decide whether outputs from that tool can be part of a training license.
Check the data terms in the other direction as well. If an assistant's provider could train on the company's prompts and code during the period of use, parts of that history may already sit in another model, which a buyer looking for distinctive data will want to know. Some plans offer settings that stop the provider training on customer code; confirm which settings applied and when.
How to split a repository by date range#
Splitting a repository by date range means dividing its history into periods with different levels of assistant use, then deciding scope and warranties for each period. The split must rest on evidence, not on recollection.
Keep the split in a simple register: period start and end, tools in use, evidence relied on and the resulting label. The register becomes part of the package documentation and can be extended as the team's practices change.
- Find the adoption date of each coding assistant from procurement records, single sign-on logs or engineering announcements.
- Use commit metadata: co-author trailers, bot accounts and commit message conventions that indicate generated code.
- Check pull request labels, templates or checklists that asked engineers to disclose AI assistance.
- Separate generated boilerplate, such as tests, migrations and scaffolding, from core logic where the history allows.
- Mark each period as pre-assistant, mixed or heavily assisted, and record the reasoning behind each label.
Treatment by period#
Each period gets its own treatment once the split is done. The table shows a common pattern; a buyer's preferences or a tool's terms can move a period into a stricter row.
Where evidence is thin, label the period as possibly assisted rather than claiming certainty in either direction. A cautious label with a clear explanation holds up better in diligence than a confident one that a commit log later contradicts.
| Period | Evidence | Typical treatment |
|---|---|---|
| Pre-assistant | History before any tool adoption date | In scope with standard warranties |
| Mixed | Enterprise tool in use, human review on every change | In scope with disclosure |
| Heavily assisted | Generated files, bot commits, scaffolding | In scope with narrower warranties, or excluded |
| Restricted tool | Outputs from a tool whose terms limit model development | Excluded |
What to disclose to the buyer#
Disclose which periods and components involved AI assistance, which tools were used and how the split was made. Buyers training their own models often care about this, because training on another model's output can raise both contractual and quality concerns.
Put the disclosure in the data description and tie the warranties to it. A license that warrants ownership of everything while the repository holds years of assisted code leaves a gap the buyer's counsel will find.
Disclosure also protects the commercial conversation. A buyer that learns of assisted periods late may reopen terms, while one told at the start can scope around them or focus on the human-written periods and the review history.
Illustrative: an insurance agency software vendor maps its assistant periods#
Illustrative: a fictional software vendor serving insurance agencies wants to license its main repository, code reviews and Jira history. The CTO finds that the team adopted one coding assistant on an enterprise plan, and that a few engineers had earlier used personal accounts on a different tool.
Counsel reviews both sets of terms and excludes the period of personal-account use, because those terms restrict using outputs for model development. The remaining history is split into pre-assistant and assisted periods using single sign-on logs and a pull request label introduced with the enterprise rollout. Both periods stay in scope, with standard warranties for the first and disclosure plus a narrower warranty for the second. Code reviews and issues, written by people throughout, are included across the full range.
How SourceX approaches AI-assisted code#
SourceX treats AI assistance as a provenance fact. In the SourceX five-step transaction it is documented at the Rights stage, where the supplier lists tools, periods and the split method, and the SourceX Evidence Packet records that provenance with the licensing rights, permitted use, privacy record and release authorization, so the buyer knows what it is licensing before Delivery.
Frequently asked questions
Does using an AI assistant mean we lose ownership of our code?
No. Code your engineers wrote or meaningfully shaped remains protectable, and the repository as a whole is still the company's confidential material. The uncertainty is limited to portions generated with little human creative input, which is why a period-by-period view is more accurate than a blanket answer.
Can we detect AI-generated code after the fact?
Not reliably from the code alone. Detection tools are imprecise, so evidence comes from records around the code: adoption dates, account logs, commit metadata and team practices. Those records are also what a buyer's counsel will ask to see.
Does AI assistance increase the risk of open source contamination?
It can, since assistants trained on public code may reproduce recognizable snippets. Run license scans across assisted periods as well as older ones, and apply the same exclusion rules to detected copyleft code that you would apply to code copied by hand.
Should we stop using coding assistants if we plan to license data?
Not necessarily. Teams can keep using them and adopt light disclosure habits, such as pull request labels, that make later provenance records easy to produce. The goal is to know which work was assisted, not to avoid assistance.
Are code reviews of AI-generated code still human-authored?
Usually yes. Review comments, approval decisions and requested changes are written by engineers even when the code under review was generated. That is one reason review histories are often among the cleanest engineering records to license.
Do contractors' AI tools change the analysis?
They can. Contractors may have used their own assistants under their own accounts and terms. Check that contractor agreements assign work product and address AI tool use, and ask major contractors which tools they used during the engagement, so their work can be labeled the same way as the in-house team's.
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.