Skip to content

Software companies

Code copyright and AI training: where the law stands in 2026

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Code copyright and AI training law in 2026 is settled on the basics and open on the biggest question. Source code is protected expression, and owners can license it for training by contract. Whether training on code without permission is fair use remains unsettled in US courts. A documented license avoids depending on that answer.

Key takeaways

  • Source code is protected by copyright as original expression, although its functional elements receive thinner protection.
  • Whether training a model on copyrighted code without permission is fair use remains unsettled and fact-specific in US courts.
  • A license replaces the fair use question with contract terms covering permitted use, output controls, attribution, term and deletion.
  • Before licensing, confirm chain of title across employee, contractor, acquired, open-source and customer-funded code.
  • Code written mainly by AI tools may carry little or no copyright protection, because US copyright generally requires human authorship.

Code copyright in 2026 rests on a stable base with an unfinished top floor. Copyright in source code, an owner's right to license it, and the human authorship requirement are well established. The legality of unlicensed training, the treatment of model outputs that echo training code, and the reach of open-source terms are still being argued.

The table gives a working status for each question at publication in October 2026. Use it to orient a conversation with counsel, not as a ruling on your facts, and check for newer decisions before relying on it: a single ruling can move any row marked unsettled.

Which code copyright questions are settled in 2026?
QuestionStatus as of October 2026What it means for a code owner
Is source code protected by copyright?Settled: yes, as original expressionYour repositories are a protectable asset, subject to chain of title.
Can an owner license code for AI training by contract?Settled: yesScope, permitted use and restrictions are whatever the contract says.
Is training on code without permission fair use?UnsettledTrial-court decisions so far turn on facts such as how copies were obtained and market harm, they do not all agree, and there is no nationwide rule.
Can model outputs that reproduce training code infringe?UnsettledRisk rises with close, lengthy copying; licenses often add output controls.
Is code written mainly by AI tools protected?Settled in principle, unsettled at the edgesHuman authorship is required; mixed human and AI work is judged case by case.
Do open-source licenses permit training?UnsettledThe answer may turn on conditions such as attribution and copyleft; a supplier can carve these files out of a training license to avoid the question.
Do Copyright Office reports bind courts?Settled: noThey are persuasive analysis, not law, though judges and lawyers read them closely.

How does fair use apply to training on code?#

Fair use is the US doctrine that lets a court excuse some unlicensed copying after weighing four factors: the purpose and character of the use, the nature of the work, the amount used, and the effect on the market for the work. No single factor decides a case, which is why rulings on AI training have not lined up neatly.

Code sits in an unusual place on the second factor. Courts have long treated the functional parts of software, such as algorithms, methods of operation and structure dictated by efficiency, as unprotected or only thinly protected, while comments, naming, organization and creative design choices carry more protection. A training set built from full repositories copies both kinds at once.

The fourth factor is where licensing enters the argument. Disputes over market harm often turn on whether a real market exists for licensing training data, and a growing record of documented licenses is part of that debate. For a code owner, the practical point is simple: a licensed use never has to survive the four-factor test.

The US Copyright Office has studied AI and copyright, including the use of copyrighted works to train models, and has published its analysis in reports. That analysis treats training as fact-specific rather than giving a blanket yes or no: some training uses may qualify as fair use and others may not, with the purpose of the resulting model and the effect on markets for the original works weighing heavily.

Copyright Office reports do not bind courts. They matter because judges, litigants and legislators read them, and because they shape the questions buyers ask in diligence. A general counsel should read the current text directly rather than rely on summaries, and should check whether later guidance has replaced it.

Outside the US the rules differ. In the EU, the AI Act's obligations for providers of general-purpose models, including transparency about training content, and copyright rules on text and data mining may apply to model developers that serve European markets, and buyers there may ask suppliers for more provenance detail. Which laws apply to a given license is assessed deal by deal with counsel.

Because the law is moving, keep a short watch list and revisit any license still in its term when one of these items changes. The developments most likely to affect a code owner are listed below.

  • Appeals court rulings on whether training on lawfully acquired copies is fair use.
  • Decisions on whether outputs that reproduce training code infringe, and who is liable for them.
  • Cases testing whether open-source conditions such as attribution and copyleft reach model training.
  • New or revised Copyright Office guidance, and any federal or state bills on training transparency.
  • EU guidance on text and data mining opt-outs and on summaries of training content.

Why a license takes the fair use question off the table#

A license takes the fair use question off the table for the licensed code because the buyer's use rests on permission rather than on a defense. The parties decide in writing what the buyer may do, for how long and under what controls, so no court needs to weigh the four factors for that use.

That is why licensed code with documented rights can appeal to model developers that want to reduce litigation exposure. For the owner, the contract is also where the real protections live, and these clauses deserve a general counsel's close reading.

  • Permitted use: training, evaluation or both, and whether fine-tuning a named product is included.
  • Output controls: commitments to reduce verbatim reproduction of licensed code in model outputs.
  • Exclusions: repositories, directories or files that stay out, such as third-party and customer-owned code.
  • Attribution and confidentiality: whether the source can be named and how the code is secured in storage.
  • Term, deletion and trained models: what happens to the data and to models already trained when the license ends.
  • Warranties and indemnities: what each side promises about title and who bears third-party claims.

Do you own the code you would license?#

Ownership of a codebase is rarely as simple as the company name on the GitHub organization. Chain of title runs through employment agreements, contractor terms, acquisition documents and customer contracts, and a gap in any of them can carve files out of a license.

Most companies find a few gaps. The usual fix is to exclude the affected files or obtain a confirmatory assignment before scoping, rather than stall the whole license.

Do you own the code you would license?
Code sourceOwnership questionWhat to check
Employee-written codeWas it created within the scope of employment?Invention assignment and confidentiality agreements
Contractor and agency codeDid the contract assign IP to the company?Contractor agreements and statements of work; a missing assignment may leave only a license
Acquired codebasesDid title transfer cleanly at closing?Asset purchase schedules and any carve-outs for licensed components
Customer-funded custom workDoes the customer own the deliverables?Professional services terms and work-for-hire clauses
Open-source componentsWhich license applies and what does it require?Dependency manifests, vendored folders and license files
AI-assisted codeHow much human authorship is there?Tool usage policies, commit history and review records

Illustrative: a field service software company weighs a code license#

Illustrative: a fictional field service software company with a long-lived Java and TypeScript codebase is asked whether its repositories and code reviews could be licensed to train a coding assistant. Its general counsel starts with chain of title, not with fair use.

The review finds that an early mobile app was built by an outside agency under a contract with no IP assignment, that one directory vendors a copyleft library, and that a large customer funded a scheduling module under a services agreement assigning deliverables to the customer. The company excludes the copyleft directory and the customer-funded module, holds the agency-built app out of scope until the agency signs a confirmatory assignment, and documents the remaining code with its open-source inventory.

The resulting scope is smaller but clean. The counsel's memo to the CEO notes that the license does not depend on how courts eventually rule on unlicensed training, because the buyer's use rests on the contract.

How SourceX approaches code rights#

SourceX handles code rights in the Rights step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. The initial fit check uses metadata only, so no repository is shared while ownership questions are still open.

For each code package that proceeds, the SourceX Evidence Packet records provenance, licensing rights, permitted use, the privacy record and release authorization. The supplier approves every step, and code is licensed, not sold, so the company keeps ownership of its repositories.

Frequently asked questions

Does a pending AI copyright lawsuit affect a license we sign now?

A pending case can change the background law, but it does not usually undo a signed license, because the buyer's use rests on your permission. Contracts can still address change-in-law risk through review rights or termination triggers. Ask counsel how your agreement should handle new rulings during its term.

Can we license code we once published as open source?

You can usually license your own copyright again on different terms, but anyone who received the code under the open-source license keeps those rights. A training license for previously public code may carry less value, since a buyer may already have access. Check whether outside contributors added code you do not own.

Does a training license let the buyer ship our code?

Not unless the contract says so. A training license usually limits use to training and evaluation, excludes redistribution, and may require controls against reproducing licensed code in outputs. Read the permitted use and output clauses together, because together they define what can leave the buyer's environment.

Should we register copyrights before licensing code?

Registration is not required for copyright to exist. In the US it can matter if you ever need to enforce your rights in court, so some companies register key releases. Most training licenses rely on contract terms rather than infringement claims, so ask counsel whether registration is worth the effort for your codebase.

Do customer no-AI-training clauses restrict licensing our own code?

Usually those clauses cover customer data, not your source code, but wording varies. Some enterprise agreements define confidential information broadly or restrict anything derived from the customer's environment. Review clauses that mention customer-specific configurations, scripts or code written for one customer before including them in scope.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify