Skip to content

AI data market

Is training AI on copyrighted data fair use? Where US courts stand in 2026

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Whether training AI on copyrighted data is fair use is still unsettled in US courts in 2026: each case turns on its facts, weighed through the four fair use factors. Courts have separated training from how copies were obtained and given real weight to market harm. For private business records, licenses and contracts decide access, not fair use.

Key takeaways

  • No US court has adopted a rule that makes all AI training fair use or all of it infringing.
  • Courts weigh four factors, and the fourth, harm to the market for the work, has drawn the most argument.
  • How a developer obtained copies is treated as a separate question from how it used them.
  • Trial-court decisions bind only the parties, while appellate rulings carry weight across their circuit.
  • Fair use is a defense to copying published works and gives no one access to private company records.

Is AI training fair use?#

AI training is neither automatically fair use nor automatically infringement under US law; courts decide each case on its facts using the four fair use factors in the Copyright Act. Results so far have differed with the type of work, the way it was obtained, what the trained system does and the evidence offered about market harm.

That uncertainty is the main point for business leaders. A developer relying on fair use accepts litigation risk for each body of material it uses, while a developer using licensed material has a contract that states what it may do.

How the four fair use factors apply to AI training#

The four factors are weighed together rather than scored one by one, and in AI cases the first and fourth have done most of the work. The table shows what courts ask under each factor and how the argument has generally run when the use is AI training.

The phrase market for licensing training data matters for anyone who owns records. The more often material is licensed for AI use, the easier it becomes for a rights holder to argue that unlicensed copying takes away a real market rather than creating a new use.

How the four fair use factors apply to AI training
FactorWhat courts askHow it has played out in AI cases
Purpose and character of the useIs the use transformative, and is it commercial?Training is often described as transformative, but less so where the tool competes directly with the source
Nature of the workIs the work creative or mostly factual?Usually given less weight; factual works receive thinner protection
Amount usedHow much was copied, and was that reasonable for the purpose?Whole works are copied for training, and courts weigh that against the stated purpose
Effect on the marketDoes the use harm the market for the work, including licensing?The most contested factor, including whether a market for licensing training data exists

Where US courts stand in 2026#

US courts in 2026 have split on outcomes but agree that facts decide: the cases that have produced rulings involve published books and editorial content, not private business records. The table summarizes the most-cited decisions; read the opinions and current counsel commentary before relying on any one of them.

District-court rulings bind only the parties and can be appealed, so headlines about a win or loss often overstate how far a ruling reaches. The US Copyright Office added a non-binding view in its May 9, 2025 pre-publication report on generative AI training: it said commercial use of vast troves of works to produce competing content, especially through illegal access, goes beyond established fair use boundaries, and it recommended letting voluntary licensing markets keep developing. Many other cases over news, music, images and code remain pending.

Where US courts stand in 2026
CaseCourt and dateTraining holdingAcquisition holdingMeaning for licensed private data
Bartz v. AnthropicN.D. Cal., Judge Alsup, June 23, 2025Training on books was exceedingly transformative and fair use; scanning purchased print books was fair useDownloading pirated copies into a central library was not fair use; the claims settled for $1.5 billion, with final approval on July 20, 2026How material was obtained is judged separately from use
Kadrey v. MetaN.D. Cal., Judge Chhabria, June 25, 2025Fair use on the record the authors presented; the court stressed it was not a general rulingNot the basis of the ruling; market-harm evidence was found insufficientMarket-harm arguments rise or fall on evidence
Thomson Reuters v. RossD. Del., February 11, 2025; Third Circuit affirmed, reported September 2026Not fair use: a non-generative tool built as a market substitute for WestlawNot a separate holdingCopying to build a competing product is hard to defend

Acquisition and use are now separate questions#

Courts have begun to treat how a developer obtained material as a separate question from how it used the material. A use that might be defended as fair can sit alongside a claim over copies downloaded from unauthorized sources, and reported settlements in the book cases centered on exactly those copies.

For buyers, the split shifts attention to sourcing: where each dataset came from and what permission covered it. For suppliers, it means documentation of rights is part of what makes a dataset usable, alongside the content itself.

Why fair use rarely decides private business records#

Fair use rarely decides questions about private business records because those records were never published. A developer cannot lawfully collect a company's support tickets, dispatch notes or code reviews from the open web, so the practical question is access, and only the company can grant it.

Other laws also protect operational records. Trade secret law, confidentiality agreements with customers and employees, privacy laws such as CCPA where they apply, and vendor terms all shape what can be shared, and none of them is resolved by a fair use ruling about books or news articles.

Much of what makes operational records useful is factual: part numbers, timestamps, defect codes and delivery windows. Copyright protects expression, not facts, which is another reason contracts rather than copyright usually define these deals.

What founders and general counsel can do now#

Founders and general counsel can act now without waiting for the law to settle, because the most useful steps protect the company whatever courts decide. Each step below either preserves a protection or documents a right.

  • Treat public website content, help centers and documentation as already exposed to crawling, and decide whether to state AI terms for it.
  • Keep private records private through access controls, confidentiality terms and clear internal policies.
  • License records only under a written agreement that defines permitted use, term, deletion and audit.
  • Document provenance and rights for anything you license, before the buyer asks.
  • Revisit your position when appellate courts rule, since their decisions can change the risk picture for developers.

Illustrative: an architecture firm separates public from private#

Illustrative: the managing principal of a fictional architecture firm reads that courts have sided with some AI developers on fair use and asks two questions. Were the firm's public project pages and design essays used to train models, and does that mean its private project records are fair game too?

Counsel separates the answers. The public pages may well have been crawled, and the firm can add AI terms and crawler controls going forward, with limited leverage over past copying. The private records, years of RFIs, submittal reviews and internal design review notes in Procore and Bluebeam, were never public, and nobody can use them without the firm's permission.

The firm updates its website terms, leaves the fair use question with counsel, and runs a metadata-only fit check on its RFI and submittal history. Client-owned drawings and deliverables are flagged for exclusion from the start.

How SourceX approaches fair use questions#

SourceX does not rely on fair use for anything it handles. Every package is a licensed transaction under a written agreement, records are licensed rather than sold outright, and the SourceX Evidence Packet documents provenance, licensing rights, permitted use, the privacy record and release authorization.

Within the SourceX Enterprise Data Value Framework, reproducibility reduces value and rights increase it. Records anyone could scrape are reproducible; private operational records with documented rights are not, whatever courts decide about published works.

Frequently asked questions

Does a fair use win for one developer protect other developers?

Not directly. A trial-court ruling binds only the parties, and its reasoning may or may not persuade another court with different facts. Appellate rulings bind lower courts within that circuit but not elsewhere. Each developer's use is judged on its own record, including how it obtained material and what its system does.

Could the Supreme Court settle the question?

It could, if a case reaches it and the Court agrees to hear it, but there is no way to predict timing or outcome. Until then, answers will come from a growing set of trial and appellate decisions, and possibly legislation. Businesses are better served planning around uncertainty than waiting for a final answer.

If AI training is fair use, can a company still charge for its data?

Yes. Fair use, where it applies, covers copying of published works a developer can already reach. It does not provide access to private records, preparation of those records, warranties about rights or continuing supply. Those are what a license delivers, which is why licensing continues however the fair use cases resolve.

Does copyright protect our support tickets and internal documents?

Some of their content may be protected. Written explanations, design notes and replies can contain protected expression, while timestamps, part numbers and defect codes are facts. In practice, confidentiality, trade secret law and contracts often protect operational records more than copyright does, and counsel can assess which protections apply to each record family.

Should we block AI crawlers on our website?

That is a business decision. Blocking crawlers and adding AI terms may limit future collection of public pages by developers that honor those signals, but it does not undo past crawling. Weigh what the public content does for marketing and customer support, and keep private records behind access controls either way.

Sources

  • On June 23, 2025, Judge William Alsup held in Bartz v. Anthropic that using books to train large language models was exceedingly transformative and therefore fair use. Source
  • Judge Alsup also held that scanning purchased print books, with the originals discarded, was fair use. Source
  • Judge Alsup ruled that downloading pirated books to build a permanent central library was not fair use and left that claim for trial. Source
  • On July 20, 2026, the court granted final approval of the $1.5 billion Bartz v. Anthropic class settlement. Source
  • On June 25, 2025, Judge Vince Chhabria granted Meta partial summary judgment in Kadrey v. Meta, holding training on the plaintiffs' books was fair use on the record presented. Source
  • The Kadrey ruling turned on market harm: the authors had not produced enough evidence of harm, and the court rejected lost AI licensing revenue as a cognizable harm. Source
  • On February 11, 2025, Judge Bibas granted partial summary judgment to Thomson Reuters and rejected Ross Intelligence's fair use defense. Source
  • In late September 2026 the Third Circuit affirmed that Ross's copying of Westlaw headnotes to train its AI was not fair use. Source
  • On May 9, 2025, the US Copyright Office released a pre-publication version of its report on generative AI training. Source
  • The report states that commercial use of vast troves of copyrighted works to produce competing content, especially through illegal access, goes beyond established fair use boundaries, and recommends letting voluntary licensing markets develop. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify