Skip to content

Rights and contracts

What the 2025 AI fair-use rulings mean for companies licensing data

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

The 2025 AI fair-use rulings, Thomson Reuters v. Ross, Bartz v. Anthropic and Kadrey v. Meta, were trial-court decisions about copying published works, and they reached different results. None gives anyone access to your private support tickets, CRM histories or code. For companies licensing data, they mainly raise the value of documented provenance and precise permitted-use terms.

Key takeaways

  • Ross rejected fair use for a competing tool, while Bartz and Kadrey found training on books fair on their records, with limits.
  • Fair use is a defense to copying works someone already has; it does not authorize access to private business systems.
  • Bartz separated lawful training from a pirated library, and the case later settled with destruction of the pirated files.
  • Kadrey rejected lost licensing fees as market harm, so suppliers should not rely on that argument to protect their records.
  • Contract terms, not copyright doctrine, remain a supplier's main control over how licensed data is used.

What did the 2025 AI fair-use rulings decide?#

The 2025 AI fair-use rulings were three US trial-court decisions on whether copying published works to build AI tools was fair use, and they reached different results on different records. Thomson Reuters v. Ross Intelligence rejected fair use for a competing legal-research tool, while Bartz v. Anthropic and Kadrey v. Meta Platforms found training on books to be fair use on the records before those courts, with important limits.

Each ruling bound only its parties, and the judges weighed the same factors differently. A general counsel should treat them as signals of how courts approach the questions, not settled law, and should read the opinions or ask counsel for case-level detail before relying on any one.

What did the 2025 AI fair-use rulings decide?
CaseCourt and dateHoldingWhat it turned on
Thomson Reuters v. Ross IntelligenceD. Del., February 11, 2025, Judge Bibas sitting by designationCopying 2,243 Westlaw headnotes to build a competing legal-research tool was not fair useRoss meant to build a market substitute; the tool was not generative AI
Bartz v. AnthropicN.D. Cal., June 23, 2025, Judge AlsupTraining on lawfully acquired books was fair use; keeping pirated copies in a central library was notHow the training copies were obtained
Kadrey v. Meta PlatformsN.D. Cal., June 25, 2025, Judge ChhabriaPartial summary judgment for Meta: training Llama models on the plaintiffs' books was fair use on the record presentedThe authors' failure to prove market harm; the court stressed the ruling was narrow

What has happened since the rulings?#

Since the rulings, one case has reached an appellate court and another has settled, which changes how the 2025 decisions should be read. In late September 2026, a Third Circuit panel affirmed in Ross that the headnotes are copyrightable and that copying them to train a legal-research AI was not fair use; LawNext reported it as the first federal appellate decision on fair use in AI training. Because Ross involved a non-generative tool, how far it reaches generative model training is still debated.

Bartz did not go to trial on the pirated library. The parties agreed to a $1.5 billion class settlement that received final approval on July 20, 2026. It requires destruction of the downloaded pirated files and copies derived from them, and its release covers only past conduct through August 25, 2025, not claims about AI outputs or future conduct. A settlement is not a ruling on fair use, but it shows what unlawfully acquired training material can cost.

Many other AI copyright suits are still moving through the courts, and decisions on other kinds of works, on market dilution and on licensing markets could shift the picture again. For operating companies those developments matter mostly at the margin, for the reasons below.

Where the rulings left questions open: licensing markets and pirated copies#

The 2025 rulings left two questions open that matter to data suppliers: whether lost licensing fees count as market harm, and how much the source of the training copies matters.

On licensing, Judge Chhabria in Kadrey rejected the argument that lost fees for licensing books as AI training data were a cognizable market harm, calling that reasoning circular. The Copyright Office's May 2025 report took the opposite view where a licensing market exists or is likely, and in Ross the market-effect factor favored Thomson Reuters because Ross meant to build a substitute product. Judge Chhabria also wrote that it is hard to imagine fair use where copying enables a potentially endless stream of competing works, a signal rather than a holding that plaintiffs with better evidence of market dilution could win.

On sources, Judge Alsup held that building a central library from pirated books was not fair use even though training itself was. In Kadrey, Meta acknowledged downloading books from shadow libraries, and the court granted Meta only partial summary judgment on the training claim on the record presented; the ruling did not resolve every claim in the case. Because the courts did not apply one rule on acquisition, developers increasingly ask suppliers for documented, lawful provenance instead of relying on a defense.

The fair-use factors, read from a data supplier's seat#

The four fair-use factors are the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the potential market. Read from a data supplier's seat, each factor maps to a practical question about operational records.

The second factor deserves attention. Facts such as a part number, a delivery time or a defect code are not protected by copyright, while the expression in a support agent's reply or an engineer's design note can be. Much of the value of operational data sits in that mix, and its protection often comes more from confidentiality and contract than from copyright.

The fair-use factors, read from a data supplier's seat
FactorWhat courts weighWhat it means for private operational data
Purpose and characterWhether the use is transformative and whether it is commercialTraining is often argued to be transformative, but access to private records still requires permission
Nature of the workCreative works get stronger protection than factual onesBusiness records mix unprotected facts with written expression in emails, tickets and documents
Amount usedHow much was copied relative to the purposeTraining usually uses whole records, so this factor rarely settles the question alone
Market effectHarm to sales of the work or to licensing markets for itKadrey rejected lost licensing fees as harm, so protect value through access control and contract rather than doctrine

Three implications for companies licensing data#

The rulings carry three practical implications for operating companies that license data, and none of them make licensing less relevant.

Confidentiality is often the stronger shield. Records kept confidential and released only under contract are better placed to keep trade secret protection where it applies than records that have been published or shared loosely.

  • Fair use does not open your systems. The doctrine can excuse certain copying of works someone already has; it does not give a developer access to your Zendesk instance, Salesforce history or private GitHub repositories. Private records reach developers through a license, and unauthorized access raises separate legal issues.
  • Provenance matters more. Courts looked at how training copies were obtained, and the Bartz settlement required destruction of pirated files, so developers ask suppliers to document where data came from, who held rights in it and what was removed before delivery.
  • Contracts carry the controls. Copyright doctrine will not limit retention, model release or downstream use of your licensed records; the license terms do, so permitted use, deletion and distribution clauses deserve careful drafting.

Which license terms matter more now?#

The license terms that matter more after the rulings are the ones that describe where data came from and limit what happens to it after delivery. Because copyright doctrine leaves so much to case-by-case judgment, a supplier gets certainty by writing its boundaries into the agreement.

Which license terms matter more now?
License termWhy the rulings raise its importance
Provenance representationsDevelopers want evidence that training data was lawfully obtained and authorized for release
Permitted useNames training, evaluation or both, so the buyer's rights do not depend on a fair-use argument
No redistribution of the datasetKeeps prepared records from becoming public material others could copy
Model and output limitsAddresses model release and outputs that could compete with your own services
Deletion or return at term endSets an end point that does not depend on how copyright law develops

Illustrative: a distributor's general counsel reads the headlines#

Illustrative: the general counsel of a fictional industrial distributor reads coverage of the fair-use rulings and asks whether the company's records could simply be used for training without its involvement. The records in question are years of order exceptions, supplier delay notes and customer service resolutions stored in Epicor and a shared helpdesk.

Counsel concludes that the rulings do not reach these records, because no developer can access them without the company's permission. The decision shifts to whether to license them at all. The company chooses a narrow, non-exclusive license covering prepared exception records with customer names removed, a permitted-use clause limited to model training and evaluation, no redistribution of the dataset, and deletion at the end of the term.

How SourceX approaches rights after the rulings#

SourceX treats every package as a licensed transaction rather than a fair-use question. Each package goes through the Rights step of the SourceX five-step transaction, Supply, Rights, Preparation, Approval and Delivery, where ownership, contract restrictions and privacy limits are reviewed with the supplier.

For each approved package, a SourceX Evidence Packet sets out provenance, licensing rights, permitted use, the privacy record and release authorization. That documentation answers the provenance questions the rulings brought forward, and the supplier approves each step before anything is delivered.

Frequently asked questions

Do the rulings mean AI companies can use any data they find?

No. The decisions addressed specific facts about published works, and outcomes differed between cases. They do not authorize access to private systems, override contracts or displace privacy law. Even for published material, the Bartz court held that a library of pirated copies fell outside the defense, and later courts may weigh acquisition differently.

Are our internal records protected by copyright at all?

Partly. Facts and data points are not protected, but written expression in emails, tickets, documents and code usually is, and the company typically owns what employees wrote in their jobs. Messages written by customers may belong to the customers. Confidentiality and trade secret law often protect business records more effectively than copyright.

Should we wait for appeals before licensing data?

Waiting rarely helps a private-records supplier, because your records are not reachable without a deal either way. Appeals mainly affect questions about published content; the Third Circuit's September 2026 Ross decision was the first appellate ruling, and more will follow. The decision to license rests on your rights, privacy limits and terms, and shorter terms with clear permitted-use limits address legal uncertainty directly.

Does a license protect us if a buyer misuses the data?

A license gives contractual remedies such as termination, deletion obligations, audit rights and indemnities, depending on what was negotiated. It cannot prevent every misuse, which is why preparation removes personal and confidential details before delivery and why permitted-use and distribution terms should be drafted narrowly.

Do these rulings apply outside the United States?

No. They interpret US copyright law. Other jurisdictions use different doctrines, including text and data mining exceptions with opt-out mechanisms in some places, and developers training models for global markets may face several regimes. Ask counsel which laws may apply to a given buyer and use.

Sources

  • On February 11, 2025, Judge Stephanos Bibas, a Third Circuit judge sitting by designation in the U.S. District Court for the District of Delaware, granted partial summary judgment to Thomson Reuters and rejected Ross Intelligence's fair-use defense for copying Westlaw headnotes to build a competing AI legal-research tool. Source
  • In the February 2025 Ross decision, the court held that 2,243 Westlaw headnotes were original enough for copyright protection, and that fair-use factor one (purpose and character) and factor four (market effect) favored Thomson Reuters because Ross meant to compete with Westlaw by developing a market substitute. Source
  • The Ross case involved a non-generative AI tool, and commentators noted that the ruling left the questions of transformativeness and market effect in generative-AI training to other courts. Source
  • In late September 2026 (reported as September 29, 2026), a Third Circuit panel, in an opinion by Judge Tamika Montgomery-Reeves, affirmed that the 2,243 Westlaw headnotes are copyrightable and that ROSS's copying of them to train a legal-research AI was not fair use. LawNext reported it as the first federal appellate decision on fair use in AI training. Source
  • On June 23, 2025, Judge William Alsup of the U.S. District Court for the Northern District of California held in Bartz v. Anthropic PBC (No. 3:24-cv-05417) that using books to train large language models was 'exceedingly transformative' and therefore fair use. Source
  • Judge Alsup ruled that Anthropic's downloading of pirated books from sources including Books3, Library Genesis and Pirate Library Mirror to build a permanent central library was not fair use, and he left that infringement claim and the resulting damages for trial. Source
  • On July 20, 2026, Judge Araceli Martínez-Olguín of the Northern District of California granted final approval of the $1.5 billion Bartz v. Anthropic class settlement and entered final judgment, finding the deal 'fair, reasonable, and adequate' and overruling all 53 objections; JURIST called it a record settlement. Two later appeals concern only the attorneys' fee award, and in a September 2, 2026 status report the parties agreed that the settlement's Effective Date had passed. Source
  • The Bartz settlement class covers copyright owners of the 482,460 works on a Works List drawn from the LibGen and PiLiMi versions Anthropic downloaded, each with an ISBN or ASIN and a US Copyright Office registration made within five years of first publication and either before Anthropic downloaded it or within three months of publication. Anthropic must destroy the original torrented files and any copies derived from them. Source
  • The release in the final Bartz v. Anthropic settlement covers only past conduct through August 25, 2025, and expressly does not release claims about AI outputs or future conduct. Source
  • On June 25, 2025, Judge Vince Chhabria of the Northern District of California granted Meta partial summary judgment in Kadrey v. Meta Platforms (No. 3:23-cv-03417). He held that Meta's use of the plaintiff authors' books to train its Llama models was fair use on the record the plaintiffs presented. Source
  • Judge Chhabria's Kadrey ruling turned on market harm. The authors had not produced enough evidence that Meta's copying harmed the market for their books, and the court rejected lost licensing revenue for AI training as a cognizable harm, calling that reasoning circular. Source
  • In Kadrey v. Meta, Judge Chhabria wrote that 'it's hard to imagine that it can be fair use to use copyrighted books to develop a tool to make billions or trillions of dollars while enabling the creation of a potentially endless stream of competing works that could significantly harm the market for those books.' Source
  • In Kadrey v. Meta, Meta acknowledged downloading the plaintiffs' books from shadow libraries via BitTorrent. The plaintiffs, who included Sarah Silverman and Junot Díaz, also lost their DMCA claim for removal of copyright management information. Source
  • On May 9, 2025, the US Copyright Office released a pre-publication version of 'Copyright and Artificial Intelligence, Part 3: Generative AI Training'. It covers training activities that implicate copyright, when training may or may not be fair use, and the licensing of copyrighted works for AI training. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify