Skip to content

AI data market

What the Thomson Reuters v. Ross appeal means for licensing training data

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

The Thomson Reuters v. Ross appeal matters because the Third Circuit reportedly affirmed in September 2026 that Westlaw headnotes are copyrightable and that copying them to train a competing AI legal research tool was not fair use. For data owners, it strengthens licensing as the expected route to training material, though it leaves generative AI open.

Key takeaways

  • In February 2025 the trial court rejected Ross's fair use defense for copying 2,243 Westlaw headnotes to build a competing legal research tool.
  • LawNext reported the Third Circuit's September 2026 affirmance as the first federal appellate decision on fair use in AI training.
  • The tool was non-generative and built as a market substitute for its source, which limits how far the reasoning reaches.
  • In Kadrey v. Meta another court rejected lost AI licensing revenue as a cognizable harm, so the market question is not settled.
  • Precise permitted-use terms and documented licenses matter more as courts weigh competing products and licensing markets.

What was Thomson Reuters v. Ross about?#

Thomson Reuters v. Ross is a copyright dispute between the owner of the Westlaw legal research service and a startup that built an AI legal search tool. The publisher claimed the startup used material derived from its editorial headnotes, the short summaries of legal points that editors attach to court opinions, to train the tool.

The training material reached the startup through question-and-answer memos prepared by a contractor, so the case is as much about sourcing as about use. The tool returned existing judicial opinions in response to legal questions; it did not generate new text, a detail commentators flagged early because it limits how far the reasoning reaches.

On February 11, 2025, Judge Stephanos Bibas, a Third Circuit judge sitting by designation in the District of Delaware, granted partial summary judgment to Thomson Reuters and rejected the fair use defense, reversing his own 2023 denial of summary judgment. He held that 2,243 headnotes were original enough for copyright protection, found that fair use factors one and four favored Thomson Reuters because Ross meant to build a market substitute for Westlaw, and certified the questions for an interlocutory appeal.

What did the Third Circuit decide?#

The Third Circuit affirmed. A panel of Judges Restrepo, Montgomery-Reeves and Bove heard argument on June 11, 2026, and in an opinion by Judge Montgomery-Reeves, reported as issued on September 29, 2026, held that the headnotes are copyrightable and that Ross's copying of them to train its legal research tool was not fair use. LawNext reported it as the first federal appellate decision on fair use in AI training.

Reports differ on exactly when and how the opinion was first filed, so read the opinion itself, and current counsel commentary, before quoting any holding. The table separates what was decided at each stage from what it means for a company that owns records.

What did the Third Circuit decide?
StageWhat was decidedWhat it means for data owners
District court, February 20252,243 headnotes protected; fair use rejected; factors one and four favored Thomson ReutersCurated written summaries can carry protection that raw facts lack
Third Circuit argument, June 11, 2026Panel considered originality of the headnotes and Key Number System and whether the use was fairAppellate courts were willing to decide AI training questions early
Third Circuit opinion, September 2026Affirmed copyrightability and the rejection of fair useCopying to build a competing product is hard to defend; licensing is the expected route
Not decidedGenerative models, private records, courts in other circuitsOperational records still rest on contracts, confidentiality and access controls

What the decision leaves open#

The decision leaves open how its reasoning applies to generative models trained on large, mixed collections, because the tool in this case was a narrow search product that competed with its source. Courts hearing cases about books, news, music and code will decide how much of the reasoning carries over.

Several other questions remain for later courts and later stages of this case. General counsel should treat the list below as the watch list before relying on the decision in a negotiation or policy.

  • Generative models trained on many sources rather than on one competitor's content.
  • Private, unpublished records, which raise access and contract questions rather than fair use.
  • Courts in other circuits, which are not bound by the Third Circuit and may reason differently.
  • Remedies and damages, which depend on later proceedings.
  • How future courts measure whether a licensing market is established enough to count.

Why a recognized licensing market matters to data owners#

A recognized licensing market matters because the fourth fair use factor asks whether copying harms the market for the work, including licensing markets the owner could reasonably develop. When a court treats licensing for AI training as a real market, unlicensed copying looks more like taking a sale than creating a new use. Not every court agrees: in Kadrey v. Meta, Judge Chhabria rejected lost licensing revenue for AI training as a cognizable harm and called that reasoning circular, so the weight of a licensing market remains contested.

That reasoning has a practical side for owners of operational records. Each documented license, with defined permitted use and provenance, is evidence that the market exists and works. It also makes licensed data the cleaner route for developers that do not want to argue fair use for every source.

The effect is strongest for curated written material such as editorial summaries, playbooks and annotated records. Raw facts remain unprotected by copyright, so licenses for factual operational data rest mainly on access, confidentiality and contract.

What general counsel should take from the case#

General counsel can take practical steps from the case without waiting for its later history, because each one improves the company's position whatever happens next. Most of them fit into work legal teams already do on contracts and data inventories.

  • Separate curated content from raw facts in the data inventory; written summaries, notes and playbooks may carry protection that transaction data does not.
  • Keep permitted-use terms precise about training, evaluation and fine-tuning, and address use in competing products.
  • Control how licensees obtain data and through whom, since contractors and intermediaries can break the chain of permission.
  • Review website and API terms that govern access to any published content.
  • Track the decision's later history and rulings in other circuits before relying on it.

Illustrative: a distributor's application notes#

Illustrative: a fictional industrial distributor employs application engineers who have written years of short notes explaining which bearing, seal or lubricant suits a customer's operating conditions, each linked to the original inquiry in its CRM. A developer building a product-selection assistant asks to license the notes.

The distributor's general counsel reads commentary on the Ross decision and draws two points. The notes are curated expert summaries, the kind of written work that may carry more protection than the order data around them, so they should not circulate on loose terms. And a written license documents a market for the notes that informal sharing would not.

The distributor licenses the notes and linked inquiries after removing customer names and contact details, limits permitted use to training and evaluation, and bars use in a product that would compete with its own selection service.

How SourceX approaches the questions in Ross#

SourceX handles each package as a written license, so the questions in Ross about unlicensed copying do not arise for the records it handles. The SourceX Evidence Packet records provenance, licensing rights, permitted use, the privacy record and release authorization, the documentation a functioning licensing market relies on.

The supplier approves each step of the SourceX five-step transaction, Supply, Rights, Preparation, Approval and Delivery, and SourceX's own rights in a deidentified dataset are set out in the signed supplier agreement. Questions about copyright protection for specific records, and about how the decision applies to them, belong with the supplier's counsel.

Frequently asked questions

Does the Ross decision apply outside the Third Circuit?

An appellate decision binds federal courts only within its own circuit, which covers Delaware, New Jersey, Pennsylvania and the Virgin Islands. Courts elsewhere may find its reasoning persuasive or reach different results. Expect lawyers in other AI cases to cite it, and watch for later decisions that agree or disagree.

Is the case about generative AI?

Not directly. The tool at issue retrieved existing opinions in response to legal questions rather than generating new text. Commentators have said this narrows how far the reasoning applies to generative models trained on broad collections, although its market-harm analysis may still be cited in those cases.

Does the decision make our operational records more valuable?

Not by itself. Value depends on the records: uniqueness, domain expertise, human-generated signal, recency, cleanliness and rights. What the case may change is the backdrop. If licensing for AI training is treated as a real market, developers have more reason to license documented sources than to rely on contested defenses.

Should we rewrite our existing data licenses after the ruling?

Review them with counsel rather than rewriting them by reflex. Useful checks include whether permitted use clearly covers or excludes training, evaluation and fine-tuning, whether use in competing products is addressed, and whether the licensee may pass data to contractors. Precise terms matter more as courts pay attention to licensing markets.

Where can we read the decision?

The opinion is available from the Third Circuit and through standard legal research services, and many law firms have published summaries. Read the opinion itself for the holding and any later history, because alerts and press coverage often simplify what a court decided on each issue.

Sources

  • On February 11, 2025, Judge Stephanos Bibas, sitting by designation in the District of Delaware, granted partial summary judgment to Thomson Reuters and rejected Ross Intelligence's fair use defense, reversing his own September 2023 denial of summary judgment. Source
  • The court held that 2,243 Westlaw headnotes were original enough for copyright protection and that factors one and four favored Thomson Reuters because Ross meant to develop a market substitute; the tool was non-generative. Source
  • On June 11, 2026, the Third Circuit heard oral argument in the certified interlocutory appeal before Judges Restrepo, Montgomery-Reeves and Bove. Source
  • In late September 2026 a Third Circuit panel, in an opinion by Judge Montgomery-Reeves, affirmed that the headnotes are copyrightable and that Ross's copying to train its AI was not fair use; LawNext reported it as the first federal appellate decision on fair use in AI training. Source
  • In Kadrey v. Meta, the court rejected lost licensing revenue for AI training as a cognizable market harm, calling that reasoning circular. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify