Leadership and readiness
How 2025 AI copyright rulings strengthened the case for licensed business data
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
The 2025 AI copyright rulings strengthened the case for licensed business data by making how training material was obtained a central legal question. Courts found training on lawfully bought books could be fair use, but pirated copies led to a $1.5 billion settlement and a competing legal-research tool lost on fair use. Licensed records with documented provenance avoid both problems.
Key takeaways
- In 2025, US trial courts reached different fair use results on different facts: Bartz and Kadrey favored developers on training, while Thomson Reuters v. Ross did not.
- How material was obtained became its own question: scanning purchased print books was fair use in Bartz, while books downloaded from pirate libraries led to a $1.5 billion settlement.
- Kadrey rejected lost AI-training licensing revenue as market harm, so the case for licensing rests on legal certainty and provenance, not on courts protecting a licensing market.
- Operational business records sit behind logins, so a written license is the practical route for a developer to use them.
- For a company licensing records, provenance and rights documents are now part of what a buyer evaluates.
What did the 2025 rulings actually decide?#
The 2025 rulings decided narrow questions in specific lawsuits brought over books and published content, not business records, and they did not reach a single answer on whether AI training is fair use. The timeline below lists the main public decisions and what followed.
Two caveats matter. The 2025 US decisions were trial-court rulings tied to their records; Judge Chhabria stressed in Kadrey that his ruling did not hold AI training lawful generally. And commentators noted that Ross involved a non-generative tool, leaving generative-AI training questions to other courts. Treat this summary as a starting point for a conversation with counsel, not a statement of current law.
| Date | Decision or event | What it held or required |
|---|---|---|
| February 11, 2025 | Thomson Reuters v. Ross Intelligence (D. Del.) | Judge Bibas rejected Ross's fair use defense for copying Westlaw headnotes to build a competing AI legal-research tool |
| June 23, 2025 | Bartz v. Anthropic (N.D. Cal.) | Judge Alsup held training on books was exceedingly transformative and fair use, and that scanning purchased print books was fair use |
| June 25, 2025 | Kadrey v. Meta Platforms (N.D. Cal.) | Judge Chhabria found fair use on the record presented and rejected lost AI-training licensing revenue as a cognizable market harm |
| September 25, 2025 | Bartz settlement, preliminary approval | A $1.5 billion class settlement covering about 482,460 works downloaded from pirate libraries, with the downloaded files to be destroyed |
| November 4, 2025 | Getty Images v Stability AI (High Court, England and Wales) | Secondary infringement claim dismissed because the model weights do not store Getty's works; permission to appeal granted in December 2025 |
| July 20, 2026 | Bartz settlement, final approval | Final approval and final judgment entered |
| September 2026 | Thomson Reuters v. Ross (Third Circuit) | Affirmed that copying the headnotes to train the tool was not fair use, reported as the first federal appellate decision on fair use in AI training |
The questions courts kept returning to#
Courts kept returning to four questions, and each one maps to something a business licensing its records can document. The table pairs the legal question with its practical meaning for a supplier.
The common thread is uncertainty. Where the law is unsettled, a developer that can point to a written license and a documented chain of custody does not need to win a fair use argument for that material.
| Question courts examined | Why it matters to AI developers | What a licensing company can document |
|---|---|---|
| How was the material obtained? | Copies from unauthorized sources were treated as a separate problem from training | A signed license and a record of how the copy was produced and delivered |
| Does the AI system substitute for the original? | Uses that compete with the source's own product weighed against the developer | Permitted use terms that define training, evaluation and any limits |
| Does lost licensing revenue count as market harm? | Kadrey rejected it as circular reasoning, so developers cannot count on that argument either way | A licensed use does not depend on winning that argument |
| Who can grant permission? | A license is only as good as the licensor's rights | Rights review, signer authority and a release authorization |
Why business records sit outside the fair use debate#
Business records sit largely outside the fair use debate because most of them were never public in the first place. Support tickets, dispatch logs, nonconformance reports and engineering reviews live behind logins, so a developer cannot lawfully collect them by crawling the web.
The legal protection around operational records also rests on more than copyright. Confidentiality obligations, trade secret law, customer contracts, vendor terms and privacy law all apply, and none of them is resolved by a fair use ruling. For these records, a license from the company that holds them is the clean route.
A dispatch log from a regional HVAC contractor, for example, is protected less by copyright than by the contractor's customer agreements, its confidentiality practices and the privacy rights of the homeowners named in it. Each of those layers needs its own check before any license.
That is why the 2025 decisions matter to business owners indirectly. They raised the cost of uncertainty for developers, which makes well-documented licensed sources more attractive, and operational records rarely reach developers any other way.
Provenance records are now part of the product#
Provenance records are now part of the product because a developer needs to show where each dataset came from and what it may be used for. A strong dataset with weak paperwork is harder to use than a smaller one with clear documentation.
Public standards and research reflect this. The Data & Trust Alliance's Data Provenance Standards include Use metadata elements for license to use, intended data use, consent documentation location, confidentiality classification and copyright, patent and trademark status. The Data Provenance Initiative audited 44 data collections spanning more than 1,800 fine-tuning text datasets, documenting their sources, licenses and creators. In California, AB 2013 required developers of generative AI systems to post training-data documentation by January 1, 2026, including whether datasets were purchased or licensed.
For a company preparing records, the paperwork usually includes the items below.
- A description of the source systems and the date range of the records.
- The license grant: permitted uses, term, territory and exclusivity.
- The rights review: customer contracts, vendor terms and employee notices checked.
- The privacy record: what personal and confidential details were removed, and how.
- A release authorization signed by someone with authority to bind the company.
What the rulings do not change for business owners#
The rulings do not change the basic work a company must do before licensing records. They do not make every archive valuable, and they do not override the contracts and laws that already govern the records.
Keep these limits in view when someone cites the 2025 decisions as a reason to move quickly.
- Customer contracts that limit reuse still apply, whatever courts say about fair use.
- Privacy laws still require care with personal information inside the records.
- The decisions do not settle whether a given set of business records is protected by copyright.
- No AI developer is obliged to license any particular dataset, and value depends on buyer demand.
- Appeals and new cases can shift the legal picture, so license terms should not rely on one reading of the law.
Illustrative: a software company answers an inbound request#
Illustrative: a fictional B2B software company receives an email from a model developer asking to use its support tickets under a short research agreement. The draft has no permitted-use definition, no deletion terms and nothing about how the developer will document the source.
After reading commentary on the 2025 decisions, the CEO decides the company will share nothing without a written license that defines training and evaluation use, requires deletion at the end of the term and records provenance on both sides. Counsel reviews customer agreements first and carves out tickets from customers whose contracts restrict reuse.
The process is slower than the developer hoped, but the company ends with a clear decision record. Whether or not a license is signed, it now knows which records it could license, who must approve and what it will ask any buyer to confirm.
How SourceX approaches provenance and rights#
SourceX approaches provenance and rights as the core of the transaction rather than an afterthought. Every package moves through the SourceX five-step transaction, Supply, Rights, Preparation, Approval and Delivery, and the supplier approves each step.
The SourceX Evidence Packet records provenance, licensing rights, permitted use, the privacy record and release authorization for each package. Records are licensed, not sold outright, so the company keeps ownership, and SourceX's own rights in a deidentified dataset are set out in the signed supplier agreement.
For a business owner, the practical lesson of 2025 is that paperwork is part of the deal. Preparing the rights review and provenance record before talking to any buyer puts the company in a stronger position, whatever happens next in the courts.
Frequently asked questions
Do the 2025 rulings mean AI developers can use any data they find?
No. The decisions addressed specific facts, mostly involving books and published works, and they treated unlawfully obtained copies as a separate problem. Business records behind logins also carry contract, confidentiality and privacy protections that a fair use ruling does not touch. Ask counsel how current case law applies to your records.
Are our support tickets protected by copyright?
Possibly in part. Written replies, documentation and code can have copyright protection, while bare facts and data fields generally do not. In practice, protection for operational records rests on several layers at once, including confidentiality, contracts and trade secret law, which is why a rights review looks at all of them.
Should we wait for appeals before considering a license?
Waiting for the law to settle is not required, because a license rests on contract rather than on fair use. What matters is that your terms do not assume one reading of the law, and that rights, privacy and approval work is finished before anything is delivered.
Why would an AI developer want business records rather than published content?
Published content shows finished writing, while operational records show how work actually gets done: requests, decisions, tool use, exceptions and outcomes. That makes them useful for training and evaluating systems meant to perform business tasks, a gap that public web text fills poorly. Value still depends on the specific records and on buyer demand.
Does licensing our records protect us if a buyer misuses them?
A license gives you contractual remedies, not a guarantee. Permitted use, audit rights, deletion obligations and indemnities define what the buyer may do and what happens if it breaches. Clear terms and a documented delivery record make enforcement more practical.
Sources
- On February 11, 2025, Judge Stephanos Bibas granted partial summary judgment to Thomson Reuters and rejected Ross Intelligence's fair-use defense for copying Westlaw headnotes to build a competing AI legal-research tool. Source
- The Ross case involved a non-generative AI tool, and commentators noted the ruling left generative-AI training questions to other courts. Source
- On June 23, 2025, Judge William Alsup held in Bartz v. Anthropic PBC that using books to train large language models was exceedingly transformative and therefore fair use. Source
- In the June 23, 2025 Bartz order, Judge Alsup held that the purchase and scanning of print books was fair use. Source
- On June 25, 2025, Judge Vince Chhabria granted Meta partial summary judgment in Kadrey v. Meta Platforms, finding fair use on the record presented, and stressed the ruling does not hold AI training lawful generally. Source
- The Kadrey court rejected lost licensing revenue for AI training as a cognizable market harm, calling that reasoning circular. Source
- Judge Alsup granted preliminary approval of the $1.5 billion Bartz v. Anthropic settlement on September 25, 2025, requiring destruction of the pirated book files; final approval and final judgment were entered on July 20, 2026. Source
- The Bartz settlement class covers the 482,460 works on a Works List drawn from the LibGen and PiLiMi versions Anthropic downloaded, and Anthropic must destroy the original torrented files and derived copies. Source
- On November 4, 2025, the High Court of England and Wales dismissed Getty Images' secondary copyright infringement claim against Stability AI, and in December 2025 granted Getty permission to appeal. Source
- In September 2026 a Third Circuit panel affirmed that ROSS's copying of Westlaw headnotes to train a legal-research AI was not fair use, reported as the first federal appellate decision on fair use in AI training. Source
- California AB 2013 requires developers of generative AI systems to post training-data documentation on or before January 1, 2026, stating whether the datasets were purchased or licensed. Source
- The Use group of the Data & Trust Alliance Data Provenance Standards includes elements for confidentiality classification, consent documentation location, license to use, intended data use, and copyright, patent and trademark status. Source
- The Data Provenance Initiative released an audit covering 44 data collections spanning more than 1,800 fine-tuning text-to-text datasets, documenting their sources, licenses, creators and other metadata. Source
Related resources
- QuestionDo AI labs buy legal documents?
- QuestionWhat is data provenance and why do buyers care?
- InsightHow do I de-identify contracts and legal documents for AI training?
- InsightIndemnification in data licenses: who covers which claims
- InsightOpt-in vs opt-out for AI training in B2B SaaS contracts
- IndustryConstruction data
See if your company qualifies
A short company assessment. No data uploads are needed.