Skip to content

Rights and contracts

The US Copyright Office report on AI training: what it says about licensing

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

The US Copyright Office report on AI training, Part 3 of its Copyright and Artificial Intelligence series released in pre-publication form on May 9, 2025, treats fair use as case by case and recommends letting voluntary licensing markets keep developing. It is persuasive analysis, not binding law, and its licensing themes matter more to private-record holders than its fair-use analysis.

Key takeaways

  • Part 3 of the Copyright Office's AI series, released May 9, 2025 in pre-publication form, is analysis and recommendation, not binding law.
  • It treats some AI training as likely fair use and some as beyond fair use, depending on purpose, source and market effect.
  • It treats lost licensing opportunities as a relevant harm, a view the Kadrey v. Meta court rejected weeks later.
  • It recommends letting voluntary licensing markets develop without government intervention, which leaves deal terms to the parties.
  • For private operational records, the report's licensing themes matter more than its fair-use analysis.

The Copyright Office report on AI training is Copyright and Artificial Intelligence, Part 3: Generative AI Training, which the agency released in pre-publication form on May 9, 2025. It covers which training activities implicate copyright, when training may or may not be fair use, and the licensing of copyrighted works for AI training. Part 1, on digital replicas, appeared on July 31, 2024, and Part 2, on copyrightability, on January 29, 2025.

The report is not a law or a regulation, and courts are not bound by it. Judges may consider the Office's views and policymakers often cite them, but the outcome of any dispute still depends on the court and the facts.

Its status carries one footnote. The report came out less than 24 hours before Register of Copyrights Shira Perlmutter was dismissed, and the Office said a final version would follow without substantive changes expected. Check the Office's AI page for whether a final version has been issued before quoting specific passages.

Most of the report concerns published creative works such as books, news, music and images. Its reasoning still reaches operating companies, because it frames how buyers and courts think about licensing as the alternative to unlicensed copying.

The report's licensing-related findings can be summarized in five points. These summaries are simplified, and the full text contains qualifications that counsel should read before anyone relies on them.

  • Fair use is decided case by case. The report does not treat all training as fair use or all training as infringement; the purpose of the use, the source of the works and the market effect drive the answer.
  • Some uses go too far. The report concluded that commercial use of vast troves of copyrighted works to produce expressive content that competes with them in existing markets, especially through illegal access, goes beyond established fair use boundaries.
  • Licensing markets count in the market-harm analysis. Where licenses for training use exist or are reasonably likely to develop, the report treats lost licensing opportunities as a relevant harm.
  • Competing outputs can be harm. The report discusses market dilution, where large volumes of AI output compete with the kinds of works a model was trained on.
  • Voluntary licensing should keep developing. The report calls government intervention premature and suggests options such as extended collective licensing only if gaps persist.

Where the report and the 2025 court rulings diverge#

The report and the 2025 court rulings diverge most on licensing markets, the part of the debate that bears directly on data suppliers.

Weeks after the report, Judge Chhabria in Kadrey v. Meta rejected lost fees for licensing books as AI training data as a cognizable market harm, calling that reasoning circular, while signaling that stronger evidence of market dilution could change the result. In Thomson Reuters v. Ross, the court found market harm where an AI tool was built as a substitute for the original, and the Third Circuit affirmed in September 2026. A supplier should read the report's licensing view as persuasive and contested, not as settled law.

Where the report and the 2025 court rulings diverge
QuestionCopyright Office reportCourt rulings so farPractical reading for a supplier
Do lost licensing fees count as harm?Yes, where a training-license market exists or is likelyKadrey rejected the argument as circularProtect value through access control and terms, not doctrine
Do competing outputs count?Market dilution is a relevant harmKadrey signaled it could, with better evidenceLimit outputs that would compete with your services where relevant
Does illegal access matter?It makes commercial use harder to defendBartz: a pirated library was not fair use; Kadrey: shadow-library downloads did not change that record's resultKeep a record of how each dataset was obtained and who approved it

What the findings mean for a company with private records#

For a company holding private operational records, the findings shift attention from fair use to the quality of the license. The table connects each finding to a practical point for a supplier.

Access is the point that matters most. The report's fair-use discussion concerns works developers can obtain. Support tickets, CRM histories, dispatch records and private code are not obtainable without the company's permission, so for these records licensing is the orderly route however fair use develops.

What the findings mean for a company with private records
FindingPractical point for a data supplier
Fair use is case by caseSet boundaries in the license instead of relying on doctrine
Source of data mattersKeep records of where each dataset came from and who approved its release
Licensing markets countDocumented licenses show a market exists for a record type, though courts differ on how much weight that carries
Market dilutionConsider whether outputs trained on your records could compete with your own services, and limit use if so
Voluntary licensing preferredScope, term, exclusivity and price are negotiated, so drafting quality matters

What the report does not settle#

The report does not settle who owns or controls business records, how privacy law applies to training data, or what terms a data license should contain. Those questions sit in contract law, privacy law and trade secret law, and they are assessed deal by deal.

It also does not set prices, recommend rates or tell any particular company whether a particular use is permitted. For a general counsel, its value is as background: a clear statement of how one federal agency reads the law, useful in board discussions and in negotiations where a buyer raises fair use as leverage.

Using the report in a licensing negotiation#

The report is most useful in a negotiation as context, not as a bargaining chip. A buyer arguing that it could train on material without a license is usually talking about published content, not about records it cannot reach, and the report's own analysis concerns works that developers can obtain.

Where the report helps is in explaining why a supplier asks for documentation and conditions. A federal agency has described licensing as a developing market and singled out illegal access, so requests for provenance records, a defined permitted use and a deletion date read as ordinary market practice rather than special demands. Because courts have not uniformly accepted the licensing-market view, use it to set the tone of a negotiation, not as a legal threat.

Exclusivity and term length remain commercial questions. A non-exclusive license keeps the supplier free to license the same prepared package elsewhere, which fits the report's picture of a developing market with many participants.

Questions to ask internally after reading the report#

The questions to ask internally after reading the report are about your records, not about copyright doctrine. The report explains how one agency sees training on published works; your decision depends on what you hold, what you promised and what you would allow a buyer to do.

Answering these questions in writing, per record family, gives the board and the authorized signer a basis for approval that does not shift with each new court decision or policy statement.

  • Which of our records are written expression, which are mostly facts, and which belong to customers or clients?
  • Have any of these records ever been published, or are they kept confidential under access controls?
  • What did our privacy policies, customer contracts and employee notices say when the records were created?
  • Could a model trained on these records produce outputs that compete with our own services?
  • Would we license the same package to more than one developer, or only on exclusive terms?

Illustrative: a consulting firm's board asks about the report#

Illustrative: a fictional operations consulting firm is considering licensing its internal project review notes and anonymized playbooks. A board member asks whether the Copyright Office report means AI developers could use the material without paying for it.

The managing partner explains that the playbooks have never been published and sit in the firm's Notion workspace and SharePoint, so no developer can reach them without the firm's permission. Counsel adds that the report's emphasis on licensing markets and data sources supports insisting on clear provenance and permitted-use terms. The board approves a narrow, non-exclusive license that excludes client deliverables and anything that identifies a client.

How SourceX documents licensing terms#

SourceX structures each package as a license, with the supplier keeping ownership, and runs it through the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. The SourceX Enterprise Data Value Framework is used to discuss what makes a record family useful to buyers before terms are set.

Each package that proceeds is documented in a SourceX Evidence Packet covering provenance, licensing rights, permitted use, the privacy record and release authorization. That record speaks directly to the questions about data sources and permitted use that the report highlights.

Frequently asked questions

Is the Copyright Office report binding on courts?

No. The Copyright Office administers registration and advises Congress, and its reports are analysis rather than law. Courts may find its reasoning persuasive, but they decide fair-use cases on their own reading of the statute and the facts before them. Treat the report as informed guidance, not as a rule.

Does the report say AI developers must license training data?

No. It says the answer depends on the facts, that some uses are likely fair and others go beyond fair use, and that licensing markets are relevant to the analysis. It encourages voluntary licensing to develop but creates no requirement to license, and at least one court has since declined to treat lost licensing fees as harm.

Does the report cover private business data?

Not directly. It focuses on copyrighted works such as books, news and music. Many business records contain copyrightable writing, but their protection often depends more on confidentiality, contracts and trade secret law. The report's licensing themes still shape how buyers approach private data.

Where can I read the report?

The report is published on the US Copyright Office website as part of its series on copyright and artificial intelligence. The version released on May 9, 2025 was labeled pre-publication, so check whether a final version has since been posted, and ask counsel to point out the sections most relevant to your record types.

Should our license agreement cite the report?

Usually not. A license should state its own terms clearly: permitted use, restrictions, term, deletion and remedies. Citing agency reasoning in the contract adds little and could create interpretation questions if the law moves. Use the report in internal discussion and negotiation instead.

Sources

  • On May 9, 2025, the US Copyright Office released a pre-publication version of 'Copyright and Artificial Intelligence, Part 3: Generative AI Training'. It covers training activities that implicate copyright, when training may or may not be fair use, and the licensing of copyrighted works for AI training. Source
  • The Part 3 pre-publication report came out less than 24 hours before Register of Copyrights Shira Perlmutter was dismissed, and the Office said a final version would follow 'without any substantive changes expected.' Source
  • Part 1 of the US Copyright Office's Copyright and Artificial Intelligence report (Digital Replicas) was published on July 31, 2024, and Part 2 (Copyrightability) on January 29, 2025. Source
  • The Copyright Office's Part 3 pre-publication report states that 'making commercial use of vast troves of copyrighted works to produce expressive content that competes with them in existing markets, especially where this is accomplished through illegal access, goes beyond established fair use boundaries,' and recommends letting voluntary licensing markets for AI training keep developing without government intervention. Source
  • Judge Chhabria's Kadrey ruling turned on market harm. The authors had not produced enough evidence that Meta's copying harmed the market for their books, and the court rejected lost licensing revenue for AI training as a cognizable harm, calling that reasoning circular. Source
  • In Kadrey v. Meta, Judge Chhabria wrote that 'it's hard to imagine that it can be fair use to use copyrighted books to develop a tool to make billions or trillions of dollars while enabling the creation of a potentially endless stream of competing works that could significantly harm the market for those books.' Source
  • In the February 2025 Ross decision, the court held that 2,243 Westlaw headnotes were original enough for copyright protection, and that fair-use factor one (purpose and character) and factor four (market effect) favored Thomson Reuters because Ross meant to compete with Westlaw by developing a market substitute. Source
  • In late September 2026 (reported as September 29, 2026), a Third Circuit panel, in an opinion by Judge Tamika Montgomery-Reeves, affirmed that the 2,243 Westlaw headnotes are copyrightable and that ROSS's copying of them to train a legal-research AI was not fair use. LawNext reported it as the first federal appellate decision on fair use in AI training. Source
  • Judge Alsup ruled that Anthropic's downloading of pirated books from sources including Books3, Library Genesis and Pirate Library Mirror to build a permanent central library was not fair use, and he left that infringement claim and the resulting damages for trial. Source
  • In Kadrey v. Meta, Meta acknowledged downloading the plaintiffs' books from shadow libraries via BitTorrent. The plaintiffs, who included Sarah Silverman and Junot Díaz, also lost their DMCA claim for removal of copyright management information. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify