Skip to content

Regulation and governance for data buyers

What the US Copyright Office's AI Training Report Means for Licensing Training Data

Quick answer

The US Copyright Office's Part 3 report on generative AI training, still a pre-publication version as of October 2026, concludes that collecting, curating and training on copyrighted works implicates the reproduction right, that fair use is a case-by-case question weighted by source, purpose and output controls, and that voluntary licensing markets should be left to develop [1][2]. For buyers, the practical effect is that provenance and a written license are now the cleanest way to take fair use off the critical path.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Status: guidance from a pre-publication report, not a rule of law

The report is persuasive agency guidance, not a statute or regulation, and courts decide fair use case by case. The Office released Part 3 on May 9, 2025 as a "pre-publication version" in response to congressional inquiries, saying a final version would follow without substantive changes expected to its analysis or conclusions [1]. As of October 2026, the Office's AI initiative page still lists Part 3 as pre-publication, alongside Part 1 (digital replicas) and Part 2 (copyrightability of outputs) [2]. Counsel should re-check that page before citing the report in a memo, a board paper or a disclosure.

Treat the report as a map of how a sophisticated federal reviewer reasons about training, not as a safe harbor. Judges are not bound by it. In Kadrey v. Meta, for example, the court granted Meta partial summary judgment on fair use on the specific record presented, and the case is not final [6]. In Bartz v. Anthropic, the class settlement received final approval in July 2026; a settlement is not a merits ruling [7].

Where the report finds reproduction in the training pipeline

The Office's core finding is that several steps of building a generative system make copies, so each step needs either a license or a defense. Commentary on the report highlights that compiling a training dataset implicates the reproduction right [3]. In engineering terms, that covers the crawl or bulk download, the raw object store, deduplicated and filtered shards, tokenized training files (for example Parquet, JSONL or WebDataset tar shards), evaluation holdouts and any backup copies.

Three further points matter for scoping a license:

  • Training itself. Copies made during training runs fall within the analysis, so a license that covers only "storage" or "internal analytics" can leave the training step unaddressed [1].
  • Model weights. The Office considered whether weights can themselves be infringing copies where the model can output material substantially similar to training inputs, which turns on memorization [1][3]. Quote the report's own language on this point rather than secondary summaries.
  • Retrieval-augmented generation. The report treats RAG as involving reproduction, because documents are copied into an index and passages are reproduced at inference time [3]. A license built for pre-training does not automatically cover a vector index or verbatim retrieval into outputs.

How the Office weighs fair use for training

The Office describes fair use for training as fact-specific, turning on which works were used, where they came from, what the model is for and what controls sit on outputs [1][4]. Analysis of the report notes that knowingly using pirated or illegally accessed copies weighs against fair use [4]. That makes acquisition method a first-order legal fact, not just an ethics question, which is the subject of our page on lawful access and pirated sources.

On purpose, the report distinguishes uses that are more transformative, such as training a model for analysis or a narrowly deployed tool, from uses where a model generates expressive content that competes with the works it was trained on [1]. Output-side controls, such as filters that block regurgitation of protected text and refusal of prompts requesting verbatim works, are treated as relevant facts [4]. On market harm, the Office looks at lost licensing revenue where a licensing market exists or is developing, which is why the state of licensing markets feeds directly back into the fair use analysis [1][5].

Failure modes counsel see in diligence follow from this framework:

  • Shadow-library corpora (books3-style dumps, torrent mirrors) mixed into "general web" shards with no source field.
  • Paywalled news or journal content collected by circumventing access controls.
  • Deduplication that runs after tokenization, so the provenance of each shard is lost.
  • RAG indexes that ingest whole licensed PDFs under a license that only covered training.

What the report says about licensing markets and collective licensing

The Office recommended no new legislation for now, preferring to let voluntary licensing markets develop, and it discussed collective licensing as a way to aggregate rights [5]. Commentary notes the Office viewed the AI licensing market as still forming, with direct deals for some content types and real friction for others, such as diffuse, low-value-per-work text [5]. It also considered government-led options and did not recommend them at this stage [1][5].

For buyers, this has three consequences:

  1. In the Office's view, an existing market cuts against fair use. If licenses are reasonably available for a category of content, an unlicensed user has a harder fourth-factor argument [1].
  2. Collective licenses fill gaps, not everything. Blanket licenses from collecting societies can cover large catalogs of published works, but scope, opt-outs and AI-specific terms vary; see collective licenses for AI training.
  3. Non-public data has no fair use shortcut in practice. Operational records such as support tickets, engineering logs or contract workflows have no publicly accessible copy to rely on, so a license from the holder is usually the only route to access at all. Whether such records are protected in the first place is covered in copyright in operational business records.

Build versus license: a decision table for counsel

The report pushes high-exposure categories toward licensing and leaves fair use as a defense for narrower, well-controlled uses. The table below translates the factors into a first-pass sorting tool for a data acquisition review. It is a triage aid, not a legal conclusion.

Illustrative example: invented to show structure; it does not describe an available dataset.

Data sourceAccess routeIntended useOutput riskReport-informed posture
Public web text, robots-respecting crawlLawful, no loginPre-training, general modelMedium (generation competes with some sources)Fair use arguable; document crawl policy and filters; license high-value publishers
Books or articles from shadow librariesUnlawfulAnyHighExclude; acquisition weighs against fair use [4]
Licensed news or journal archiveContractPre-training plus RAGHigh for RAG verbatimLicense must name training and retrieval separately [3]
Non-public company records (tickets, CRM, contracts)Only via holderFine-tuning, evals, RAGLow to medium after redactionLicense required for access; add privacy controls
Open-licensed corpora (CC BY, CC BY-SA)LawfulPre-trainingMedium (attribution, share-alike)Check license elements; see Creative Commons and AI training

License and diligence checklist derived from the report

Each reproduction point the Office identified should map to a granted right, a record, or a deliberate exclusion. Use this checklist when reviewing a supplier license or a data acquisition memo.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Grant covers each copy: ingestion, storage, preprocessing, training, evaluation and backups are named uses, not implied ones.
  • RAG and indexing are explicit: embedding, vector storage and display of retrieved passages are either granted or excluded.
  • Weights and derivatives: the license states whether trained weights, fine-tuned checkpoints and distilled models may be retained and distributed after term.
  • Provenance field per record: each file or shard carries a source ID, acquisition date and license reference, preserved through deduplication (for example a source_id and license_id column carried into Parquet shards).
  • Lawful access evidence: the supplier represents how it obtained the data and that no access controls were circumvented.
  • Output controls documented: regurgitation filters, memorization tests on held-out licensed text and prompt-level refusals are recorded in the model card.
  • Personal data handled separately: copyright clearance does not resolve privacy; see PII redaction for LLM training data.
  • Jurisdictional overlay: if the model is placed on the EU market, the copyright policy required by AI Act Article 53 and described in the GPAI Code of Practice copyright chapter must cover the same data [8]; see the GPAI Code of Practice copyright chapter.

How the US report fits with EU and UK rules

The US approach rests on fair use as a flexible defense, while most other major regimes rely on statutory exceptions plus licensing, so one dataset can sit under different rules depending on where the model is trained and placed on the market. The EU layers text and data mining opt-outs onto Article 53 duties for general-purpose model providers [8]. The UK's computational analysis exception remains limited to non-commercial research, which we cover in UK commercial AI training and copyright. For a country-by-country view, see text and data mining exceptions by country, and for the wider regulatory picture, start at the AI training data compliance hub.

The practical answer for a multi-jurisdiction program is to license to the strictest regime you will ship into and keep one provenance record per dataset that serves US fair use analysis, EU summaries and customer audits. Our enterprise data licensing explainer and the comparison of licensed, synthetic and scraped data cover how buyers structure those terms.

Where SourceX fits for licensed operational data

SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents and finance or legal workflows, and manages licensing and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, with diligence materials on source, rights, preparation and allowed use prepared per dataset. SourceX does not source scraped web content. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. If you need non-public records for fine-tuning, evaluation or RAG, you can describe the data you need.

Request licensed training data with a documented rights trail

If the Copyright Office's analysis has moved a dataset from "arguable fair use" to "license it," describe the records, intended uses and delivery needs rather than the businesses that might hold them. SourceX looks for US companies holding that data, every release is approved by the supplying company, and personal details are removed or replaced before delivery under a recorded method. Start a request at sourcex.si/buyers.

Sources

  1. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  2. U.S. Copyright Office, "Copyright and Artificial Intelligence". https://www.copyright.gov/ai/
  3. Jenner & Block, "US Copyright Office Releases \"Pre-Publication Version\" of Report on Copyright Issues in Generative AI Training" (2025). https://www.jenner.com/en/news-insights/client-alerts/us-copyright-office-releases-pre-publication-version-of-report-on-copyright-issues-in-generative-ai-training
  4. Manatt, Phelps & Phillips, "Copyright Office Releases Pre-Publication Report on Copyrighted Works in Generative AI Training" (2025). https://www.manatt.com/insights/insight/copyright-office-releases-pre-publication-report-on-copyrighted-works-in-generative-ai-training
  5. Mishcon de Reya, "US Copyright Office report Part 3: generative AI training" (2025). https://www.mishcon.com/news/us-copyright-office-report-part-3-generative-ai-training
  6. Akin Gump Strauss Hauer & Feld, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
  7. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  8. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data