Skip to content

Data licensing for AI training

License or rely on fair use? How AI teams decide whether to license training data

Quick answer

You do not always need a license to train on copyrighted works in the US, but fair use is a case-by-case defense, not a permission, and it never gives you access to data you cannot lawfully obtain. As of October 2026, courts have split: training was held fair use on some records and not on others, and pirated acquisition drew settlement-scale exposure. License when the data is non-public, when your model competes with the source's market, when acquisition is questionable, or when you ship outside the US.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What the 2025-2026 rulings actually decided

The rulings decided narrow, record-specific questions, and none of them establishes that AI training is categorically fair use. Read each for the factor that tipped it, because those factors are what your own facts will be tested against.

  • Thomson Reuters v. ROSS (3d Cir., 29 September 2026). The Third Circuit affirmed that Westlaw headnotes were original enough to be protected and that copying them to build a competing legal research tool was not fair use [2]. The tool was not generative, and the court's reasoning leaned on the substitute-product relationship between the two products. Lesson: when your output serves the same buyers as the source, transformative-purpose arguments weaken sharply.
  • Kadrey v. Meta (N.D. Cal., June 2025). The court granted Meta partial summary judgment on fair use for training Llama on books, but expressly because the plaintiffs had not built a record of market harm, including market dilution [3]. The case is ongoing. Lesson: the win was evidentiary, and a better-prepared plaintiff could change the outcome.
  • Bartz v. Anthropic. The class settlement of $1.5 billion received final approval [4].

The U.S. Copyright Office's Part 3 report, still a pre-publication version as of October 2026, reaches a similar shape: some training uses will be fair and others will not, and the existence of a functioning licensing market for the works is relevant to the fourth factor [1]. Do not predict outcomes from any single ruling; build the decision on the factors.

Where fair use cannot help: access, contracts and non-public data

Fair use is a defense to infringement, so it does nothing for data you have no lawful way to obtain. Much of the data that matters for enterprise and agentic use cases sits in this category.

Support ticket histories in Zendesk or Salesforce Service Cloud, engineering records in Jira and GitHub Enterprise, contract repositories, finance workflows in NetSuite: none of it is on the open web. Getting it means a company decides to release it, which means a contract. Even where content is technically reachable, three non-copyright regimes can block reliance on fair use:

  • Contract and terms of service. Click-through or API terms that prohibit model training create breach claims that fair use does not answer.
  • Trade secret. Proprietary code, playbooks and internal documents may be protected regardless of copyrightability; see source code license terms for AI training.
  • Privacy and consent. Personal data in operational records triggers obligations under privacy law that a copyright defense never touches.

The practical rule for counsel: fair use is a question only for public, lawfully accessed content. For proprietary operational data, the question is which license terms, not whether to license. The comparison in licensed vs synthetic vs scraped data covers the alternatives.

Jurisdiction changes the answer outside the US

Outside the US, there is usually no open-ended fair use defense, so the analysis turns on specific statutory exceptions with conditions. A model trained in reliance on US fair use can still face exposure where it is placed on the market.

JurisdictionMechanismKey condition for commercial trainingAs of October 2026
United StatesFair use, 17 U.S.C. 107Four-factor, fact-specific; market effect weighs heavily [1]Rulings split by record [2][3]
European UnionDSM Directive Art. 4 TDM exception; AI Act Art. 53(1)(c) requires honoring its opt-outsLawful access; rightsholder opt-outs must be honored; GPAI providers need a copyright policy [5]GPAI duties in force [5]
United KingdomCDPA s29ALawful access and non-commercial research only [7]Government report published 18 March 2026 [8]; s29A unchanged
SingaporeCopyright Act 2021 s244 computational data analysisLawful access to the work [9]In force

For EU market access, the GPAI Code of Practice copyright chapter asks signatories to maintain a copyright policy, avoid circumventing access controls when crawling, and respect machine-readable rights reservations such as robots.txt [6]. A license with the rightsholder is the cleanest way to show that a reserved work was used with permission. In the UK, commercial training on unlicensed copies has no statutory exception to rely on.

Disclosure duties make provenance visible

Transparency laws now force developers to describe training data publicly, which turns an undocumented fair-use bet into a visible one. California AB 2013 required developers of generative AI systems available to Californians to post training-data documentation by 1 January 2026, including dataset sources and whether they contain copyrighted material [10]. EU Article 53 requires GPAI providers to publish a sufficiently detailed summary of training content using the Commission's template [5].

Once you disclose that a corpus includes copyrighted books or news, plaintiffs know where to look. Licensed data lets you disclose a source and a grant rather than a defense theory. Keep a provenance record per dataset: source, acquisition channel, license or legal basis, opt-out checks performed, and retention location.

A decision framework for counsel

Decide per dataset, not per model, because a single training run can mix content with very different risk profiles. Score each dataset on acquisition, market substitution, data type, distribution footprint and contractual overlay, then route it to rely, license or exclude.

Illustrative example: invented to show structure; it does not describe an available dataset.

Dataset situationFair use postureRecommended route
Public-domain or permissively licensed text (CC0, CC BY with attribution handled)Not neededRely on the license; check terms in the open data license compatibility matrix
Purchased print books, scanned in-house, general-purpose pre-training, US-only releaseArguable on current US records [3][4]Counsel judgment; document acquisition and output filtering
Shadow-library or pirated copiesWeak; acquisition exposure independent of training [4]Exclude and purge, including derived shards
Headnotes, summaries or annotations used to build a competing toolWeak after ROSS [2]License; see licensing legal content for AI
News or reference content behind paywalls with opt-outsWeak in the EU; contested in the USLicense; consider grounding APIs vs content licenses
Non-commercial-only datasets (CC BY-NC, research licenses)Not a fair use question; license limits applyExclude from commercial runs; see non-commercial datasets
Proprietary operational data (tickets, code, contracts, workflow logs)Not available without the holder's releaseLicense with a defined training grant

Factors that push toward licensing. Any one of these usually decides it:

  1. Your model's outputs compete with the source in the same market, such as legal research, news summaries or stock imagery.
  2. You cannot document lawful acquisition for every file in the corpus.
  3. The model will be placed on the EU or UK market, or offered to enterprise customers who demand indemnities.
  4. The content is behind access controls, API terms or rights reservations.
  5. You need ongoing refreshes; litigation risk compounds with every new crawl, while a license can cover future deliveries.
  6. Regurgitation risk is material: verbatim memorization of licensed text is a contract issue, while for unlicensed text it is direct evidence of copying.

Scholars have argued that training should generally be allowed because licensing millions of works one by one is impractical [11]. That argument has the most force for diffuse web text and the least for concentrated, high-value sources where a single rightsholder or a small group controls the content and a licensing market already exists, which is exactly where courts weigh market harm [1].

Cost-benefit: what licensing buys beyond litigation risk

Licensing is a procurement decision as much as a legal one, and the benefits extend past avoiding suits. Weigh the license fee against these line items:

  • Expected litigation cost. Defense spend through summary judgment, plus statutory damages exposure per registered work, multiplied by the number of works.
  • Remediation cost. Retraining or unlearning if a dataset must be removed; see what happens to trained models when a license ends.
  • Commercial friction. Enterprise buyers increasingly ask for training-data representations and IP indemnities; see IP indemnities for licensed training data.
  • Data quality. Licensed operational data typically arrives with schema, metadata and context that scraped copies lack.
  • Access. For non-public data, a license is the only route, so the comparison is license versus not having the data.

When you license, make the grant match the use. Pre-training, fine-tuning and retrieval need different rights; the pre-training rights grant guide and the AI training rights grant clause cover drafting.

Pre-signature checklist for a licensed dataset

A license only reduces risk if it actually covers what you will do and the licensor actually holds the rights. Before signing, confirm:

  • The licensor owns or controls the content and has the consents needed to release it for AI training.
  • The grant names the uses: pre-training, fine-tuning, evaluation, retrieval, and successor models.
  • The record scope is defined: tables, date ranges, fields, file formats and volume.
  • Personal data handling is specified: what was removed or replaced, by what method, and how a sample was checked.
  • Term, post-termination model rights and deletion obligations are explicit.
  • Warranties and indemnity are sized to the exposure the license is meant to remove.
  • Delivery is through an access-controlled channel, with a manifest you can reconcile.
  • The provenance record supports your AB 2013 and Article 53 disclosures.

The AI data license negotiation checklist has fallback positions for each item, and the licensing hub maps the rest of the cluster. For the regulatory side, start at AI training data compliance and the SourceX legal framework.

How SourceX fits when licensing is the answer

When your analysis lands on licensing proprietary operational data, SourceX sources it on request from US companies that hold it, including support and sales histories, engineering records, documents, and finance and legal workflows. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery, and every release is approved by the supplying company and delivered under a license that defines records, uses, term and delivery. Describe the data you need on the SourceX buyer page; a request does not guarantee a match.

License the training data you cannot rely on fair use for

If your dataset review flags content where fair use is weak or unavailable, describe the data and the uses you need. SourceX finds US businesses that hold it, assesses data and licensing permissions, and nothing is contracted until a supplier agrees. Start a data request for licensed training data.

Sources

  1. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  2. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (3d Cir.)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  3. Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
  4. Authors Guild, "Court Grants Final Approval of Anthropic Copyright Settlement" (2026). https://authorsguild.org/news/court-grants-final-approval-anthropic-copyright-settlement/
  5. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  6. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  7. UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 (consolidated), section 29A". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  8. UK Government (GOV.UK), "Report on Copyright and Artificial Intelligence" (2026). https://www.gov.uk/government/publications/report-and-impact-assessment-on-copyright-and-artificial-intelligence/report-on-copyright-and-artificial-intelligence
  9. Government of Singapore, "How does Singapore law treat the use of copyright works for AI training". https://isomer-user-content.by.gov.sg/61/7709dff4-3cdd-438d-b767-df5155aa943d/How does Singapore law treat the use of copyright works for AI training.pdf
  10. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  11. Mark A. Lemley and Bryan Casey (SSRN), "Fair Learning" (2020). https://papers.ssrn.com/abstract=3528447

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data