Skip to content

Data licensing for AI training

AI training data licensing: a buyer's guide to rights, terms and pricing

Quick answer

AI training data licensing is the contract route to training, fine-tuning, evaluating or grounding a model on data your team does not own. A workable license settles four things: which uses are granted (training, keeping and shipping the model, using its outputs), who carries the risk if the supplier's rights prove defective, how the price is counted, and whether an open license already covers the use. This hub maps each block to its detailed guide.

By SourceX Editorial · Updated

The four blocks every AI data license has to settle

Every training data license, from click-through terms to a negotiated agreement, must settle four blocks: the rights grant, model and output rights, risk allocation, and price. Disputes usually trace back to one left implicit.

BlockQuestions the contract must answerClauses to look forDetailed guides
Rights grant and scopeWhich acts (copy, preprocess, train, evaluate, retrieve), which models, which affiliates and contractors, how long, where?Grant and definitions, field of use, term and territory, affiliate and processor accessPre-training rights grant, fine-tuning-only licenses, rights grant clause
Model and output rightsDo weights survive termination? May you release them, distill successors or let customers fine-tune? Who owns outputs?Survival, derivative models, sublicensing, output limitsModel retention after termination, derivative and successor models, output ownership
Risk allocationDoes the supplier control the data? Were consents adequate? Who pays if a third party sues?Warranties, IP indemnity and caps, audit, deletion and returnData warranties, IP indemnities, audit rights
Price and paymentWhat unit is priced, how is it counted, what triggers further payment?Fee schedule, unit definitions, refreshes, minimum guaranteesPricing structures compared, per-token pricing

Standard dataset licenses such as Creative Commons, CDLA or research-only terms come as fixed texts you cannot negotiate, usually disclaim warranties, and may say little about models or outputs; they get their own section below. For clause definitions, see SourceX's AI data license terms explained and the licensing terms index.

Train, keep, use: the three rights questions buyers conflate

Permission to train does not imply permission to keep the model after the license ends, to release its weights, or to show outputs that reproduce the data. Each needs its own clause.

  • May we train? The grant should name the acts your pipeline performs: copying into storage, tokenizing, chunking, embedding, annotating, filtering, building derived datasets, and training or evaluating models. A right to access or display content is not, by itself, a right to copy it into a training corpus. The grant should also name the models covered and who may touch the data (affiliate, contractor and cloud processor access).
  • May we keep and ship the model? Weights cannot be "returned" like files, so the contract should state whether they survive termination, whether open-weight release of models trained on licensed data is allowed, and whether customers may fine-tune on the data or receive it under a sublicense.
  • May we use the outputs? Suppliers worry about regurgitation because researchers have recovered thousands of training examples from aligned production chat models [1]. Expect output-side terms: the 2024 HarperCollins book-licensing program was described as an opt-in, per-title license that included a commitment to limit verbatim reproduction [2].

Retrieval is a fourth, separate use. A grounding license pays for content fetched and shown at query time, and industry coverage describes publishers moving toward usage-based grounding deals priced separately from training [3]. See grounding licenses vs training licenses and the RAG content licensing hub.

Why buyers license instead of relying on fair use or a TDM exception

As of October 2026, commercial AI developers in the US, EU and UK have no settled, unconditional right to train on content they can access, so a license is how buyers replace legal uncertainty with defined permissions:

  • United States. The Copyright Office's Part 3 report on generative AI training, still a May 2025 pre-publication version, concludes that copying works into training datasets may be prima facie infringing absent a defense such as fair use, and recommends no new legislation for now, leaving the licensing market to develop [4]. On 29 September 2026 the Third Circuit held in Thomson Reuters v. Ross (No. 25-2153, precedential) that Westlaw headnotes were copyrightable and that copying them to train a non-generative legal-research tool was not fair use [5]. In Bartz v. Anthropic, the class settlement received final approval in July 2026 [6]. Kadrey v. Meta is ongoing: a June 2025 ruling found fair use for training on that record while other claims continued [7].
  • European Union. The text and data mining exception in Article 4 of Directive (EU) 2019/790 does not apply where rightsholders have reserved their rights, for example by machine-readable means, and AI Act Article 53(1)(c) requires general-purpose AI model providers to keep a copyright policy that identifies and honors those reservations [8] (see EU opt-out checks under Article 4).
  • United Kingdom. The text and data analysis exception in section 29A of the Copyright, Designs and Patents Act 1988 covers non-commercial research only [9].

Open-web permission is also shrinking: an audit of about 14,000 domains found that between 2023 and 2024, new restrictions made about 5% of all tokens in the C4 corpus, and more than 28% of its most actively maintained critical sources, fully restricted [10]. See license or rely on fair use.

Risk terms: chain of title, warranties, indemnities and audit

Risk terms decide who pays when the supplier's rights prove narrower than the license claims; the classic gap is chain of title, where the signer does not hold every right it grants. The Authors Guild notes that typical trade publishing contracts reserve ungranted rights to the author, so publishers need authors' permission before including books in AI licensing deals [2]. The same gap appears when a business licenses records containing its customers' content.

Require, at minimum:

  1. Title and authority: the supplier owns or controls the data and may grant the stated uses (warranty of title).
  2. Lawful acquisition: no data taken from infringing sources or by circumventing access controls.
  3. Consents: notices and consents cover the licensed use, and any de-identification followed the stated method.
  4. IP indemnity: scope (third-party claims arising from the data as delivered), cap, and carve-outs for your own modifications.
  5. Audit, deletion and return: usage reporting, and destruction of raw copies, derived datasets and embeddings at termination (deletion and return clauses).

The license must also leave room for disclosure duties. California's AB 2013 requires developers of generative AI systems offered to Californians to post a summary of their training datasets, including sources or owners, copyright status, whether the data was purchased or licensed, and whether it contains personal information; postings were due by 1 January 2026 and before each later release or substantial modification [11]. EU AI Act Article 53(1)(d) requires a public training-content summary [8], using the template the Commission published on 24 July 2025 [12]. A clause that forbids naming the source can collide with both; see confidentiality clauses vs AI transparency duties and what buyers need from suppliers for EU training-data summaries.

How AI training data licenses are priced

Training data has no public price list; what recurs is a small set of pricing structures, each needing a precise definition of the unit counted.

StructureFits whenPin down in the contract
Flat feeOne delivery with a known manifestManifest (counts, dates, fields); re-delivery cost
Per unit (record, document, audio hour, image)Volume is selected from a poolUnit definition; duplicates and rejected units
Per tokenText for pre-training or fine-tuningWhich tokenizer; counted before or after deduplication and filtering
Subscription or refreshData that keeps changingCadence; schema-change notice; rights in past deliveries after cancellation
Usage-based (per crawl, retrieval or display)Retrieval and grounding rather than training [3]Metering method, reporting and audit
Revenue share or minimum guaranteeValue depends on your product's successRevenue definition; attribution across many datasets

For what moves the price level, see what drives the price of licensed enterprise data, revenue-share data deals and negotiating exclusivity as a buyer.

When an open license is enough for commercial AI training

An open license is enough only when its full text permits your use, it covers every upstream source in the dataset, and it traces back to the party that created the data. Dataset-hub metadata is a weak guide: the Data Provenance Initiative's audit of more than 1,800 text datasets found license omission rates above 70% and error rates above 50% on popular dataset hosting sites [13]. A study of six widely used image datasets found potential license-violation risks in five if used to build commercial AI software, partly because one dataset can mix sources under different licenses [14].

The BEIR retrieval benchmark's paper lists SciFact under CC BY-NC 2.0, several datasets under CC BY-SA, and four of its 19 datasets with no reported license [15]. Microsoft states that MS MARCO datasets are intended for non-commercial research purposes only [16]. Openly licensed corpora do exist: Common Pile v0.1 assembles about 8 TB of public domain and openly licensed text from 30 sources, and its authors report 7B-parameter models trained on it are competitive with models trained on unlicensed text at similar compute [17].

Before relying on an open dataset, check these points (the open data license compatibility matrix compares common licenses):

Licensing operational business records differs from licensing published content

Operational records such as support tickets, engineering histories, contracts or finance workflows raise problems that news, books or stock images rarely do: they hold other people's personal and confidential information, and the company's earlier promises to those people limit what it may license. Three checks matter:

  • Prior privacy commitments. In a February 2024 staff blog post, the FTC's Office of Technology warned that adopting more permissive practices, such as using data for AI training or sharing it with third parties, through a surreptitious, retroactive change to terms of service or a privacy policy may be unfair or deceptive [18].
  • De-identification the law recognizes. Under the CCPA, data counts as deidentified only if, among other conditions, the business contractually binds recipients to the definition's requirements, so expect a no-reidentification clause [19]. For protected health information, HIPAA recognizes two de-identification methods, expert determination and safe harbor [20]; SourceX requires one of them before health records are considered for a license.
  • Sector redisclosure limits. Under Regulation P, a company that receives nonpublic personal information from a nonaffiliated financial institution may reuse and redisclose it only within limits tied to how it was received [21].

In SourceX's process, rights review checks that the business owns or may share the records and that required consents are in place; names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked. No method is perfect, so keep no-reidentification terms anyway. See the de-identified data hub and industry operational data hub.

Choosing a route: open license, direct license or managed sourcing

Ask, in order, whether usable terms already exist, who holds the rights, and whether someone must find the holder first.

  1. An open dataset whose license text permits your use. Use it, keep license texts and attribution records with the training run, and document training data provenance.
  2. A known rights holder that licenses its content (a publisher, data vendor or research consortium). Negotiate directly, starting from a data license term sheet and the negotiation checklist with fallback positions.
  3. Data that exists only inside operating businesses and has never been offered for license. Someone must find holders, obtain their approval, run rights review and de-identification, and paper the deal. SourceX sources operational datasets from US companies and manages that commercial process, including licensing agreements and ongoing purchases; buyers describe the data, not the businesses. Datasets are sourced on request, so a request does not guarantee a match (how SourceX works with data buyers).
  4. Data that does not exist yet. Commission collection and decide between an IP assignment and a license for commissioned data.

For the step-by-step process once a route is chosen, see SourceX's guide to licensing proprietary data for AI training. Adjacent questions have their own guides: compliance duties for training data, procurement from requirements to renewal, evaluation-only license terms and in-house counsel's license review. The AI data buyer hub maps every cluster.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

License operational data from US businesses

If the data you need sits inside US companies rather than in a public corpus, describe it on the buyers page: the records, fields, history and uses you need licensed. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, manages the license, and coordinates delivery and future purchases. Start a data request with SourceX.

Guides in this section

Sources

  1. Nasr et al., ICLR 2025, "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
  2. Authors Guild, "HarperCollins AI licensing deal" (2024). https://authorsguild.org/news/harpercollins-ai-licensing-deal/
  3. Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?" https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
  4. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  5. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  6. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  7. Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
  8. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  9. UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 (consolidated), section 29A". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  10. Longpre et al., NeurIPS 2024, "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
  11. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  12. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  13. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  14. Rajbahadur et al., "Can I use this publicly available dataset to build commercial AI software? Most likely not" (2021). https://arxiv.org/abs/2111.02374v4
  15. Thakur et al., "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
  16. Microsoft, "MS MARCO Datasets". https://microsoft.github.io/msmarco/Datasets.html
  17. Kandpal et al., "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
  18. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  19. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  20. eCFR / U.S. Department of Health and Human Services, "45 CFR 164.514 (de-identification of protected health information)". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  21. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data