Data licensing for AI training
AI training data licensing: a buyer's guide to rights, terms and pricing
Quick answer
AI training data licensing is the contract route to training, fine-tuning, evaluating or grounding a model on data your team does not own. A workable license settles four things: which uses are granted (training, keeping and shipping the model, using its outputs), who carries the risk if the supplier's rights prove defective, how the price is counted, and whether an open license already covers the use. This hub maps each block to its detailed guide.
By SourceX Editorial · Updated
The four blocks every AI data license has to settle
Every training data license, from click-through terms to a negotiated agreement, must settle four blocks: the rights grant, model and output rights, risk allocation, and price. Disputes usually trace back to one left implicit.
| Block | Questions the contract must answer | Clauses to look for | Detailed guides |
|---|---|---|---|
| Rights grant and scope | Which acts (copy, preprocess, train, evaluate, retrieve), which models, which affiliates and contractors, how long, where? | Grant and definitions, field of use, term and territory, affiliate and processor access | Pre-training rights grant, fine-tuning-only licenses, rights grant clause |
| Model and output rights | Do weights survive termination? May you release them, distill successors or let customers fine-tune? Who owns outputs? | Survival, derivative models, sublicensing, output limits | Model retention after termination, derivative and successor models, output ownership |
| Risk allocation | Does the supplier control the data? Were consents adequate? Who pays if a third party sues? | Warranties, IP indemnity and caps, audit, deletion and return | Data warranties, IP indemnities, audit rights |
| Price and payment | What unit is priced, how is it counted, what triggers further payment? | Fee schedule, unit definitions, refreshes, minimum guarantees | Pricing structures compared, per-token pricing |
Standard dataset licenses such as Creative Commons, CDLA or research-only terms come as fixed texts you cannot negotiate, usually disclaim warranties, and may say little about models or outputs; they get their own section below. For clause definitions, see SourceX's AI data license terms explained and the licensing terms index.
Train, keep, use: the three rights questions buyers conflate
Permission to train does not imply permission to keep the model after the license ends, to release its weights, or to show outputs that reproduce the data. Each needs its own clause.
- May we train? The grant should name the acts your pipeline performs: copying into storage, tokenizing, chunking, embedding, annotating, filtering, building derived datasets, and training or evaluating models. A right to access or display content is not, by itself, a right to copy it into a training corpus. The grant should also name the models covered and who may touch the data (affiliate, contractor and cloud processor access).
- May we keep and ship the model? Weights cannot be "returned" like files, so the contract should state whether they survive termination, whether open-weight release of models trained on licensed data is allowed, and whether customers may fine-tune on the data or receive it under a sublicense.
- May we use the outputs? Suppliers worry about regurgitation because researchers have recovered thousands of training examples from aligned production chat models [1]. Expect output-side terms: the 2024 HarperCollins book-licensing program was described as an opt-in, per-title license that included a commitment to limit verbatim reproduction [2].
Retrieval is a fourth, separate use. A grounding license pays for content fetched and shown at query time, and industry coverage describes publishers moving toward usage-based grounding deals priced separately from training [3]. See grounding licenses vs training licenses and the RAG content licensing hub.
Why buyers license instead of relying on fair use or a TDM exception
As of October 2026, commercial AI developers in the US, EU and UK have no settled, unconditional right to train on content they can access, so a license is how buyers replace legal uncertainty with defined permissions:
- United States. The Copyright Office's Part 3 report on generative AI training, still a May 2025 pre-publication version, concludes that copying works into training datasets may be prima facie infringing absent a defense such as fair use, and recommends no new legislation for now, leaving the licensing market to develop [4]. On 29 September 2026 the Third Circuit held in Thomson Reuters v. Ross (No. 25-2153, precedential) that Westlaw headnotes were copyrightable and that copying them to train a non-generative legal-research tool was not fair use [5]. In Bartz v. Anthropic, the class settlement received final approval in July 2026 [6]. Kadrey v. Meta is ongoing: a June 2025 ruling found fair use for training on that record while other claims continued [7].
- European Union. The text and data mining exception in Article 4 of Directive (EU) 2019/790 does not apply where rightsholders have reserved their rights, for example by machine-readable means, and AI Act Article 53(1)(c) requires general-purpose AI model providers to keep a copyright policy that identifies and honors those reservations [8] (see EU opt-out checks under Article 4).
- United Kingdom. The text and data analysis exception in section 29A of the Copyright, Designs and Patents Act 1988 covers non-commercial research only [9].
Open-web permission is also shrinking: an audit of about 14,000 domains found that between 2023 and 2024, new restrictions made about 5% of all tokens in the C4 corpus, and more than 28% of its most actively maintained critical sources, fully restricted [10]. See license or rely on fair use.
Risk terms: chain of title, warranties, indemnities and audit
Risk terms decide who pays when the supplier's rights prove narrower than the license claims; the classic gap is chain of title, where the signer does not hold every right it grants. The Authors Guild notes that typical trade publishing contracts reserve ungranted rights to the author, so publishers need authors' permission before including books in AI licensing deals [2]. The same gap appears when a business licenses records containing its customers' content.
Require, at minimum:
- Title and authority: the supplier owns or controls the data and may grant the stated uses (warranty of title).
- Lawful acquisition: no data taken from infringing sources or by circumventing access controls.
- Consents: notices and consents cover the licensed use, and any de-identification followed the stated method.
- IP indemnity: scope (third-party claims arising from the data as delivered), cap, and carve-outs for your own modifications.
- Audit, deletion and return: usage reporting, and destruction of raw copies, derived datasets and embeddings at termination (deletion and return clauses).
The license must also leave room for disclosure duties. California's AB 2013 requires developers of generative AI systems offered to Californians to post a summary of their training datasets, including sources or owners, copyright status, whether the data was purchased or licensed, and whether it contains personal information; postings were due by 1 January 2026 and before each later release or substantial modification [11]. EU AI Act Article 53(1)(d) requires a public training-content summary [8], using the template the Commission published on 24 July 2025 [12]. A clause that forbids naming the source can collide with both; see confidentiality clauses vs AI transparency duties and what buyers need from suppliers for EU training-data summaries.
How AI training data licenses are priced
Training data has no public price list; what recurs is a small set of pricing structures, each needing a precise definition of the unit counted.
| Structure | Fits when | Pin down in the contract |
|---|---|---|
| Flat fee | One delivery with a known manifest | Manifest (counts, dates, fields); re-delivery cost |
| Per unit (record, document, audio hour, image) | Volume is selected from a pool | Unit definition; duplicates and rejected units |
| Per token | Text for pre-training or fine-tuning | Which tokenizer; counted before or after deduplication and filtering |
| Subscription or refresh | Data that keeps changing | Cadence; schema-change notice; rights in past deliveries after cancellation |
| Usage-based (per crawl, retrieval or display) | Retrieval and grounding rather than training [3] | Metering method, reporting and audit |
| Revenue share or minimum guarantee | Value depends on your product's success | Revenue definition; attribution across many datasets |
For what moves the price level, see what drives the price of licensed enterprise data, revenue-share data deals and negotiating exclusivity as a buyer.
When an open license is enough for commercial AI training
An open license is enough only when its full text permits your use, it covers every upstream source in the dataset, and it traces back to the party that created the data. Dataset-hub metadata is a weak guide: the Data Provenance Initiative's audit of more than 1,800 text datasets found license omission rates above 70% and error rates above 50% on popular dataset hosting sites [13]. A study of six widely used image datasets found potential license-violation risks in five if used to build commercial AI software, partly because one dataset can mix sources under different licenses [14].
The BEIR retrieval benchmark's paper lists SciFact under CC BY-NC 2.0, several datasets under CC BY-SA, and four of its 19 datasets with no reported license [15]. Microsoft states that MS MARCO datasets are intended for non-commercial research purposes only [16]. Openly licensed corpora do exist: Common Pile v0.1 assembles about 8 TB of public domain and openly licensed text from 30 sources, and its authors report 7B-parameter models trained on it are competitive with models trained on unlicensed text at similar compute [17].
Before relying on an open dataset, check these points (the open data license compatibility matrix compares common licenses):
- Non-commercial (NC) terms: whether training a model you sell or use internally counts as commercial (CC BY-NC datasets and company training).
- Share-alike and copyleft: whether obligations could reach derived datasets or weights (Creative Commons element by element, CDLA-Permissive and CDLA-Sharing).
- Attribution: how you will credit every source in model documentation.
- Gated, custom or research-only terms: click-through conditions on gated and custom-licensed datasets, and whether commercial rights for a research-only dataset can be bought.
Licensing operational business records differs from licensing published content
Operational records such as support tickets, engineering histories, contracts or finance workflows raise problems that news, books or stock images rarely do: they hold other people's personal and confidential information, and the company's earlier promises to those people limit what it may license. Three checks matter:
- Prior privacy commitments. In a February 2024 staff blog post, the FTC's Office of Technology warned that adopting more permissive practices, such as using data for AI training or sharing it with third parties, through a surreptitious, retroactive change to terms of service or a privacy policy may be unfair or deceptive [18].
- De-identification the law recognizes. Under the CCPA, data counts as deidentified only if, among other conditions, the business contractually binds recipients to the definition's requirements, so expect a no-reidentification clause [19]. For protected health information, HIPAA recognizes two de-identification methods, expert determination and safe harbor [20]; SourceX requires one of them before health records are considered for a license.
- Sector redisclosure limits. Under Regulation P, a company that receives nonpublic personal information from a nonaffiliated financial institution may reuse and redisclose it only within limits tied to how it was received [21].
In SourceX's process, rights review checks that the business owns or may share the records and that required consents are in place; names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked. No method is perfect, so keep no-reidentification terms anyway. See the de-identified data hub and industry operational data hub.
Choosing a route: open license, direct license or managed sourcing
Ask, in order, whether usable terms already exist, who holds the rights, and whether someone must find the holder first.
- An open dataset whose license text permits your use. Use it, keep license texts and attribution records with the training run, and document training data provenance.
- A known rights holder that licenses its content (a publisher, data vendor or research consortium). Negotiate directly, starting from a data license term sheet and the negotiation checklist with fallback positions.
- Data that exists only inside operating businesses and has never been offered for license. Someone must find holders, obtain their approval, run rights review and de-identification, and paper the deal. SourceX sources operational datasets from US companies and manages that commercial process, including licensing agreements and ongoing purchases; buyers describe the data, not the businesses. Datasets are sourced on request, so a request does not guarantee a match (how SourceX works with data buyers).
- Data that does not exist yet. Commission collection and decide between an IP assignment and a license for commissioned data.
For the step-by-step process once a route is chosen, see SourceX's guide to licensing proprietary data for AI training. Adjacent questions have their own guides: compliance duties for training data, procurement from requirements to renewal, evaluation-only license terms and in-house counsel's license review. The AI data buyer hub maps every cluster.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
License operational data from US businesses
If the data you need sits inside US companies rather than in a public corpus, describe it on the buyers page: the records, fields, history and uses you need licensed. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, manages the license, and coordinates delivery and future purchases. Start a data request with SourceX.
Guides in this section
- AI Data License Negotiation Checklist: Asks and FallbacksRanked checklist for negotiating AI training data licenses: 14 terms, each with a first ask, a fallback and a walk-away signal, plus licensor redlines.
- AI Data License Term Sheet: A Buyer's TemplateA buyer's term sheet for AI training data licenses: 20 fields with an AI check for each, which clauses bind before signing, and a filled example.
- AI Data Licensing Pricing Models: Choosing a Fee StructureFlat fee, per-record, per-token, subscription or revenue share: how to pick an AI data license pricing model and compare quotes in different units.
- AI Training Rights Grant Clause: Definitions and Sample TextHow to draft an AI training rights grant clause: grant verbs, qualifiers, and definitions of Training, Model, Derived Data and Output, with sample text.
- CC BY-NC Datasets and Commercial Model TrainingCan a company train on CC BY-NC or research-only data? What NonCommercial covers, when terms reach model weights, and routes to commercial rights.
- Data License Termination: What Happens to Trained ModelsMust you delete a model trained on licensed data when the license ends? Compare survival, run-off, retrain and unlearning terms and what to insist on.
- Data License Warranties for AI Training: What to RequireWhich warranties to require in an AI training data license: title, rights, consents, privacy, file integrity, accuracy limits, qualifiers and survival.
- Dataset Licenses for Commercial AI Training: A MatrixWhich open dataset licenses allow commercial AI training? A matrix of CC, ODC, CDLA, MIT, Apache and RAIL terms, with attribution and share-alike rules.
- Derivative and Successor Model Rights in AI Data LicensesHow to define Model in an AI data license so rights reach successor versions and distilled, quantized or merged models: three drafting patterns compared.
- Fine-Tuning Data License: Base Models, Adapters, DeploymentWhat a fine-tuning-only data license should cover: base-model scope, adapters and distilled models, deployment tiers, weight retention and upgrade terms.
- Pre-Training Data License: Grant, Model Scope, RetentionWhat a pre-training data license must grant: pipeline processing rights, a Model definition that covers successors, weight retention and output terms.
- Training Data Indemnification: Scope, Caps and Carve-OutsNegotiating training data indemnification: which third-party claims to cover, exclusions to narrow, IP super-caps, defense control and settlement consent.
- Audit and Usage-Reporting Rights in AI Data LicensesHow AI data buyers negotiate audit rights: certification first, scoped independent audits, records to keep, trade-secret limits and cost-shifting triggers.
- Authorized Users Clause: Affiliates, Contractors, CloudsHow to draft an authorized users clause in an AI data license so affiliates, annotation vendors, eval contractors and cloud hosts can use the data.
- CDLA-Permissive-2.0 and CDLA-Sharing-1.0 for AI TrainingWhat CDLA-Permissive-2.0 and CDLA-Sharing-1.0 allow for AI training: the data grant, Results and trained models, notice duties, and CC BY compared.
- Collective Licensing for AI Training: Scope and GapsHow collective and blanket licenses from collecting societies cover AI use of published content, what they exclude, and when a direct license is needed.
- Commissioned Training Data: IP Assignment vs LicenseWho owns custom-collected AI training data? Compare IP assignment, exclusive and non-exclusive licenses, contributor chain of title and vendor reuse.
- Creative Commons and AI Training: BY, SA, NC, ND ExplainedHow CC BY, SA, NC and ND apply when CC-licensed text and images train models: when conditions bite, attribution at scale, versions and buyer checks.
- Data License Confidentiality vs Training Data DisclosureHow to draft data license confidentiality terms so you can still describe licensed sources in AB 2013 notices, EU AI Act summaries and model cards.
- Data License vs Data Use Agreement for AI Data DealsData license, DUA, data sharing agreement, DPA or NDA: which fits an AI training data deal, what each must contain, and when you need several.
- Data-for-Model-Access Deals vs Cash Data LicensesHow data-for-model-access partnerships differ from cash AI data licenses: model ownership, exclusivity, valuing credits and co-development, and exit terms.
- Deletion and Return Clauses for Licensed AI Training DataHow to scope deletion and return clauses in AI training data licenses: raw copies, shards, embeddings, backups, processors, certificates and weights.
- Field-of-Use Restrictions in AI Data Licenses: DraftingHow to draft field-of-use restrictions in AI data licenses: vertical, product, modality and customer fields, objective tests, and the price trade-off.
- Gated Dataset Terms of Use: A Commercial Review ChecklistHow to read gated and custom-licensed dataset terms on model hubs before training: who may accept, commercial use, model distribution and audit records.
- License or Fair Use? Deciding How to Source AI Training DataA decision framework for counsel: when to license AI training data instead of relying on fair use or TDM exceptions, given 2025-2026 rulings and EU rules.
- Licensed Data in Hosted Fine-Tuning APIs: What to CheckCan you upload licensed data to a fine-tuning API? Check the license grant, processor terms, retention, region and model ownership before the first upload.
- Licensing Case Law, Headnotes and Treatises for Legal AIWhich legal materials are public domain and which editorial layers need a license: case law, headnotes, annotations, treatises and filings for AI.
- Licensing Images for AI Training: Terms and ReleasesWhich image licenses actually permit AI training: stock vs training grants, editorial limits, model and property releases, logos, captions and metadata.
- Market Data Licenses for AI Training: Non-Display RulesHow exchange and vendor market data licenses treat AI training: non-display use, derived data, model outputs and audits, plus a rider checklist.
- Marketplace Dataset Licenses: Checking AI Training RightsHow to read data marketplace subscription and listing terms for AI training rights, spot silent or restrictive clauses, and negotiate a private offer.
- Master Data License Agreements and Order Forms for AI DataHow to structure a master data license agreement with order forms and dataset schedules so repeat AI training data purchases do not need renegotiation.
- Negotiating Exclusive AI Data Licenses: Buyer StructuresHow AI data buyers negotiate exclusivity: field-limited, time-boxed lockout, first refusal and holdback structures, with verification, remedies and MFN.
- Open-Weight Release Rights for Models on Licensed DataCan you publish open weights trained on licensed data? The license clauses, memorization tests and flow-down terms licensors ask for before weight release.
- Per-Token Training Data Pricing: Count, Audit, CompareHow to price licensed training text per token: pin the tokenizer, define the dedup and filtering basis, verify counts within tolerance and compare quotes.
- Perpetual vs Term AI Training Data Licenses and TerritoryHow to split an AI training data license into access term, use term and perpetual model rights, and draft territory for globally served models.
- Prohibited-Use Clauses in AI Data Licenses: Buyer GuideHow to review prohibited-use lists in AI data licenses: competitor, surveillance, re-identification and open-weights bans, plus a redline checklist.
- Research-Only Dataset to Commercial License: 4 RoutesPrototyped on a research-only corpus? See 4 routes to commercial AI training rights: rights-holder license, consortium tier, re-collection, replacement.
- Revenue-Share AI Data Deals: Base, Attribution, ReportingHow AI buyers should structure revenue-share data licenses: the revenue base, attributing one dataset among many, reporting duties and when to say no.
- Share-Alike Data Licenses and Trained AI ModelsDoes share-alike reach model weights or outputs? How CC BY-SA, ODbL and CDLA-Sharing triggers work in AI training and how model teams manage the risk.
- Sublicensing Licensed Training Data to Platform CustomersHow AI platforms get rights to let customers fine-tune or ground on licensed data: sublicense grants, flow-down terms, reporting and the liability chain.
- Subscription Data Licenses for AI Training RefreshesWhat rights each delivered batch keeps after an AI training data subscription ends, how refresh licenses are priced, and the clauses to negotiate first.
- Synthetic Data From Licensed Data: Derived Dataset RightsCan licensed data seed synthetic data, and who owns the output after term? Clause terms, generator-model limits, privacy tests and a Derived Data rider.
- Who Owns Model Outputs Under a Training Data LicenseDoes a data licensor own outputs from a model trained on its data? How to draft output rights, regurgitation limits and confidentiality carve-outs.
Sources
- Nasr et al., ICLR 2025, "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
- Authors Guild, "HarperCollins AI licensing deal" (2024). https://authorsguild.org/news/harpercollins-ai-licensing-deal/
- Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?" https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 (consolidated), section 29A". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
- Longpre et al., NeurIPS 2024, "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Rajbahadur et al., "Can I use this publicly available dataset to build commercial AI software? Most likely not" (2021). https://arxiv.org/abs/2111.02374v4
- Thakur et al., "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
- Microsoft, "MS MARCO Datasets". https://microsoft.github.io/msmarco/Datasets.html
- Kandpal et al., "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- eCFR / U.S. Department of Health and Human Services, "45 CFR 164.514 (de-identification of protected health information)". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.