Data licensing for AI training
Writing the AI training rights grant clause: definitions and sample language
Quick answer
An AI training rights grant clause should list the acts it permits instead of relying on the word "train." Name who may exercise the rights, the operative verbs (reproduce, store, modify, combine, create derived data), the qualifiers (exclusivity, territory, duration, sublicensing) and the purpose. Then define Training, Model, Derived Data and Output so every pipeline step and artifact falls clearly inside or outside the grant, and keep the reservation of rights from taking back what the grant gives.
By SourceX Editorial · Updated
Why "the right to train" is not a grant
A one-line right to "train AI models on the Data" leaves the copying and transformation around training unaddressed and says nothing about which artifacts you keep. The U.S. Copyright Office's Part 3 report, which as of October 2026 is still the May 2025 pre-publication version, concludes that many acts in AI training, including copying and organizing works into datasets, may be prima facie infringing unless a license or an exception such as fair use covers them [1]. That copying happens in storage and preprocessing, before any gradient step.
Unwritten scope becomes a dispute. In one practitioner case study of an imagery-for-training license, the parties had not written down what training use covered or whether the trained model was a derivative work, and the fix was to define the scope of training use in the contract [2]. Relying on fair use instead is a fact-specific bet: on 29 September 2026 the Third Circuit held that ROSS's use of Westlaw headnotes to train a non-generative legal-research tool was not fair use [3] (see license or rely on fair use).
The grant sentence, element by element
Every training grant combines the same elements (grantee, operative verbs, qualifiers, rights base and purpose), and a licensor can narrow any of them without touching the word "train."
| Element | Typical licensor draft | Buyer position | Why it matters |
|---|---|---|---|
| Grantee | "Licensee" | Licensee, its affiliates, and contractors and cloud providers acting for them | Annotation vendors and cloud GPUs sit outside "Licensee" (affiliates and contractors clause) |
| Operative verbs | "use" | Use, reproduce, store, modify, annotate, combine with other data, create Derived Data | Copying and modification happen before training |
| Exclusivity | Non-exclusive | Non-exclusive, unless you pay for a time- or field-limited exclusive license | See negotiating exclusivity as a buyer |
| Territory and site | "At Licensee's premises" | Worldwide, or named cloud regions | Cloud training runs elsewhere |
| Duration | "Revocable," "during the Term" | Data access for the Term; rights in trained Models perpetual and irrevocable | A revocable grant can strand a deployed model (models after termination) |
| Transfer | "Non-transferable, non-sublicensable" | Sublicensable to Permitted Users; assignable with the business | Acquisitions and platform customers (change of control; sublicensing to customers) |
| Rights base | "Under Licensor's copyrights" | Under all intellectual property and other rights Licensor holds | Ticket fields or transaction logs may carry thin copyright, so the license works mainly as a contract |
| Purpose | "Internal research" | Training, Evaluation, and development, deployment and commercial use of Models | "Research" can exclude production (field-of-use drafting) |
Published licenses show these qualifiers at work. NVIDIA's sample data license for evaluation grants a limited, non-exclusive, revocable, non-transferable, non-sublicensable license solely for evaluating and testing NVIDIA technologies [4]; that shape fits a sample, not production training (evaluation-only license terms). The Linguistic Data Consortium's for-profit membership agreement licenses data for linguistic and language-based research and technology development, solely at sites listed in an exhibit [5], a limit that could block training in an unlisted cloud region. A satellite-imagery license hosted by Planet makes its grant irrevocable except as provided in a named article [6], the pattern a buyer wants for trained models.
Defining Training by what changes the model
Define Training functionally, as any process that uses the licensed data to compute, adjust or select a model's parameters, then list acts as examples. Post-training alone involves distinct acts: the InstructGPT pipeline fine-tuned on labeler-written demonstrations, trained a reward model on human rankings of outputs, then optimized against it with reinforcement learning [7], while Direct Preference Optimization (DPO) fits the policy to preference data without a separate reward model [8]. A grant that names only "fine-tuning" may not clearly reach the reward model.
List these acts after "including, without limitation":
- Preparation: copying, format conversion (for example to JSONL or Parquet), filtering, de-identification, deduplication, tokenization and mixing with other data.
- Pre-training and continued pre-training, including tokenizer training (pre-training rights).
- Supervised fine-tuning, including adapters (fine-tuning-only licenses).
- Preference optimization, reward models and reinforcement learning.
- Distillation and synthetic example generation (synthetic data from licensed data).
- Search and selection: hyperparameter sweeps, ablations and checkpoint selection. "Select" matters because a validation split used for early stopping shapes the model without any gradient update.
Define Evaluation separately, because scoring models on held-out records changes no parameters and may involve third-party models. Keep retrieval out of Training too: indexing records or their embeddings for lookup at inference time is a different use, which the Copyright Office report discusses as involving reproduction [1] (embedding and vector index rights).
Model, Derived Data and Output: draft by how each artifact is made
These definitions decide what you may keep after training, so define each artifact by how it is produced, not by whether copyright law would call it a derivative work.
Model covers weights, checkpoints, adapters and reward models produced through Training, "whether or not a derivative work under applicable law." Whether it reaches successor generations, distilled students and merges is a scope choice compared in derivative and successor model rights.
Derived Data is everything generated from the records that is not a model: annotations, labels, quality scores, embeddings, tokenized shards, filtered subsets and evaluation splits. Split it in two. Most subword tokenizers are reversible, so tokenized shards and filtered subsets can reproduce the records and belong with the licensed data for deletion; scores, labels and aggregate statistics cannot, and a buyer can ask to own them. Licensors define this layer too: the Planet-hosted license defines a "Value-Added Product" derived from imagery through technical manipulation or added data [6].
Output is content a Model generates in operation. Lemley and Henderson argue that model outputs lack the human authorship copyright requires, so restrictions on their use rest on contract [9]. A workable buyer position: Outputs are neither Licensed Data nor Derived Data and the licensor claims no rights in them, except Output that reproduces a licensed record verbatim or nearly so. See who owns model outputs.
Sample grant and definitions for a buyer's first draft
This is a buyer's opening position for a non-exclusive training license; expect pushback on survival, Evaluation and Permitted Users.
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
1. DEFINITIONS
"Licensed Data" means the records listed in the Delivery Manifest by record
ID and SHA-256 hash, including corrected and replacement deliveries, in any
format.
"Permitted Users" means Licensee, its Affiliates, and contractors and cloud
service providers acting on behalf of either, each bound by written terms at
least as protective of the Licensed Data as this Agreement.
"Training" means any process that uses Licensed Data to compute, initialize,
adjust or select the parameters of a machine learning model, including
without limitation:
(a) copying, storing, converting, filtering, de-identifying, deduplicating,
tokenizing, sampling and mixing Licensed Data with other data;
(b) training tokenizers, pre-training and continued pre-training;
(c) supervised fine-tuning, including adapter and low-rank methods;
(d) preference optimization, reward-model training and reinforcement
learning;
(e) distillation and generating synthetic training examples; and
(f) hyperparameter search, ablations and checkpoint selection.
"Evaluation" means measuring the performance, safety or behavior of any
machine learning model, including a third party's model, using Licensed Data
held out from Training, and reporting aggregate results.
"Model" means any machine learning model, including its weights, parameters,
checkpoints, adapters and any reward model, created through Training by or
for a Permitted User, whether or not it is a derivative work under
applicable law.
"Derived Data" means data generated from Licensed Data other than Models,
including annotations, labels, quality scores, embeddings, tokenized shards,
filtered subsets and evaluation splits. "Reconstructive Derived Data" means
Derived Data from which any Licensed Data record can be reproduced; it is
treated as Licensed Data.
"Output" means content generated by a Model in operation. Output is not
Licensed Data or Derived Data, except Output that reproduces a Licensed
Data record verbatim or near-verbatim.
"Retrieval Use" means storing Licensed Data, or embeddings of it, in an
index that a model queries at inference time.
"Purpose" means Training, Evaluation, and the development, deployment and
commercial use of Models and Outputs, excluding Retrieval Use unless
Schedule 3 grants it.
2. LICENSE GRANT
2.1 Licensor grants Licensee a non-exclusive, worldwide license, under all
intellectual property and other rights Licensor holds in the Licensed
Data, for Permitted Users to use, reproduce, store, modify, annotate,
translate, combine with other data and create Derived Data from the
Licensed Data, in each case for the Purpose during the Term.
2.2 Licensee may sublicense the rights in Section 2.1 only to Permitted
Users, and is responsible for their compliance.
2.3 Licensee's rights in Models, Outputs and Derived Data other than
Reconstructive Derived Data are perpetual, irrevocable and fully
paid-up, and survive expiry or termination of this Agreement.
2.4 Licensor reserves all rights not expressly granted in this Agreement,
which includes its Schedules, Order Forms and the Delivery Manifest. No
terms of use, click-through or portal terms presented with access to or
delivery of the Licensed Data apply to Licensee or limit this Section 2.
Why it is drafted this way:
- Functional lead-in. The lead-in carries the definition and the lettered list illustrates it, so an unlisted act that adjusts parameters is still covered.
- Reconstructive Derived Data. Tokenized shards follow the deletion clause; scores and labels stay with you.
- Retrieval excluded by default. Its price is negotiated, not assumed.
- Likely counter-proposals. Evaluation limited to Licensee's own models, survival lost on termination for uncured breach of the restrictions, and a named contractor list instead of a class. Fallback positions are in the AI data license negotiation checklist.
Worked example: testing the draft against one pipeline
Before the first redline, walk your pipeline through the draft and point to the words that cover each step. In this invented scenario, a buyer licenses three years of business-to-business support tickets to build a support agent.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Pipeline step | Covered in the sample by | Under "use the Data to train Licensee's AI models" |
|---|---|---|
| Copy exports into a cloud bucket in another region | 2.1 (reproduce, store); the cloud provider is a Permitted User | Copying and third-party hosting not addressed |
| Redact personal data and remove near-duplicates | Training (a) | Modifying the records not addressed |
| Fine-tune an open-weight base model with LoRA adapters | Training (c); Model | Covered |
| Train a reward model on rated agent replies, then run reinforcement learning | Training (d); Model includes reward models | Unclear whether a reward model is one of "Licensee's AI models" |
| Score three vendors' models on a held-out split | Evaluation | Not training, and the models are not Licensee's |
| Distill the agent into a small on-device model | Training (e); Model | Disputed: the student never saw the records |
| Send a sample to an annotation vendor for intent labels | Permitted Users; Derived Data | Disclosure to a third party not authorized |
| Index ticket resolutions for a retrieval assistant | Excluded until Schedule 3 grants Retrieval Use | Unclear either way |
| Keep serving the agent after the term ends | 2.3 | The model may have to be retired |
Two rows need business decisions, not drafting: whether retrieval is in the deal, and whether a distilled student counts as a Model. Record both in the AI data license term sheet.
Reservations, incorporated terms and other clauses that shrink the grant
A reservation of rights is reasonable, but check that it, and every other clause, leaves the express grant intact. "All rights not expressly granted are reserved" is standard; confirm that "this Agreement" includes schedules, order forms and the delivery manifest, so rights granted there count as express.
Watch for terms that arrive with access: Stack Exchange placed its data dump behind a login and an agreement not to use the content to train AI models, which critics said conflicted with the Creative Commons license on the content [10]. A delivery portal, API or click-through can add a no-training rule the same way, so state that no access terms limit the grant.
Then check three clauses against the grant:
- Confidentiality. Article 53(1)(d) of the EU AI Act requires providers of general-purpose AI models to publish a sufficiently detailed summary of training content [11]. California's AB 2013 requires developers to post documentation, including whether datasets contain copyrighted or licensed material, first by 1 January 2026 and before later releases [12]. Carve out legally required disclosures (confidentiality vs transparency duties).
- Deletion. "Delete all copies and derivatives" can be read to reach Models; exclude Models and non-reconstructive Derived Data (deletion and return clauses).
- Restrictions. A prohibited-use list should restrict named uses, not redefine the grant (prohibited-use clauses).
The licensor can only grant what it holds
A grant passes only the rights the licensor has, so pair it with an authority representation and ask how the records were collected. The Authors Guild notes that typical trade publishing agreements grant rights for publication in book and excerpt forms and reserve all other rights to the author, so publishers need authors' permission before including books in AI licensing deals [13]. Businesses licensing customer-submitted content face the same question.
Customer terms can also undercut the grant. In a February 2024 Tech@FTC blog post, FTC technology staff wrote that it may be unfair or deceptive for a company to adopt more permissive data practices, such as AI training, while telling users only through a surreptitious, retroactive change to its terms of service or privacy policy [14]. Ask when the relevant terms took effect relative to when the records were collected.
Privacy law adds conditions that travel with the grant. Under California Civil Code section 1798.140(m) as codified in October 2026, information counts as deidentified only if the business holding it publicly commits not to re-identify it and contractually obligates recipients to comply with the same requirements [15], so expect re-identification prohibition language beside the grant. Back authority with data warranties and an IP indemnity for licensed training data.
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Every dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and diligence materials on source, rights, preparation and allowed use are prepared per dataset for the buyer's review. State the training uses you need when you submit a data request to SourceX.
Phrases to search for before signature
Most narrow grants trace back to a few phrases; search the draft for each and replace it with a defined term.
- "use the Data to train AI": no copying, modification or Derived Data rights.
- "internal research purposes": may exclude production deployment and customer-facing features.
- "Licensee's models": may exclude reward models, distilled students and affiliates' models.
- "solely at Licensee's facilities": collides with training in cloud regions.
- "in accordance with Licensor's policies, as amended from time to time": lets the licensor narrow the grant on its own.
For plain-language background, see SourceX's guide to AI data license terms, its clause-by-clause agreement guide and its glossary of licensing terms. The AI training data licensing hub maps the other clauses.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Know which training uses you need licensed?
Describe the data you need and the training, evaluation and deployment uses the license must cover. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license, delivery and future purchases. Datasets are sourced on request, so a request does not guarantee a match. Submit your licensing requirements.
Sources
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- Terms.law, "AI and data licensing" (practitioner case study). https://terms.law/case-studies/ai-data-licensing-archive-imagery.html
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- NVIDIA, "NVIDIA Sample Data License for Evaluation (version 2026.01.19)" (2026). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
- Linguistic Data Consortium (University of Pennsylvania), "LDC For-Profit Membership Agreement". https://Catalog.Ldc.Upenn.Edu/license/ldc-for-profit-membership.pdf
- Planet (hosted license document), "Geomatics EULA SPOT PLEIADES 2023.1" (2023). https://assets.planet.com/docs/Geomatics_EULA_SPOT_PLEIADES_2023.1.pdf
- Ouyang et al. (OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Rafailov et al. (Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- SpicyIP, "Discussing Lemley and Henderson's 'The Mirage of Artificial Intelligence Terms of Use Restrictions'" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
- DevClass, "Stack Exchange restricts access to dump of user-contributed data as critics complain license permits reuse for any purpose" (2024). https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Authors Guild, "HarperCollins AI licensing deal" (2024). https://authorsguild.org/news/harpercollins-ai-licensing-deal/
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.