Data licensing for AI training
AI data license negotiation checklist for buyers, with fallback positions
Quick answer
A data license negotiation checklist for AI training should rank terms by what you cannot fix once a model is trained. Settle five model-defining terms first: the training rights grant, the definition of licensed models, model survival after termination, title and collection warranties, and an IP and privacy indemnity outside the general cap. Then negotiate deletion, takedowns, audit, disclosure carve-outs, acceptance, access and assignment. Leave exclusivity and price for last. Each term below has a first ask, a fallback and a walk-away signal.
By SourceX Editorial · Updated
Why model rights come first and price comes last
Model-defining terms come first because their failures cannot be reversed after training: a weak grant or deletion clause can force you to retire a model that products depend on, while a weak price can be fixed at renewal. Rank the remaining terms by reversibility, size of exposure and verifiability.
The legal backdrop explains why model rights are not worth trading for a discount. The US Copyright Office's Part 3 report, still a May 2025 pre-publication version as of October 2026, concludes that many acts in AI training, including copying and organizing works into datasets, may be prima facie infringing unless an exception such as fair use applies [1]. Courts have not settled the question. On 29 September 2026 the Third Circuit held that ROSS's use of Westlaw headnotes to train a non-generative legal-research tool was not fair use [2], while a June 2025 district court ruling in Kadrey v. Meta found training on books fair use on that record, with other claims continuing [3].
How the licensor obtained its content is a separate exposure. In Bartz v. Anthropic, the class settlement received final approval in July 2026 [4]. Regulators add model-level risk: the FTC's 2021 order against Everalbum required deletion of the models and algorithms built from users' photos and videos [5].
Decide first whether you are licensing or relying on fair use. Once you license, the grant has to name every act your pipeline performs.
The ranked checklist: first ask, fallback, walk-away
Rows 1 to 5 decide whether your model is usable, rows 6 to 12 set what operating the license costs, and rows 13 and 14 are commercial. The fallback is the most counsel should concede without escalating; a walk-away signal is wording that stops signature until the deal owner re-approves the risk.
| # | Term | First ask | Fallback | Walk-away signal |
|---|---|---|---|---|
| 1 | Training rights grant | Named acts: copy, store, transform, train, fine-tune and evaluate, then deploy and commercialize trained models and outputs | The same acts within a defined field of use | "Use" or "internal research," with no training or deployment verb |
| 2 | Which models are licensed | Any model whose parameters were updated with the data, plus fine-tuned, distilled, quantized and merged variants | A named model family and successors released within a set period | One named checkpoint; each new version needs a new license |
| 3 | Models after the license ends | Perpetual, irrevocable rights in models trained before the cut-off date | Survival plus no further training; run-off only for your uncured breach | Weights fall under "derivatives" in the deletion duty, or "unlearning" on request |
| 4 | Title and collection warranties | Unqualified warranties of rights, lawful acquisition, and notices and consents that cover disclosure for training | A knowledge qualifier limited to scheduled third-party content, plus copies of upstream licenses | "As is," with no warranty of title |
| 5 | IP and privacy indemnity | Defense and indemnity for third-party IP and privacy claims, including retraining costs, outside the cap | A super-cap set as a multiple of total fees (see liability caps and super-caps) | Indemnity inside a 12-months-of-fees cap, or excluding training claims |
| 6 | Deletion and return | Data copies only (raw, staged, tokenized, embedded), backups on rotation, written certificate | A longer deletion window and an agreed certificate form | Deletion reaches weights or the records behind legal disclosures |
| 7 | Record takedowns during the term | Removal from data copies and future runs, with replacement records or credit | A volume cap and batched takedowns at fixed intervals | Every takedown obliges you to retrain |
| 8 | Audit and usage reporting | An annual officer's certificate of compliance | One independent audit a year, on notice, limited to logged training inputs | Inspection of weights, source code or unrelated datasets |
| 9 | Confidentiality and disclosures | Carve-out for legally required disclosures and training-data documentation | Category-level descriptions without the licensor's name, where the law allows | An absolute ban on identifying the source |
| 10 | Acceptance and remedies | Written criteria, an inspection window, re-delivery or credit for defective records | Re-delivery first, then a refund for the defective portion | Deemed acceptance on delivery or on silence |
| 11 | Affiliates, contractors and cloud | Affiliates, contractors and cloud processors may process the data under confidentiality | Named processor categories, with notice | No third-party processing, which rules out cloud training |
| 12 | Assignment and change of control | Assignment to affiliates or an acquirer without consent; the licensor's change of control leaves the license intact | Consent not to be unreasonably withheld | The licensor may terminate if you are acquired |
| 13 | Exclusivity | If needed: exclusivity limited by field and time, or a right of first refusal | Notice when the data is licensed to others in your field | An exclusivity premium with no field definition or remedy |
| 14 | Price structure | A fixed fee for a defined record set and delivery schedule; an MFN clause for repeat purchases | Per-record pricing with a cap, or a minimum guarantee credited against fees | A revenue share on model revenue with no attribution method |
Redlines licensors send, and how to answer them
Licensor first drafts often narrow the grant, widen deletion, qualify warranties and pull the indemnity inside the general cap; answer each with precise defined terms rather than broader promises. The markup below answers four of them; Trained Models and Retained Records are defined as in the survival language for trained models.
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
Licensor draft with buyer markup: [- deleted -] [+ inserted +]
2.1 Grant. Licensor grants Licensee a non-exclusive, worldwide license
to [- use the Licensed Data for internal research purposes -]
[+ copy, store, process, annotate and transform the Licensed Data;
train, fine-tune, evaluate and test Models on it; and use, host,
deploy and commercialize Trained Models and their Outputs +]
during the Term [+ and, for Trained Models and their Outputs,
perpetually and irrevocably +].
8.2 Title. [- To Licensor's knowledge, -] Licensor has all rights needed
to grant Section 2.1 [+ and collected the Licensed Data under notices
and consents that permit its disclosure to Licensee for the Permitted
Purpose. No Licensed Data was obtained from unauthorized copies.
Knowledge-qualified warranties apply only to the Third-Party Content
listed in Schedule C +].
11.3 Deletion. Within [X] days after the Term ends, Licensee will delete
all copies of the Licensed Data [- and all derivatives -] [+ ,
excluding Trained Models and Retained Records, +] and certify
deletion in writing.
12.4 Cap. Each party's total liability is limited to the fees paid in
the 12 months before the claim[- . -][+ , except that Licensor's
obligations under Section 10 (IP and Privacy Indemnity) are subject
to a separate cap of [X] times total fees and include reasonable
costs to retrain or replace affected Trained Models. +]
Other common moves:
- "Unlearn on request." No agreed test shows that a dataset's influence has left the weights. Offer no further training after the cut-off, removal at the next scheduled model version, or priced removal: one practitioner case study describes a tiered removal-on-request mechanism whose cost depends on the method [6].
- Audit "of all systems and models." Offer the officer's certificate, then an independent audit limited to logged training inputs, the scope used in the same case study [6].
- "No output similar to Licensed Data." Accept a measurable limit on verbatim reproduction instead. In HarperCollins's 2024 opt-in program for selected nonfiction titles, the AI company agreed to limit verbatim reproduction, according to the Authors Guild [7].
- "Licensee relies on the dataset card." Hosting-site metadata is not a warranty: the Data Provenance Initiative reported license omission above 70% and license error rates above 50% on popular dataset hosting sites [8]. Ask for upstream licenses and a disclosure schedule.
- Warranties on ownership but not collection. FTC staff warned in February 2024 that adopting more permissive data practices, such as AI training, through a surreptitious, retroactive change to terms or a privacy policy may be unfair or deceptive [9]. Markup 8.2 therefore reaches the notices in force at collection, which consent and notice records let you check.
- Licensor termination for convenience. Strike it, or tie it to model survival and a refund of prepaid fees.
Statutory terms neither side can negotiate away
Some clauses come from statute, so negotiate their scope and flow-down, not their existence. As of October 2026, these US and EU rules may apply; check each against your dataset and model.
Terms the licensor must impose on you, or that bind you as the recipient:
- California deidentified data. Under Civil Code 1798.140(m), information counts as deidentified only if the business, among other conditions, commits not to reidentify it and contractually obligates recipients to comply with the subdivision [10]. Accept the no-reidentification covenant; negotiate how it flows down to your contractors.
- HIPAA limited data sets. A limited data set under 45 CFR 164.514(e) is still protected health information, may be used only for research, public health or health care operations, and requires a data use agreement that restricts further disclosure and prohibits identifying or contacting individuals [11]. Confirm your training purpose fits first. For health records, SourceX requires HIPAA de-identification (Safe Harbor or Expert Determination) before anything is considered for a license.
- Financial data under Regulation P. A recipient of nonpublic personal information from a nonaffiliated financial institution under an exception may use it only for the purpose for which it was received, whether or not the recipient is a financial institution [12]. No permitted-use clause can license more.
Duties that fall on you, which row 9 must preserve:
- EU AI Act Article 53(1)(d). Providers of general-purpose AI models must publish a sufficiently detailed summary of training content using the AI Office template [13].
- California AB 2013. Developers of generative AI systems made available to Californians must post training-data documentation, including dataset sources or owners and whether datasets include copyrighted or licensed material, first by 1 January 2026 and again before each later release or substantial modification [14].
- Colorado SB26-189. From 1 January 2027, developers of automated decision-making technology that materially influences consequential decisions must give deployers documentation that includes training data categories [15].
Sequence, owners and sign-off gates
Give each stage a named owner and a gate that must close before the next stage opens. A realistic licensing timeline covers durations, and internal approvals for a data purchase covers each reviewer's checklist.
| Stage | Owner | Output | Gate before the next stage |
|---|---|---|---|
| 1. Define uses | ML lead with counsel | Training stages (pre-training, fine-tuning, evaluation, retrieval), model families, deployment channels | Rows 1 to 3 asks match the stated uses |
| 2. Set positions | Counsel, privacy, security | Walk-away list and fallback authority for each row | Deal owner approves what counsel may concede alone |
| 3. Term sheet | Counsel and licensor | Rows 1 to 5 plus price structure in plain language | No drafting until rows 1 to 5 are agreed in principle |
| 4. Diligence | Data lead, privacy | Sample review, questionnaire answers, upstream licenses, consent terms | Findings populate the warranty disclosure schedule |
| 5. Redlines | Counsel | Marked drafts and a concession log | Any walk-away signal escalates to the deal owner |
| 6. Approval | Legal, privacy, security, finance | Signed approval record | Signature |
| 7. Implement | ML data engineering | Pipeline controls for permitted uses, deletion and takedowns | First delivery accepted against written criteria |
Stage 3 uses the AI data license term sheet, and stage 4 the data provider due diligence questionnaire. The FISD Alternative Data Council's industry questionnaire, for example, asks for the consent terms agreed with the individuals concerned and, for data bought from others, the contract terms that allow resale [16]. For counsel's own triage, see reviewing an AI data license as in-house counsel.
If SourceX sources the dataset, rights review checks that the business owns or may share the records and that required consents are in place, and diligence materials on source, rights, preparation and allowed use are prepared for your review. State your row 1 to 3 positions when you describe the data you need to SourceX; nothing is contracted until a supplier agrees.
Trades that keep rows 1 to 5 intact
The cheapest concessions give the licensor something it can verify in exchange for rights that protect your model:
- A finite term for holding the data, in exchange for perpetual rights in models trained during it.
- A measurable verbatim-reproduction limit, in exchange for a broad model definition.
- Certification plus a log-limited audit, in exchange for no model inspection.
- Non-exclusivity or a narrower exclusivity field, in exchange for price.
- Batched takedowns with replacement records, in exchange for no retraining duty.
Mistakes that cost buyers the model:
- Agreeing price before scope. A number fixed early becomes leverage against rows 1 to 5, because price follows scope, volume, rights and exclusivity. So does a commercial lead trading a fallback for a discount without counsel.
- Training on an evaluation license. An evaluation grant that does not name training does not license it; see evaluation-only data license terms.
- Signing the licensor's analytics paper. Agreements written for dashboards or market data feeds often license only "internal business purposes," which does not name training.
SourceX's supplier-side guide to preparing a commercial brief asks licensors to write down their positions on allowed uses, exclusivity, term, territory, payment structure and deletion before negotiating. The deal terms checklist and common data licensing mistakes show the licensor's side; AI data license terms explained and the AI training data licensing hub cover the rest.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Turn your ranked terms into a data request
Describe the data you need and the uses that must be licensed, starting with your rows 1 to 3 positions. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license in which pricing and allowed uses are agreed. Submit your licensing requirements.
Sources
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- terms.law, "AI and data licensing (case study: archive imagery licensed for AI training)". https://terms.law/case-studies/ai-data-licensing-archive-imagery.html
- Authors Guild, "HarperCollins AI Licensing Deal" (2024). https://authorsguild.org/news/harpercollins-ai-licensing-deal/
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version 2024). https://arxiv.org/abs/2310.16787
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- eCFR (HHS), "45 CFR 164.514(e) - Limited data set and data use agreements". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
- FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ), with generative AI questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.