Data licensing for AI training
Derivative and successor model rights in AI data licenses
Quick answer
Derivative model rights in a data license reach only the models its definition of "Model" covers. Copyright's "derivative work" concept is an unreliable test of whether a distilled, quantized or merged model, or next year's successor, is licensed. Counsel should pick one of three scope patterns (named model, model family with successors, or any model trained on the data), mark each derivation technique in or out, settle synthetic outputs separately, and keep a lineage record that shows which models trace back to the licensed data.
By SourceX Editorial · Updated
Why copyright's "derivative work" test cannot set model scope
Copyright asks whether protected expression was copied or adapted; model scope asks which models were built from the licensed data, and the copyright answer for trained models is unsettled. As of October 2026, the U.S. Copyright Office's Part 3 report is still a May 2025 pre-publication version. It finds that compiling training datasets implicates the reproduction right and considers whether model weights can themselves be infringing copies where outputs closely resemble training inputs [1]. The builders of the GPT-NL corpus excluded NonCommercial and ShareAlike data to keep it usable for commercial training, and describe the law on whether a trained LLM is a derivative work as unclear [2].
US courts decide training disputes on their own facts. On 29 September 2026 the Third Circuit, in a precedential opinion, affirmed that ROSS's use of Westlaw headnotes to train a non-generative legal-research tool was not fair use [3]. In Kadrey v. Meta, a June 2025 district court ruling found fair use on its record for training on the plaintiffs' books; as of October 2026, other claims continue [4]. Neither outcome settles whether a distilled student derives from your corpus.
A license that uses "derivative works of the Data" as its boundary imports that uncertainty in both directions:
- Licensor reading. Every descendant model is a derivative work, so all fall under restrictions and deletion duties.
- Licensee reading. A student trained only on a teacher's generated outputs holds little or none of the licensed expression, so it escapes every restriction, though its capability came from the data.
- Output status. Lemley and Henderson argue model outputs lack the human authorship copyright requires, so limits on their use rest on contract [5].
Instead, define Model, Successor Model and Derived Outputs by lineage (weights inherited, data used, outputs used) and add "whether or not a derivative work under applicable law."
Three ways to define the licensed Model
Most model-scope clauses follow one of three patterns, and the right one depends on how long the licensed data must keep contributing to your roadmap.
| Pattern | What it covers | Typical fit | Licensor's position | Buyer's exposure |
|---|---|---|---|---|
| A. Named model | One model identified by name, base checkpoint hash and training run ID, plus listed derivatives | A single task model, often under a fine-tuning-only license | Easy to price and audit | Next base model or generation needs a new license |
| B. Family with successors | A named family plus successors that meet objective tests, usually time-boxed | Product lines with regular version releases | Wants a generation cap or renewal trigger | Disputes over what counts as a successor |
| C. Any model trained on the data | Every model trained in whole or in part, directly or indirectly, on the data or its outputs | Foundation-model programs; see the pre-training rights grant | Value spreads across unlimited models | Likely higher fee; still needs output and sublicensing terms |
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
PATTERN A
"Licensed Model" means the model identified in Schedule 2 by name, base
checkpoint (SHA-256 of the weight files) and training run ID, and any
Permitted Derivative. "Permitted Derivative" means a copy of the Licensed
Model that is quantized, pruned, converted to another file format or adapted
with low-rank adapters, where no Licensed Data is used after the Schedule 2 run.
PATTERN B
"Successor Model" means a model that (i) is initialized in whole or in part
from parameters of a Covered Model, (ii) is trained on a data mixture that
includes Licensed Data, or (iii) is released to replace a Covered Model in
the same product or endpoint, in each case whose training run starts before
[date].
PATTERN C
"Model" means any machine learning model trained in whole or in part,
directly or indirectly, on Licensed Data or Derived Outputs, including by
distillation, whether or not it is a derivative work under applicable law.
"Derived Outputs" means data generated by a Model, including synthetic
examples, preference labels and logits.
"In whole or in part" is also how the FTC describes models and algorithms it has required companies to delete when developed with unlawfully obtained data [6]. If your grant reaches less far, a challenge to the data can reach models your license does not cover. Fuller grant language is in writing the AI training rights grant.
Which descendants each pattern reaches
List your stack's derivation techniques in a schedule and mark each in or out, because a model can inherit from licensed data through weights, retraining on the data or another model's outputs.
| Descendant | Inherits through | A: named | B: family and successors | C: any model |
|---|---|---|---|---|
| Full fine-tune, SFT or preference tuning (DPO, RLHF) of a covered model | Weights | Only if listed | Yes (weights test) | Yes |
| LoRA adapters, or adapters merged into the weights | Weights | List expressly | Yes | Yes |
| Quantized (int8 or int4, e.g. exported as GGUF) or pruned copy | Transformed weights | List expressly | Yes | Yes |
| Merge of a covered model with an outside model | Part of the weights | Unclear | Unclear under a new name | Yes ("in part") |
| Next generation trained from scratch on a mixture that includes the data | Data | No | Yes (data test) | Yes |
| Next generation warm-started from a covered checkpoint, data not reused | Weights | No | Yes (weights test) | Only with "indirectly" |
| Student distilled from a covered model's logits or generated text | Outputs | No | Only with an outputs test | Only if Derived Outputs are included |
| Reward, classifier or embedding model trained on the data | Data | No | Unclear: name auxiliary models | Yes |
| Customer fine-tunes of your model through an API or open weights | Weights, by a third party | No | No | No, if the grant says "by Licensee" |
Merges and customer-built models cause the most surprises. A merge averaging a covered fine-tune with an outside model can ship under a new name, outside a name-based family. Models your customers build are not "trained by Licensee," so they need a flow-down or sublicensing clause for platform customers. Open-weight release puts every later derivative outside your control; see releasing open-weight models trained on licensed data.
Distillation and synthetic outputs need their own clause
Distillation transfers capability through outputs rather than weights, so a definition built only on weight inheritance misses it. In one published method, a differentially private teacher generates synthetic text that is then used to distill a student, which learns from those outputs rather than from the teacher's weights [7].
Model providers already draft for this. Anthropic's help center says its terms do not allow using outputs to train models that compete with its own, while non-competing uses such as sentiment analysis or content categorization tools are allowed [8]. Commentators question whether such terms bind parties who never accepted them [5].
Expect a data licensor to ask for comparable limits on outputs of models trained on its data. Settle three questions:
- Generation. May a covered model produce synthetic data, preference labels or logits for training?
- Downstream training. May those Derived Outputs train models outside the covered set, such as small on-device students?
- Survival. Do Derived Outputs, and models trained on them, survive termination, or must synthetic sets be deleted with the data copies?
A workable buyer position treats students distilled from covered models as covered, with the same rights and survival, and asks the licensor to disclaim rights in ordinary outputs except verbatim reproduction of licensed records. See synthetic data generated from licensed data and who owns model outputs under a training data license.
Successor tests: weights, data, replacement and time
A successor clause holds up when it relies on facts your training logs can prove, not on product names that marketing can change.
- Weights test. The model is initialized in whole or in part from covered parameters. Checkpoint lineage proves it.
- Data test. Licensed manifest IDs appear in the run's data mixture. Data-loader manifests prove it.
- Replacement test. The model replaces a covered model in the same product or endpoint. It catches renamed models but is harder to prove; use it as a backstop.
- Time box. Tie coverage to the date a training run starts, not the release date, so a run that begins during the term and ships later is covered. Add a run-off period for runs in progress at expiry. For termination, see what happens to trained models when a data license ends.
Rights to past training do not extend forward on their own. In Bartz v. Anthropic, the class settlement received final approval in July 2026 [9]. A license that names one model generation works the same way. Successors built by an acquirer or affiliate need assignment and change-of-control terms.
How model scope changes the price
Broader model scope moves more of the data's long-term value to the buyer, so expect licensors to price it higher or attach a renewal trigger. Agree the structure before the number:
- a per-model fee with a pre-priced option for the next generation;
- a family grant capped at an agreed number of successors, then a renewal fee;
- a time-boxed grant covering every run that starts within the term;
- an enterprise-wide flat fee.
Price distilled and quantized variants together with their teacher, or per-model counts multiply quickly. Structures are compared in AI data license pricing structures and per-model vs enterprise-wide training licenses. SourceX does not publish prices; terms depend on scope, volume, history, rights and exclusivity, and are agreed per deal in writing.
The lineage register that makes model scope provable
Whatever pattern you sign, keep a record linking every released model to its parents, derivation method and licensed datasets, because disclosure rules and auditors will ask. California's AB 2013 requires developers of generative AI systems offered to Californians to post training-data documentation before each release or substantial modification, covering dataset sources or owners, whether datasets include licensed material and whether synthetic data was used [10]. AI Act Article 53(1)(d) requires a public training-content summary for each general-purpose AI model [11].
License terms get lost downstream: one audit found license omission above 70% and license errors above 50% on popular dataset hosting sites [12]. NIST's Generative AI Profile lists value chain and component integration among generative AI risks [13]. FTC orders have also reached derived models: the 2021 Everalbum order required deletion of models and algorithms developed using users' photos and videos [14], and the 2023 proposed Rite Aid order reached data, models or algorithms derived from collected images [15]. A register shows which models fall inside a lineage and which do not.
Illustrative example: invented to show structure; it does not describe an available dataset.
model_id: support-agent-v3-int4
parents: [support-agent-v3]
derivation: quantize # fine_tune | adapter | distill_outputs | distill_logits | quantize | prune | merge | continued_pretrain | from_scratch
training_run_id: null # no new training run
licensed_lineage:
- manifest_id: LD-0142
license_id: DLA-2026-017
scope_pattern: B
reached_via: weights # weights | data | outputs
synthetic_inputs: []
release_channels: [api, on_device]
disclosures: { ca_ab2013_doc: "2026-09", eu_art53_summary: "v3" }
license_status: covered # covered | option_needed | out_of_scope
See audit and usage-reporting rights for audit clauses and the AI data provenance guide for dataset-level records.
Drafting sequence and who owns each step
Start from the model roadmap rather than the licensor's template, so the definition describes models you will actually build.
- Research lead. List descendants planned within the term: next base generation, distilled students, quantized on-device builds, merges, reward and embedding models.
- ML platform. Confirm the derivation methods in use and what logs capture: checkpoint hashes, data-loader manifests, generator IDs for synthetic sets.
- Counsel. Choose pattern A, B or C and draft Model, Successor Model and Derived Outputs.
- Counsel and research lead. Attach the technique schedule, marked in or out.
- Procurement. Price the options and record them in the AI data license term sheet.
- Privacy reviewer. If the data contains personal information, confirm whether distilled or synthetic outputs inherit de-identification and no-re-identification duties.
- ML platform. Turn on the lineage register before the first training run.
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Every dataset goes through rights review and is delivered under a license that defines which records are included, what they can be used for, how long the license runs and how delivery happens. Datasets are sourced on request, so a request does not guarantee a match; state the scope pattern you need when you describe your data needs to SourceX.
Wording that turns a derivative clause into a dispute
Most model-scope disputes trace back to a few phrases; search the draft for them before signing.
| Wording in the draft | Problem | Replace with |
|---|---|---|
| "Licensee shall not create derivative works of the Data" | Training arguably creates one, so the clause can bar the licensed use | An express training grant; restrict redistribution of the data instead |
| "the model developed under this Agreement" | Singular; the next generation falls outside | A pattern B or C definition |
| "Models and all derivatives" in the deletion clause | Weights fall within the deletion duty | Exclude Models and add a survival clause |
| "trained on the Data" alone | Merges, warm-started successors and distilled students are disputed | "in whole or in part, directly or indirectly" |
| "by Licensee" with no affiliates clause | Affiliate and customer-built models are excluded | Affiliates clause plus customer flow-down |
| Successor defined by product name | Renaming moves models in or out | Weights, data and time tests |
Fallback positions are in the AI data license negotiation checklist. For general terms, see SourceX's guide to AI data license terms and its clause-by-clause agreement guide; the AI training data licensing hub maps other clauses.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Need data licensed for future model versions?
Describe the data you need and the models it must cover, from one fine-tuned model to successor generations and distilled students. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license, delivery and future purchases. Specify your data and model-scope requirements.
Sources
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- arXiv:2604.00920, "GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training" (2026). https://arxiv.org/pdf/2604.00920
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
- SpicyIP, "Discussing Lemley and Henderson's 'The Mirage of Artificial Intelligence Terms of Use Restrictions'" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- arXiv:2403.00932, "Differentially Private Knowledge Distillation via Synthetic Text Generation" (2024). https://arxiv.org/html/2403.00932v2
- Anthropic (Claude Help Center), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- Federal Trade Commission, "Coming face to face with Rite Aid's allegedly unfair use of facial recognition technology" (2023). https://www.ftc.gov/business-guidance/blog/2023/12/coming-face-face-rite-aids-allegedly-unfair-use-facial-recognition-technology
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.