Data licensing for AI training
Who owns model outputs under a training data license
Quick answer
A training data licensor owns model outputs only if the license says so, and a buyer should make sure it says the opposite. Copyright law on AI outputs is unsettled, so ownership is allocated by contract, and that allocation binds the parties even where outputs may not be copyrightable [1]. Negotiate an express statement that the licensor claims no rights in outputs. Accept only narrow exceptions for verbatim reproduction of licensed records and disclosure of the licensor's confidential information.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Three artifacts the output clause must keep apart
An output clause works only if the license defines weights, outputs and output-derived datasets as separate things. Some licensor drafts fold all three into one word, "Derivatives." That can turn an ordinary generation product into a licensed derivative, along with whatever royalty, audit or deletion obligation attaches to derivatives.
- Model weights and checkpoints. These are the trained parameters. Whether you can keep, fine-tune or ship them belongs in the derivative and successor model rights and post-termination retention clauses, not in the output clause.
- Outputs. These are the tokens, images, embeddings, classifications or actions the deployed model produces in response to inputs. They are the product your customers buy.
- Output-derived datasets. These are synthetic corpora, distillation sets or preference pairs built by sampling the model. They sit between the other two, because they can carry licensed patterns into a new model without carrying any licensed record.
Model-provider licenses show why the third category matters. Google's Gemma Terms of Use define "Model Derivatives" to include any model trained to behave like Gemma by transferring patterns from its outputs, explicitly covering distillation and training on synthetic data [5]. A data licensor can copy that drafting pattern. Treat any definition that reaches "models trained on Outputs" as a restriction on synthetic data, and negotiate it under the synthetic data rights terms rather than letting it ride in on the output clause.
Why ownership is settled by contract, not by copyright
Contract governs output ownership because copyright law gives no reliable default in either direction. Practitioner guidance notes that output ownership provisions remain enforceable between the parties even where the output itself may not qualify for copyright protection [1], and that the US Copyright Office's human-authorship position limits protection for purely machine-generated material [3]. Ownership questions in AI data deals remain fact-specific [2].
The training side is no clearer as of October 2026. The Copyright Office's Part 3 report on generative AI training is still a pre-publication version from May 2025 [9]. That is why practitioners recommend a training data license carry an explicit grant, a warranty package, an indemnity and an allocation of output ownership, so the deal holds whatever courts decide on fair use [4]. Whether your outputs are themselves protectable is a separate compliance question, covered under copyright guidance in the AI data hub.
What a buyer-favorable output clause contains
A buyer-favorable output clause has four parts: a disclaimer of licensor rights, an ownership statement for the licensee, a narrow reproduction exception, and an explicit carve-out from royalty and audit. Each part closes a specific path by which a licensor could later reach into revenue from your product.
- No licensor claim. The licensor acknowledges that it acquires no right, title, interest, lien or royalty in outputs, and that outputs are not "Licensed Data," "Derivatives" or "Confidential Information" merely because the model was trained on the data.
- Licensee ownership, as between the parties. Outputs belong to the licensee or its customers as between the parties, to the extent any rights exist. The "to the extent" phrasing avoids warranting copyrightability.
- Narrow verbatim-reproduction exception. The only retained claim covers outputs that reproduce a licensed record, or a substantial portion of one, verbatim or near-verbatim. That exception should trigger a remediation duty, not an ownership transfer.
- No pass-through economics. Per-output fees, revenue share on generated content and audits of output logs are excluded unless the pricing structure says otherwise in plain words.
Watch for three drafting failure modes. A definition of "Licensed Data" that includes "any data derived from" the records captures outputs by implication. A confidentiality clause covering "information derived from Confidential Information" captures model behavior. A field-of-use restriction that limits "use of the Licensed Data" to internal purposes can be read to bar commercial deployment of the model.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Clause element | Buyer position | Common licensor ask | Reasonable fallback |
|---|---|---|---|
| Ownership of outputs | Licensor claims no rights; licensee owns as between parties | Licensor owns outputs "derived from" Licensed Data | Licensor claims no rights except reproduced records |
| Definition of Derivatives | Weights only; outputs and output datasets excluded | Includes outputs and models trained on outputs | Outputs excluded; output-derived training sets handled in synthetic-data clause |
| Reproduction exception | Verbatim or near-verbatim copy of a licensed record | Any output "resembling" Licensed Data | Substantial similarity to an identifiable record, assessed on a stated method |
| Remedy for reproduction | Suppress, filter and notify | Licensor ownership plus damages | Remediation duty plus cap tied to the IP indemnity |
| Output filtering | Commercially reasonable controls, described in an annex | Licensor-approved filters and audit rights | Annual written summary of controls; no log access |
| Confidential facts in outputs | Covered only where an output discloses non-public licensor facts | All outputs are Confidential Information | Disclosure test plus prompt-level suppression |
| Survival | Output rights survive termination | Outputs must be deleted at termination | Outputs generated during the term survive; see deletion and return |
Sample language for element 1, for counsel to adapt: "As between the parties, Licensor acquires no right, title or interest in any Output, and no Output constitutes Licensed Data, a Derivative or Licensor Confidential Information solely because the Model was trained on Licensed Data, except as set out in Section [x] (Reproduced Records)."
Outputs that resemble licensed data
Outputs that merely resemble licensed data are usually the lower-risk case commercially, while outputs that reproduce licensed records are the real risk and should be handled by an engineering control, not an ownership transfer. Resemblance is the point of training: a model fine-tuned on support tickets should write like a support agent. Reproduction is different, because a copied record carries the licensor's copyright, any third-party rights and any personal data in it.
Memorization is measurable, not hypothetical. Researchers extracted verbatim training examples, including personal information, from GPT-2, a publicly released language model [6], and showed that image diffusion models can regenerate individual training images [7]. Deduplicating the training set reduced how often models emitted memorized text [8]. That gives both sides an objective basis for the exception: tie it to records an extraction test can actually find, and specify the test.
Licensors may ask for regurgitation controls, which is reasonable if they are scoped. Acceptable obligations include deduplication before training, n-gram or embedding-similarity filters against the licensed corpus at inference, canary records seeded by the licensor to test extraction, and a takedown path that suppresses a flagged record within the product. Resist obligations to share raw prompt and output logs. They expose your customers' data and can conflict with your own privacy commitments.
Confidential facts that surface in outputs
An output can leak licensed information without copying a single record, so the confidentiality clause needs its own output test. A model trained on a licensor's sales histories might state a customer's renewal price, or a model trained on engineering tickets might reveal an unreleased product name. Neither is verbatim reproduction, yet both disclose the licensor's non-public facts.
Draft the confidentiality carve-out around disclosure, not derivation. An output breaches confidentiality only if it discloses specific non-public information identifiable to the licensor or its counterparties, and the licensee's duty is to suppress such outputs once notified. Pair it with pre-delivery redaction of identifiers in the data itself, and ask suppliers what de-identification method was applied and how a sample was checked; the de-identification evidence package lists the documents. Where the training data includes customer conversations, the same analysis applies to agent deployment logs.
Regulatory duties that touch outputs
Regulation adds output-side duties that sit with the model provider regardless of what the data license says. Under Article 53(1)(c) of the EU AI Act, providers of general-purpose AI models must keep a policy to comply with Union copyright law; those obligations have applied since 2 August 2025 [10]. The Code of Practice's Copyright chapter, a voluntary route to showing compliance, includes measures aimed at mitigating copyright-infringing outputs [11].
The practical consequence for a data license is alignment. If you are a GPAI provider or build on one, your output filtering controls already exist for regulatory reasons, so describe those same controls in the license annex rather than accepting a second, licensor-specific regime. Your downstream customer terms should also not promise more about outputs than your data licenses support.
Outputs used to build new datasets
Output-derived datasets need an express right, because many licensors treat them as a way around scope limits. Distilling a model trained on licensed data into a smaller model, or sampling it to build a fine-tuning set, can reproduce the value of the licensed corpus without the license. Diligence on synthetic data increasingly asks about its ownership and provenance [12].
Decide which side of the line you need before negotiating. If your product only serves outputs to end users, accept a restriction on using outputs to train competing general-purpose models in exchange for a clean ownership statement. If your roadmap includes distillation, self-training or a synthetic data product, negotiate that right explicitly. A fine-tuning-only license is usually too narrow for it, and the negotiation checklist has fallback positions. If you are sourcing new data for these uses, describe the data and intended uses to SourceX, since allowed uses are agreed in the license and nothing is contracted until the supplier agrees.
Sourcing training data with output rights in view
SourceX sources operational datasets from US companies on request and manages the commercial process, including the licensing agreement that sets pricing and allowed uses. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, allowed uses, term and delivery, and nothing is contracted until the supplying company agrees. For clause-by-clause background, see AI data license terms explained and the licensing hub, then describe the data you need to SourceX.
Frequently asked questions
Does the data licensor own model outputs by default?
No default rule gives a training data licensor ownership of outputs, but a broad definition of "Licensed Data" or "Derivatives" can give it a contractual claim. Read the definitions section before the ownership section, because that is where the capture happens.
Can we give customers an output ownership statement if our data licenses are silent?
Silence is weaker than an express disclaimer, because a licensor can later argue that outputs fall within a derivatives or confidentiality definition. Map every active data license against the output terms you give customers, and close gaps at renewal. The general ownership rules for the underlying data are covered in do I keep ownership of licensed data.
Should output terms differ for pre-training and fine-tuning data?
Yes, in practice. Fine-tuning data has a stronger influence on output style and content per record, so licensors of small fine-tuning sets may push harder on reproduction and confidentiality terms than licensors of pre-training data.
Sources
- Lathrop GPM, "Navigating AI Ownership in Commercial and IP License Agreements: Key Considerations for Tech Providers and Customers". https://www.lathropgpm.com/insights/navigating-ai-ownership-in-commercial-and-ip-license-agreements-key-considerations-for-tech-providers-and-customers/
- Reed Smith, "Entertainment and Media Guide to AI, Three Years On: The Thorny Issue of Data Ownership". https://www.reedsmith.com/articles/entertainment-media-guide-to-ai-three-years-on/the-thorny-issue-of-data-ownership/
- Vorys, "Key Contract Terms and Conditions for AI Products and Services, Part 1: Data Ownership and Licensing" (2024). https://sitepilot10.firmseek.com/client/vorys/www/printpilot-publication-key-contract-terms-and-conditions-for-ai-products-and-services-part-1-data-ownership-and-licensing.pdf?1714235896
- terms.law, "AI and Data Licensing: A Usable Agreement (Memo)". https://terms.law/insights/ai-training-data-licensing-usable-agreement.html
- Google (archived copy hosted by Wind River), "Gemma Terms of Use (version dated 24 March 2025)" (2025). https://open.windriver.com/info/uni-license-list/licenses/gemma-tou-2025-03-24.html
- Carlini et al., arXiv, "Extracting Training Data from Large Language Models" (2020). https://arxiv.org/pdf/2012.07805
- Carlini et al., arXiv, "Extracting Training Data from Diffusion Models" (2023). https://arxiv.org/abs/2301.13188
- Lee et al., arXiv, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- Mayer Brown, "Synthetic Data as a Deal Asset: Ownership, Provenance and Diligence Considerations in AI Acquisitions" (2026). https://www.mayerbrown.com/en/insights/publications/2026/07/synthetic-data-as-a-deal-asset-ownership-provenance-and-diligence-considerations-in-ai-acquisitions
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.