Provenance, rights and permitted use
Training on Other Models' Outputs: Checking Provider Output Terms Before You Use the Data
Quick answer
Often yes, but not freely. Major API providers let customers own their outputs while barring use of those outputs to develop competing models [1][2], and open-weight licenses attach their own conditions, from Gemma's "Model Derivatives" definition to Llama's naming rule [3][4]. Because outputs probably carry little copyright, these limits mostly rest on contract, so who accepted which terms matters [5]. Before training, identify which records are model-generated, which model produced them, under which terms version, and whether your model would "compete."
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why model-generated records hide inside acquired datasets
Model-generated text now appears in datasets that were never sold as synthetic. Support exports contain agent replies drafted by an assistant and sent with light edits, sales histories hold AI-written follow-up emails, and engineering records include commit messages and code produced by coding assistants. Labeling vendors may also have used a model to draft annotations or rationales that humans only approved.
Distillation corpora are the obvious case: teacher completions collected to train a student, as covered in building distillation datasets. Training one model on another's outputs is an established technique [6], which is exactly why providers write terms about it. The provenance question is the same in both cases, but in operational data nobody flagged it.
Two separate questions follow. First, which records are model output and from which system. Second, what the terms of that system say about training, and whether they bind you. For detection methods, see detecting model-generated content in purchased "human" data; this page covers the terms check.
What the main provider terms actually restrict
Provider terms fall into three patterns: competing-model bans, derivative-model definitions, and attribution conditions. Read each family separately, because the trigger and the consequence differ.
Competing-model bans (closed API providers). OpenAI's terms say users own their Output, yet list "use Output to develop models that compete with OpenAI" among prohibited activities [2]. Anthropic's help center states that outputs may not be used to train competing AI models and lists non-competing uses that remain allowed [1]. The restriction is on the user who accepted the terms and on the purpose, not a label that travels with the text.
Derivative-model definitions (Gemma). Gemma's terms define "Model Derivatives" to include a model trained to behave like Gemma by transferring patterns from its weights, parameters, operations or Outputs, which reaches distillation and training on synthetic Outputs [3]. A student trained on Gemma outputs can therefore inherit Gemma's use restrictions and distribution conditions, rather than escaping them.
Attribution and naming conditions (Llama 3.1). The Llama 3.1 Community License permits using Llama outputs to create, train or improve another model, but if that model is distributed or made available, its name must begin with "Llama"; distributing Llama Materials or derivative works also triggers "Built with Llama" attribution and a license copy [4]. Here the output is usable for training other models; the obligations are naming, attribution and license pass-through, alongside Llama's acceptable use policy.
| Term family | Typical trigger | What it restricts | What it means for an acquired dataset |
|---|---|---|---|
| Competing-model ban (API terms) | Using Output to develop a model that competes with the provider | Purpose of use by the account holder | Ask who generated the records and whether the buyer's model competes |
| Derivative-model definition (Gemma-style) | Training a model to behave like the source model, including via Outputs | Student inherits source-model terms | Treat the trained model as a derivative and carry the use policy forward |
| Attribution or naming condition (Llama 3.1) | Distributing a model built with outputs | Name prefix, attribution notice, license copy | Plan naming and notices before release, not after |
| Non-commercial data license on outputs | Dataset publisher adds its own license | Commercial use of the dataset as published | A separate license layer on top of the model terms |
Whether output restrictions bind a dataset buyer
The answer is unsettled: the account holder who accepted the terms is clearly bound, while enforceability against a downstream buyer who never accepted them is contested. Commentary on Lemley and Henderson's analysis argues that model output likely lacks copyright protection, so providers depend on contract, and contract reaches only parties to it [5]. That argument is not a ruling, and courts have not settled it as of October 2026.
Several facts can put the buyer back inside the contract. If your organization holds an API account with the same provider, its own terms already bind you on purpose of use. If the supplier warrants compliance with provider terms and the license passes restrictions through, you may have agreed contractually anyway. If the supplier generated the data in breach, a provider dispute with the supplier can disrupt supply and draw your training run into discovery.
The practical stance is to treat enforceability as a risk weight, not a clearance. A competing-model clause you are not party to is lower risk than one you are, but it is not zero, and the reputational cost of shipping a model distilled in breach of a rival's terms can exceed the legal exposure.
Deciding whether your model "competes"
"Competes" is undefined in most terms, so assess it against your model and product, not the provider's whole business. A narrow classifier for invoice routing is a weaker competitor than a general chat model released as an API. The same dataset can be acceptable for one buyer and problematic for another.
Score each use against these factors:
- Model generality. A general-purpose instruction-following or chat model sits closest to a frontier provider's product.
- Distribution. Internal tools differ from models offered to third parties as an API, open weights or a product feature.
- Training role. Using outputs as SFT targets or distillation signal differs from using them as negatives, eval references or retrieval context; check whether the terms reach evaluation at all [1].
- Volume and centrality. A few thousand stray AI-drafted replies inside a large support corpus differ from a corpus that is mostly one teacher's completions.
- Behavior transfer. If the goal is to imitate the source model's style or reasoning traces, Gemma-style derivative language is most likely to apply [3].
Record the conclusion and the reasoning per dataset in your AI training data use register, so the call can be revisited when terms change.
The provenance record to request for model-generated content
The record that makes this check possible names the generating model, the account context and the terms version for each batch of generated content. Providers publish dated revisions of their terms, as OpenAI does [2], so a generation date lets counsel identify the version in force. Without it, you are reading today's terms against last year's data.
Ask the supplier for a generation manifest at batch level, plus a field on each record flagging whether model assistance was involved. The structure below extends the fields in provenance records for synthetic training data.
Illustrative example: invented to show structure; it does not describe an available dataset.
generation_manifest:
batch_id: gen-batch-0042
record_selector: "support_tickets where reply_origin in ['ai_drafted','ai_drafted_edited']"
generator:
provider: "ExampleAI"
model_id: "example-model-2025-03"
access_route: "API, supplier's own enterprise account"
account_holder: "supplier"
terms:
terms_name: "ExampleAI Business Terms"
terms_version_date: "2025-01-15"
terms_url_snapshot: "archived copy attached"
output_clause_summary: "Customer owns Output; no use of Output to develop competing models"
generation_window: { start: "2025-02-01", end: "2025-11-30" }
human_involvement: "agents may edit drafts before sending; edit flag per record"
record_fields:
reply_origin: [human, ai_drafted, ai_drafted_edited]
edit_distance_bucket: [none, light, heavy]
supplier_statement: "Generated for the supplier's own operations, not for resale as training data"
The reply_origin flag lets you filter or down-weight model-assisted records if counsel decides a clause applies. The account_holder field tells you whose contract governs, and supplier_statement captures purpose, which matters when terms distinguish operational use from building training sets.
A pre-training check for datasets with model-generated content
Run this check before data enters a training mix, alongside your wider chain-of-title documents.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Step | Question | Evidence to collect | Red flag |
|---|---|---|---|
| 1 | Does the dataset contain model output at all? | Supplier disclosure, origin flags, detector sample | Supplier says "100% human" but detector and metadata disagree |
| 2 | Which models generated it? | Model IDs, providers, open-weight license names | "Various models" with no list |
| 3 | Under which terms version? | Generation dates, archived terms, account type | No dates; terms read only in current form |
| 4 | Who accepted those terms? | Account holder, enterprise agreement reference | Generated through personal or shared accounts |
| 5 | Does our use compete or create a derivative? | Use-case memo scored on the factors above | General chat model trained mainly on one rival's outputs |
| 6 | Are there downstream conditions? | Naming, attribution, license copy obligations [4] | Release plan ignores required notices |
| 7 | Does the supplier license address it? | Representation on generated content, pass-through terms | License silent on model-generated records |
| 8 | Can we disclose it if required? | Synthetic-data flag for training documentation | No way to answer whether synthetic data was used |
Step 8 is not academic. California AB 2013 requires developers of generative AI systems offered to Californians to post documentation about their training data, including whether synthetic data generation was used, with postings due by 1 January 2026 [7]. A buyer that cannot tell model-generated from human records cannot answer that accurately. The data rights attestation template gives wording to request from suppliers, and SourceX for AI data buyers can take a request that specifies how model-generated records should be flagged.
How this differs from open-dataset and gated-license checks
Output terms sit on a different layer from dataset licenses, and a dataset can clear one while failing the other. A publisher may release model-generated data under CC BY-NC or a custom license; that license governs the dataset as published, while the generating model's terms governed the publisher's own conduct. Check both, as described in gated and custom-licensed datasets on model hubs.
For operational data from a company, there is usually no dataset license yet; the license you negotiate becomes that layer. That is the moment to require origin flags and generation manifests, because asking after delivery rarely produces them. For a broader comparison of licensed, synthetic and scraped sources, see licensed vs synthetic vs scraped AI training data, and for the full rights framework, the provenance buyer's guide.
Sourcing operational data with clear provenance
SourceX sources operational datasets, such as support and sales histories and engineering records, from US companies on request; categories are not inventory, and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, comes with diligence materials on source, preparation and allowed use, and is delivered under a license defining records, uses, term and delivery. Describe the data you need, including how you will treat model-generated records, at SourceX for AI data buyers.
Frequently asked questions
If I own my outputs, can I train anything on them?
Not necessarily. OpenAI's terms assign Output ownership to the user and still prohibit using Output to develop competing models [2]. Ownership and permitted use are separate questions in these contracts.
Do open-weight model outputs carry fewer restrictions?
They carry different ones. Llama 3.1 allows training other models on outputs subject to naming and attribution conditions [4], while Gemma treats models trained on its Outputs as Model Derivatives subject to its terms [3].
Can I use restricted outputs for evaluation only?
Possibly, depending on the provider. Some providers list permitted non-competing uses [1], but eval sets can also shape model development, so document the role and read the specific terms version.
Sources
- Anthropic (Claude Help Center), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- OpenAI, "Terms of Use (rest of world), revision dated 2024-12-11" (2024). https://openai.com/policies/row-terms-of-use/revisions/2024-12-11/
- Google, "Gemma Terms of Use (April 1, 2024 archive)" (2024). https://ai.google.dev/gemma/terms-archive/terms_04_01_24
- Meta (Hugging Face model repository), "Llama 3.1 70B Instruct model card and Llama 3.1 Community License" (2024). https://huggingface.co/meta-llama/Llama-3.1-70B-Instruct
- SpicyIP, "Discussing Lemley and Henderson's 'The Mirage of Artificial Intelligence Terms of Use Restrictions'" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
- DeepLearning.AI (The Batch), "When One Machine Learning Model Learns From Another". https://www.deeplearning.ai/the-batch/when-one-machine-learning-model-learns-from-another
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.