Data licensing for AI training
Can you generate synthetic data from licensed data? Rights to derived datasets
Quick answer
Only if the license says so. A standard AI training grant covers training a model on the licensed records; it does not automatically cover feeding those records to a generator, keeping the synthetic output after term, or sharing it. Buyers who plan synthetic augmentation or distillation need a defined "Derived Data" category with explicit generation rights, survival and transfer terms, and a privacy test for outputs. Those terms belong in the license itself. They also need to check the generator model's own output terms.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why a training grant does not cover synthetic generation
A training grant authorizes one use, and synthetic generation is a second use that produces a new asset. Most grants are drafted around "train, fine-tune and evaluate machine learning models" (see writing the AI training rights grant). Prompting a model with licensed support tickets to produce 200,000 paraphrased tickets is arguably "use for training," but the output is a dataset, not a model. A licensor can reasonably say it is a copy or adaptation of the licensed data.
Law-firm analysis treats synthetic data as a separate deal asset whose ownership and provenance must be traceable to its inputs [1]. In practice, rights can attach to the generator as well as to the output. Some licenses grant full rights to generated data, while others restrict redistribution or require attribution [2]. Without explicit language, you inherit the most restrictive reading at the worst moment, often during an acquisition or a model release review.
Three ways licensed records become synthetic data
The rights question changes with the generation method, so name the method in the license. Each method carries different leakage and ownership risk.
| Method | What touches licensed data | Typical output | Main rights risk |
|---|---|---|---|
| Seeded prompting (few-shot, Self-Instruct style) | Licensed records placed in prompts as examples | Instruction pairs, paraphrased tickets, dialogues | Near-copies of seed records; seeds sent to a third-party API |
| Fine-tuned generator | Generator weights trained on licensed records | Large synthetic corpora sampled from the tuned model | Generator itself is a derived model; output may memorize records |
| Distillation from a teacher trained on licensed data | Teacher model outputs, indirectly the licensed corpus | Student training sets, preference pairs | Output terms of the teacher plus the data license both apply |
| Statistical synthesis (tabular, GAN, copula, DP) | Licensed tables fit to a model | Synthetic rows with matched marginals | Re-identification and attribute disclosure from sparse rows |
For the method side, see seed data for synthetic instruction generation and the SourceX guide to combining licensed and synthetic data. For definitions, see the synthetic data glossary entry. This page covers only the contract rights.
Defining Derived Data so the clause actually works
A workable clause defines Derived Data by method and by test, not by intent. "Data derived from the Licensed Data" is too broad: it captures embeddings, labels and aggregate statistics. Negotiate four separate decisions.
- Scope. Does Derived Data include synthetic records, model-generated annotations, embeddings and summary statistics? Embeddings have their own issues; see embedding and vector index rights.
- Ownership. Licensee-owned, licensor-owned with a license back, or jointly held. Joint ownership is often the least useful because co-owners' exploitation rights differ by jurisdiction.
- Survival. Does Derived Data survive termination, or is it treated like Licensed Data and deleted? This interacts with what happens to trained models when a license ends.
- Transfer. May Derived Data be sold, shared with affiliates, given to a vendor, or used to train third-party or open-weight models? Downstream model questions are covered under derivative and successor model rights.
A common licensor compromise is a "non-reconstructive" test: Derived Data survives and is freely usable only if no record, or no substantial portion of one, can be recovered from it. Buyers should insist the test is measurable, for example exact and near-duplicate matching against the seed set at an agreed similarity threshold, rather than a judgment the licensor makes later.
Illustrative example: invented to show structure; it does not describe an available dataset.
DERIVED DATA RIDER (buyer draft)
1. "Derived Data" means synthetic records, generated instruction
or preference pairs, and machine-generated annotations created
by Licensee using Licensed Data as prompts, seeds, or training
data for a generator, excluding (a) verbatim or near-verbatim
copies of Licensed Data and (b) Licensed Data itself.
2. Permitted generation: seeded prompting; fine-tuning a generator
hosted in Licensee's environment; distillation. Sending Licensed
Data to a third-party model API requires written approval naming
the provider and its data-retention terms.
3. Non-reconstruction test: before retention or transfer, Licensee
runs exact-match and near-duplicate screening (MinHash, Jaccard
>= 0.8 on 5-gram shingles) against the seed set and removes hits.
Results are logged and available under the audit clause.
4. Ownership: Licensee owns Derived Data that passes Section 3.
5. Survival: Section 3-passing Derived Data survives termination;
failing records are deleted with the Licensed Data.
6. Transfer: internal use and affiliates permitted; sale or public
release of Derived Data requires [consent | separate fee | is
prohibited].
7. Personal data: Derived Data from records containing personal
data must meet the de-identification standard in Schedule P.
Generator-model terms are a second license
The model you use to generate synthetic data has its own terms, and they can restrict the output regardless of what the data licensor allows. Mayer Brown flags that generator terms may bar using outputs to train competing models [1]. Gemma's terms, for example, define "Model Derivatives" to include models trained to perform like Gemma by transferring patterns from its outputs, which reaches distillation and training on Gemma-generated synthetic data [3].
So a distilled dataset can be bound by two documents at once: the data license on the seeds and the generator terms on the output. Record which model, version and terms date produced each synthetic batch; the provenance records for synthetic training data page lists the fields. Detailed output-term analysis for hosted APIs belongs to the provenance and fine-tuning clusters rather than this page.
Personal data does not disappear in synthetic output
Synthetic records generated from personal data can still be personal data, so the license must say which privacy standard the output meets. Under GDPR Recital 26, information is anonymous only if a person cannot be identified by means reasonably likely to be used [4]. The EDPB's Opinion 28/2024 applies a similar analysis to AI models trained on personal data, and addresses consequences when development processing was unlawful [5].
NIST SP 800-188 describes formal privacy methods such as differential privacy alongside traditional de-identification, and cautions that traditional methods have inherent limits [6]. Where seeds are protected health information, HIPAA de-identification under 45 CFR 164.514 (Safe Harbor or Expert Determination) sets the standard, and a limited data set is disclosed only under a data use agreement that limits its uses [7]. Treat synthetic output from a limited data set as bound by that agreement unless counsel concludes otherwise. The privacy analysis is covered in depth at whether synthetic data from licensed records is still personal data.
Contract points to add:
- The de-identification standard the seeds met before delivery, and whether generation must preserve it.
- Membership-inference or nearest-neighbor distance tests on outputs before any external release.
- A prohibition on generating records that target or describe named individuals in the seed set.
- Who bears notification duties if a synthetic record is later linked to a real person.
Disclosure and provenance obligations follow the synthetic set
Regulators and acquirers increasingly ask where synthetic data came from, so keep lineage from seed to output. California AB 2013 requires developers of generative AI systems offered to Californians to post documentation about training data, and its required items include whether synthetic data generation was used; disclosures were due 1 January 2026 [8]. An acquirer's diligence team will also want synthetic, licensed and scraped data separated in records [1].
License documentation across the industry is already unreliable. The Data Provenance Initiative found license omission rates above 70% and error rates above 50% on popular dataset hosting sites [9]. A synthetic set with no recorded seed license inherits that problem. Tag every synthetic record with a batch ID linked to the seed license, generator version and run date.
Pricing derived-data rights
Derived-data rights are a separate economic term, and buyers usually see one of three structures. Which one fits depends on how much of the licensor's value the synthetic set replaces.
| Structure | How it works | Fits when | Buyer watch-out |
|---|---|---|---|
| Included | Derived Data rights bundled in the base fee | Internal SFT augmentation, no resale | Licensor may narrow scope to compensate |
| Capped volume | Rights cover up to an agreed record or token count | Distillation runs with predictable size | Define how records are counted after filtering |
| Separate derived-data fee | Additional fee for retention after term or external transfer | Synthetic set will be sold, shared or outlive the license | Avoid per-use royalties that need ongoing reporting you cannot produce |
Compare these with the base models in AI data license pricing structures, and add the derived-data line to your overall position in the licensing cluster guide. If you are still scoping seed data, you can describe the records and intended derived uses to SourceX so data and licensing permissions are assessed as part of the request.
Sourcing licensed seed data for synthetic generation
SourceX sources operational datasets from US companies on request and manages the licensing process, with each dataset rights-reviewed and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, and every release is approved by the supplying company. Describe the seed data and generation plan you need on the SourceX buyer page.
Frequently asked questions
Can I use licensed data to generate synthetic data for fine-tuning only?
Yes, if the license defines that use. A fine-tuning-only license that is silent on generation leaves the output's status open, so add a Derived Data rider before generating rather than after.
Who owns synthetic data generated from proprietary data?
Ownership follows the contract, not the generation step. Absent an ownership clause, expect the licensor to argue the output is an adaptation of its data and the generator provider to apply its output terms.
Can derived synthetic data train a third-party or open-weight model?
Only with an explicit transfer right in the data license and generator terms that allow it. Some generator terms restrict training competing models on outputs or bring models trained on outputs under the generator's own license [1][3].
Sources
- Mayer Brown, "Synthetic Data as a Deal Asset: Ownership, Provenance and Diligence Considerations in AI Acquisitions" (2026). https://www.mayerbrown.com/en/insights/publications/2026/07/synthetic-data-as-a-deal-asset-ownership-provenance-and-diligence-considerations-in-ai-acquisitions
- hoop.dev, "Licensing Models for Synthetic Data Generation". https://hoop.dev/blog/licensing-models-for-synthetic-data-generation/
- Google AI for Developers, "Gemma Terms of Use" (2024). https://ai.google.dev/gemma/terms-archive/terms_04_01_24
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- CMS (summary of European Data Protection Board opinion), "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- Electronic Code of Federal Regulations (HHS), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.