Skip to content

Data licensing for AI training

Can you generate synthetic data from licensed data? Rights to derived datasets

Quick answer

Only if the license says so. A standard AI training grant covers training a model on the licensed records; it does not automatically cover feeding those records to a generator, keeping the synthetic output after term, or sharing it. Buyers who plan synthetic augmentation or distillation need a defined "Derived Data" category with explicit generation rights, survival and transfer terms, and a privacy test for outputs. Those terms belong in the license itself. They also need to check the generator model's own output terms.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why a training grant does not cover synthetic generation

A training grant authorizes one use, and synthetic generation is a second use that produces a new asset. Most grants are drafted around "train, fine-tune and evaluate machine learning models" (see writing the AI training rights grant). Prompting a model with licensed support tickets to produce 200,000 paraphrased tickets is arguably "use for training," but the output is a dataset, not a model. A licensor can reasonably say it is a copy or adaptation of the licensed data.

Law-firm analysis treats synthetic data as a separate deal asset whose ownership and provenance must be traceable to its inputs [1]. In practice, rights can attach to the generator as well as to the output. Some licenses grant full rights to generated data, while others restrict redistribution or require attribution [2]. Without explicit language, you inherit the most restrictive reading at the worst moment, often during an acquisition or a model release review.

Three ways licensed records become synthetic data

The rights question changes with the generation method, so name the method in the license. Each method carries different leakage and ownership risk.

MethodWhat touches licensed dataTypical outputMain rights risk
Seeded prompting (few-shot, Self-Instruct style)Licensed records placed in prompts as examplesInstruction pairs, paraphrased tickets, dialoguesNear-copies of seed records; seeds sent to a third-party API
Fine-tuned generatorGenerator weights trained on licensed recordsLarge synthetic corpora sampled from the tuned modelGenerator itself is a derived model; output may memorize records
Distillation from a teacher trained on licensed dataTeacher model outputs, indirectly the licensed corpusStudent training sets, preference pairsOutput terms of the teacher plus the data license both apply
Statistical synthesis (tabular, GAN, copula, DP)Licensed tables fit to a modelSynthetic rows with matched marginalsRe-identification and attribute disclosure from sparse rows

For the method side, see seed data for synthetic instruction generation and the SourceX guide to combining licensed and synthetic data. For definitions, see the synthetic data glossary entry. This page covers only the contract rights.

Defining Derived Data so the clause actually works

A workable clause defines Derived Data by method and by test, not by intent. "Data derived from the Licensed Data" is too broad: it captures embeddings, labels and aggregate statistics. Negotiate four separate decisions.

  1. Scope. Does Derived Data include synthetic records, model-generated annotations, embeddings and summary statistics? Embeddings have their own issues; see embedding and vector index rights.
  2. Ownership. Licensee-owned, licensor-owned with a license back, or jointly held. Joint ownership is often the least useful because co-owners' exploitation rights differ by jurisdiction.
  3. Survival. Does Derived Data survive termination, or is it treated like Licensed Data and deleted? This interacts with what happens to trained models when a license ends.
  4. Transfer. May Derived Data be sold, shared with affiliates, given to a vendor, or used to train third-party or open-weight models? Downstream model questions are covered under derivative and successor model rights.

A common licensor compromise is a "non-reconstructive" test: Derived Data survives and is freely usable only if no record, or no substantial portion of one, can be recovered from it. Buyers should insist the test is measurable, for example exact and near-duplicate matching against the seed set at an agreed similarity threshold, rather than a judgment the licensor makes later.

Illustrative example: invented to show structure; it does not describe an available dataset.

DERIVED DATA RIDER (buyer draft)
1. "Derived Data" means synthetic records, generated instruction
   or preference pairs, and machine-generated annotations created
   by Licensee using Licensed Data as prompts, seeds, or training
   data for a generator, excluding (a) verbatim or near-verbatim
   copies of Licensed Data and (b) Licensed Data itself.
2. Permitted generation: seeded prompting; fine-tuning a generator
   hosted in Licensee's environment; distillation. Sending Licensed
   Data to a third-party model API requires written approval naming
   the provider and its data-retention terms.
3. Non-reconstruction test: before retention or transfer, Licensee
   runs exact-match and near-duplicate screening (MinHash, Jaccard
   >= 0.8 on 5-gram shingles) against the seed set and removes hits.
   Results are logged and available under the audit clause.
4. Ownership: Licensee owns Derived Data that passes Section 3.
5. Survival: Section 3-passing Derived Data survives termination;
   failing records are deleted with the Licensed Data.
6. Transfer: internal use and affiliates permitted; sale or public
   release of Derived Data requires [consent | separate fee | is
   prohibited].
7. Personal data: Derived Data from records containing personal
   data must meet the de-identification standard in Schedule P.

Generator-model terms are a second license

The model you use to generate synthetic data has its own terms, and they can restrict the output regardless of what the data licensor allows. Mayer Brown flags that generator terms may bar using outputs to train competing models [1]. Gemma's terms, for example, define "Model Derivatives" to include models trained to perform like Gemma by transferring patterns from its outputs, which reaches distillation and training on Gemma-generated synthetic data [3].

So a distilled dataset can be bound by two documents at once: the data license on the seeds and the generator terms on the output. Record which model, version and terms date produced each synthetic batch; the provenance records for synthetic training data page lists the fields. Detailed output-term analysis for hosted APIs belongs to the provenance and fine-tuning clusters rather than this page.

Personal data does not disappear in synthetic output

Synthetic records generated from personal data can still be personal data, so the license must say which privacy standard the output meets. Under GDPR Recital 26, information is anonymous only if a person cannot be identified by means reasonably likely to be used [4]. The EDPB's Opinion 28/2024 applies a similar analysis to AI models trained on personal data, and addresses consequences when development processing was unlawful [5].

NIST SP 800-188 describes formal privacy methods such as differential privacy alongside traditional de-identification, and cautions that traditional methods have inherent limits [6]. Where seeds are protected health information, HIPAA de-identification under 45 CFR 164.514 (Safe Harbor or Expert Determination) sets the standard, and a limited data set is disclosed only under a data use agreement that limits its uses [7]. Treat synthetic output from a limited data set as bound by that agreement unless counsel concludes otherwise. The privacy analysis is covered in depth at whether synthetic data from licensed records is still personal data.

Contract points to add:

  • The de-identification standard the seeds met before delivery, and whether generation must preserve it.
  • Membership-inference or nearest-neighbor distance tests on outputs before any external release.
  • A prohibition on generating records that target or describe named individuals in the seed set.
  • Who bears notification duties if a synthetic record is later linked to a real person.

Disclosure and provenance obligations follow the synthetic set

Regulators and acquirers increasingly ask where synthetic data came from, so keep lineage from seed to output. California AB 2013 requires developers of generative AI systems offered to Californians to post documentation about training data, and its required items include whether synthetic data generation was used; disclosures were due 1 January 2026 [8]. An acquirer's diligence team will also want synthetic, licensed and scraped data separated in records [1].

License documentation across the industry is already unreliable. The Data Provenance Initiative found license omission rates above 70% and error rates above 50% on popular dataset hosting sites [9]. A synthetic set with no recorded seed license inherits that problem. Tag every synthetic record with a batch ID linked to the seed license, generator version and run date.

Pricing derived-data rights

Derived-data rights are a separate economic term, and buyers usually see one of three structures. Which one fits depends on how much of the licensor's value the synthetic set replaces.

StructureHow it worksFits whenBuyer watch-out
IncludedDerived Data rights bundled in the base feeInternal SFT augmentation, no resaleLicensor may narrow scope to compensate
Capped volumeRights cover up to an agreed record or token countDistillation runs with predictable sizeDefine how records are counted after filtering
Separate derived-data feeAdditional fee for retention after term or external transferSynthetic set will be sold, shared or outlive the licenseAvoid per-use royalties that need ongoing reporting you cannot produce

Compare these with the base models in AI data license pricing structures, and add the derived-data line to your overall position in the licensing cluster guide. If you are still scoping seed data, you can describe the records and intended derived uses to SourceX so data and licensing permissions are assessed as part of the request.

Sourcing licensed seed data for synthetic generation

SourceX sources operational datasets from US companies on request and manages the licensing process, with each dataset rights-reviewed and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, and every release is approved by the supplying company. Describe the seed data and generation plan you need on the SourceX buyer page.

Frequently asked questions

Can I use licensed data to generate synthetic data for fine-tuning only?

Yes, if the license defines that use. A fine-tuning-only license that is silent on generation leaves the output's status open, so add a Derived Data rider before generating rather than after.

Who owns synthetic data generated from proprietary data?

Ownership follows the contract, not the generation step. Absent an ownership clause, expect the licensor to argue the output is an adaptation of its data and the generator provider to apply its output terms.

Can derived synthetic data train a third-party or open-weight model?

Only with an explicit transfer right in the data license and generator terms that allow it. Some generator terms restrict training competing models on outputs or bring models trained on outputs under the generator's own license [1][3].

Sources

  1. Mayer Brown, "Synthetic Data as a Deal Asset: Ownership, Provenance and Diligence Considerations in AI Acquisitions" (2026). https://www.mayerbrown.com/en/insights/publications/2026/07/synthetic-data-as-a-deal-asset-ownership-provenance-and-diligence-considerations-in-ai-acquisitions
  2. hoop.dev, "Licensing Models for Synthetic Data Generation". https://hoop.dev/blog/licensing-models-for-synthetic-data-generation/
  3. Google AI for Developers, "Gemma Terms of Use" (2024). https://ai.google.dev/gemma/terms-archive/terms_04_01_24
  4. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  5. CMS (summary of European Data Protection Board opinion), "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  6. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  7. Electronic Code of Federal Regulations (HHS), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  8. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  9. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data