Skip to content

Provenance, rights and permitted use

Provenance Records for Synthetic Training Data: Generator, Prompts, Seed Data and Terms

Quick answer

Synthetic data provenance means recording, for every generated record, the model and exact version that produced it, the terms that governed that model's outputs when you generated it, the prompt or template, the seed records it was conditioned on, the decoding settings, the filters it passed and any human edits. Without those fields nobody can later answer the questions counsel and acquirers ask: which model generated this dataset, what that model was trained on, and whether its output terms or the seed data's license restrict the use you intend [1].

By SourceX Editorial · Updated

Why synthetic data needs its own provenance record

Synthetic data inherits risk from two upstream sources, the generator and the seed data, and a standard dataset card captures neither at the level needed. A generated SFT example is a derived artifact: its usability depends on the output terms of the model that wrote it and on the rights in whatever real records were placed in its context window. Mayer Brown's July 2026 deal analysis advises buyers to ask which model generated the synthetic data and what that model was trained on, and warns that restrictions on outputs can leave a synthetic dataset unusable for its purpose [1].

The failure is usually discovered late. A team distills a teacher model into a student, ships the student, and only during an acquisition, a customer security review or an AB 2013 disclosure does someone ask which teacher produced the 400,000 instruction pairs. If the answer is "a mix of whatever API we had keys for that quarter," the dataset becomes a liability rather than an asset.

For background on what synthetic data is and how it compares with licensed and scraped sources, see the glossary entry on synthetic data and the guide licensed vs synthetic vs scraped AI training data. This page covers only the record itself; the cluster hub on data provenance for AI training data covers provenance for non-synthetic sources.

Generator identity: model, version, endpoint and training provenance

Record the generator precisely enough that someone could find the exact terms and model card in force on the generation date. "GPT-class model" or "Llama" is not provenance; a pinned identifier is. Hosted APIs change behavior under stable aliases, and open-weight families ship multiple checkpoints under one name with different licenses.

Capture these generator fields for each generation run:

  • Model identifier and snapshot. The dated snapshot string the API returned in the response, not the alias you requested. For open weights, the repository, revision hash and quantization.
  • Access path. Direct provider API, a cloud marketplace endpoint, or self-hosted inference. Terms can differ by channel, so record the contracting entity.
  • Generator training provenance. What the provider discloses about the model's own training data: a model card, an EU training-content summary, or an AB 2013 posting. Diligence counsel now ask this directly [1].
  • Teacher role. Whether the model generated content, judged or filtered it, or rewrote human text. A model used only as a classifier filter may carry different output terms from one that authored the text.

Output terms: the clause that decides whether the dataset is usable

The terms version in force when you generated the data is the single most consequential provenance field, because output restrictions attach at generation time. Several model licenses restrict using outputs to build other models. Google's Gemma Terms of Use, for example, define "Model Derivatives" to include models trained to perform similarly to Gemma through distillation or through training on synthetic data Outputs generated by Gemma, which pulls such student models under the Gemma terms [2].

Hosted providers publish their own positions on training with outputs; Anthropic's help center, for instance, has a dedicated article on whether outputs can be used to train an AI model [3]. Read the version that applied on your generation dates and archive it, because these pages change. The enforceability of output restrictions is debated in the literature, with Lemley and Henderson arguing that many such restrictions rest on weak legal footing [4], but a contested clause is still a diligence finding that a buyer or acquirer will price.

Store the terms as evidence, not as a link: a dated PDF or HTML snapshot with a SHA-256 hash, the clause numbers that address outputs, and your account's contracting entity. For clause-by-clause guidance on provider terms, see training on other models' outputs.

Prompts, templates and decoding settings

Prompt provenance lets you reproduce a record, explain its distribution and spot leakage of protected content. Version system prompts and templates in source control and record the template ID and commit hash on every record, not just the run. Persona libraries, topic lists and few-shot exemplars are inputs too; if the few-shot exemplars were copied from a benchmark, every generated record carries contamination risk (see contamination through synthetic data).

Decoding settings matter for both reproducibility and audit. Log temperature, top_p, max tokens, stop sequences, the random seed where the API accepts one, and the number of samples drawn per prompt. When you later find a cluster of near-duplicates or a memorized passage, these fields tell you whether it came from a low-temperature run that should be regenerated.

Seed data: where licensed rights carry into generated records

Seed data rights travel with the synthetic output, so each record must point to the seed record IDs it was conditioned on, not just the seed dataset name. If a support ticket, contract clause or engineering incident was placed in the prompt, the generated paraphrase is derived from that record and is shaped by its license, consent basis and any deletion obligation. Whether a license permits generating synthetic data at all is covered in rights to derived datasets; whether the output can still be personal data is covered in synthetic data privacy risk.

At record level, store the seed dataset ID, the license or agreement reference, the seed record IDs (or a salted hash of them if IDs are sensitive), and the transformation applied: paraphrase, extraction, question generation, persona simulation or counterfactual. The seed data for synthetic instruction generation page covers how to choose and prepare seeds; this record is what makes a later takedown or license expiry executable. When a seed license ends or a data subject's record must be removed, you need to find every synthetic descendant, which dataset-level notes cannot do (see record-level provenance).

Filters, judges and human edits

Every filter, judge and human edit changes what the dataset is, so each needs a logged identity and version. Record deduplication method and threshold (for example MinHash with a Jaccard cutoff), PII scrubbers and their configuration, toxicity or quality classifiers with versions, and any LLM-as-judge model with its own snapshot and terms. A judge model is a generator for provenance purposes when its rationales or rewrites end up in the data.

Human edits need the same treatment as annotation: who edited (an annotator pool ID, not a name), under which agreement, with which guidelines version, and whether the editor used AI assistance. The human-annotation provenance page lists the annotator agreement fields. Mark each record's final state as generated, generated-then-edited, or human-written-with-model-assistance, since buyers and diligence reviewers may treat these differently.

An illustrative record schema

The record below shows one way to encode the fields on each synthetic example as a JSON sidecar or as extra columns in a Parquet file. Dataset-level summaries can then be generated from these records into a Hugging Face dataset card, whose YAML header holds license and size metadata [9], or into Croissant JSON-LD, which describes resources and record structure in a schema.org vocabulary [8].

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "syn-sft-000184223",
  "record_state": "generated_then_edited",
  "generated_at": "2026-03-14T09:22:07Z",
  "generator": {
    "role": "author",
    "provider": "example-provider",
    "model_snapshot": "example-model-2026-02-01",
    "access_path": "direct_api",
    "contracting_entity": "Example Labs Inc.",
    "terms_snapshot_sha256": "9f2c...e41a",
    "terms_output_clauses": ["3.2", "4.1"],
    "generator_training_disclosure": "model card v2, AB 2013 posting archived 2026-01-05"
  },
  "prompt": {
    "template_id": "support-resolution-v7",
    "template_commit": "a41b9e0",
    "few_shot_ids": ["fs-022", "fs-031"],
    "decoding": {"temperature": 0.7, "top_p": 0.95, "max_tokens": 1024, "seed": 1187, "n": 1}
  },
  "seed": {
    "dataset_id": "licensed-support-tickets-2025q4",
    "agreement_ref": "LIC-2025-118",
    "record_ids_hashed": ["h:5c1e...", "h:a90d..."],
    "transformation": "paraphrase_with_resolution_steps"
  },
  "filters": [
    {"name": "minhash_dedup", "version": "1.3", "threshold": 0.8, "result": "pass"},
    {"name": "pii_scrubber", "version": "2.0", "result": "2 spans replaced"},
    {"name": "judge", "model_snapshot": "example-judge-2026-01-15", "score": 8, "result": "pass"}
  ],
  "human_edit": {"editor_pool": "pool-B", "guideline_version": "g-4", "ai_assisted": false},
  "permitted_use": ["sft"],
  "excluded_use": ["evaluation"]
}

The permitted_use and excluded_use fields are the payoff. They are computed from the generator terms and the seed agreement, and they stop a training team from moving a record into an eval split, where generated test sets cause their own problems (see synthetic evaluation data limits).

How the record feeds disclosures and diligence

A complete record turns regulatory and deal questions into queries rather than investigations. California's AB 2013 requires developers of generative AI systems offered to Californians to post training-data documentation, including whether synthetic data generation was used in development; those postings were due 1 January 2026 [7]. In the EU, Article 53 requires general-purpose AI model providers to publish a summary of training content and maintain a copyright policy [6], and the AI Office's July 2025 template sets the baseline for that summary [5].

Cross-industry metadata schemes point the same way. The Data & Trust Alliance Data Provenance Standards include categories for source, lineage, legal rights and generation method [10], which map directly onto the generator, seed and terms blocks above. The cluster's Data Provenance Standards explainer walks through each category.

Use this pre-acceptance checklist before a synthetic dataset enters a training mix or a data room:

CheckPass conditionCommon failure
Generator pinnedEvery record has a dated snapshot, not an alias"latest" alias logged; snapshot unknown
Terms archivedHashed terms snapshot per generation windowLink to current terms page only
Output clause reviewedCounsel note on distillation or competing-model language [2]Student model falls under teacher's derivative definition
Seeds traceableSeed record IDs or hashes on each recordOnly the seed dataset name recorded
Seed license allows generationAgreement reference with derived-data clauseGeneration not addressed in the license
Filters versionedFilter and judge versions loggedJudge model swapped mid-run without a log
Human edits attributedPool ID, guideline version, AI-assistance flagEdits made in a spreadsheet with no trail
Use flags setpermitted_use and excluded_use populatedEval and train splits drawn from one pool

For older corpora that fail this checklist, the options are covered in datasets without provenance.

Choosing licensed seed data with provenance you can carry forward

The cleanest synthetic pipeline starts from seed data whose rights were documented before generation. SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records and documents, and each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which gives your seed block a real agreement reference. If you need seed data for generation, describe it on the SourceX buyer page; requests are sourced rather than pulled from stock, and a request does not guarantee a match.

For mixing strategies once you have both kinds of data, see combining licensed and synthetic data.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request seed data for documented synthetic generation

SourceX finds US businesses that hold the operational data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Describe the seed data your synthetic pipeline needs at https://sourcex.si/buyers.

Sources

  1. Mayer Brown, "Synthetic Data as a Deal Asset: Ownership, Provenance and Diligence Considerations in AI Acquisitions" (2026). https://www.mayerbrown.com/en/insights/publications/2026/07/synthetic-data-as-a-deal-asset-ownership-provenance-and-diligence-considerations-in-ai-acquisitions
  2. Google, "Gemma Terms of Use (archived April 1, 2024 version)" (2024). https://ai.google.dev/gemma/terms-archive/terms_04_01_24
  3. Anthropic (Claude Help Center), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  4. SpicyIP, "Discussing Lemley and Henderson's \"The Mirage of Artificial Intelligence Terms of Use Restrictions\"" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
  5. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  6. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  7. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  8. Akhtar et al. (MLCommons), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  9. Hugging Face, "Dataset Cards". https://huggingface.co/docs/hub/en/datasets-cards
  10. Help Net Security, "Cross-industry standards for data provenance in AI" (2024). https://www.helpnetsecurity.com/2024/07/22/saira-jesani-data-trust-alliance-data-provenance-standards/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data