Skip to content

Schemas, packaging and delivery

Tracing Which Models Trained on Which Licensed Dataset

Quick answer

To track which models were trained on which dataset, give every licensed delivery an immutable version ID tied to its manifest digest, log that ID as an input on every training, fine-tuning, distillation and eval run, propagate it to every derived artifact (tokenized shards, filtered subsets, synthetic data, eval splits), and copy the resolved list into the model registry entry for each checkpoint. The test is one query: "everything that touched license X, version Y." If answering it needs interviews, the lineage does not exist yet.

By SourceX Editorial · Updated

Why lineage has to be captured at run time, not reconstructed later

Dataset-to-model lineage only exists if the training job writes it down while it runs. Once weights are trained, nothing inside the checkpoint reliably names its inputs; influence-estimation methods such as TracIn work from saved checkpoints and gradients over a selected subset of layers, and they estimate influence rather than prove usage [8]. Public dataset ecosystems show what happens without discipline: one audit of more than 1,800 text datasets found license information omitted in over 70% of cases on popular hosting sites and mislabeled in over 50% [4].

For licensed data the stakes are contractual as well as technical. When a license ends, a supplier asks for a destruction certificate, or counsel needs to know whether a product falls inside the allowed uses, the team must produce a list of affected checkpoints and derived datasets. That list is a database query or it is a guess. The contract side (what happens to trained models at termination) is covered in what happens to trained models when a data license ends; this page covers the records that make any of those terms enforceable.

The identifiers that must exist before the first training run

Traceability starts with a stable dataset version identifier that cannot be reused for different bytes. Assign it at intake, derive it from content, and bind it to the license that governs it.

  • license_id: your internal reference for the executed agreement, including amendments. Every downstream record points here.
  • dataset_version_id: immutable, one per delivery or snapshot. Our guide to versioning licensed datasets covers snapshot naming and release notes.
  • manifest_digest: a SHA-256 over the delivery manifest, so the version ID resolves to exact bytes. See dataset manifests and checksums.
  • allowed_uses: a machine-readable copy of permitted purposes (for example pre-training, SFT, eval, distillation, RAG) entered by counsel, not engineers.
  • storage_uri: the bucket prefix or table path, with access controls scoped to the license.

Dataset cards and Croissant files help, but they describe the dataset, not its consumers. A Hugging Face card carries license and size in YAML front matter [6], and Croissant-RAI adds machine-readable life cycle metadata [5]; neither records which runs read the data. Treat them as inputs to your registry, as discussed in Croissant metadata for licensed datasets.

Logging dataset versions in experiment tracking and the model registry

Every run must log the exact dataset_version_id it read, with a context label, before the first optimizer step. In MLflow this is mlflow.log_input() with a context such as "training"; in Weights & Biases it is an artifact reference; in a homegrown stack it is a row written by the data loader. For licensed data, log a reference (URI plus digest) rather than copying records into the tracking server, which would create another untracked copy.

Three habits close the common gaps:

  1. Log from the data loader, not the launch script. Config files drift; the loader knows which shards it actually opened. Emit the resolved list of dataset_version_ids plus shard-level digests.
  2. Tag with license_id. A tag such as license_id=LIC-0042 on each input lets the registry answer license questions without a join to a contracts system.
  3. Promote lineage into the registry entry. When a checkpoint is registered, copy the union of inputs from its run and all parent runs (warm starts, resumed jobs, LoRA bases) into the model version's metadata. A model version whose lineage field is empty should fail promotion.

Warm starts are the most common break. A fine-tune that starts from an internal base checkpoint inherits every dataset that base was trained on, so lineage must be transitive, walking parent checkpoints back to the first pre-training run.

Tracking derived artifacts: shards, filters, synthetic data and eval sets

Most licensed data reaches a model through a derivative, so every derived artifact needs its own version ID plus a pointer to its parents. The W3C PROV vocabulary is a ready model: an activity used an entity, a new entity wasGeneratedBy that activity, and the output wasDerivedFrom the input. You do not need an RDF store; the same three edges fit in a relational table.

The derivatives that usually escape tracking:

  • Tokenized and packed shards. Mixing sequences from several sources into one packed shard means a shard can carry three licenses. Record per-shard source proportions.
  • Filtered and deduplicated subsets. Quality filters and MinHash deduplication produce new versions; log the filter config hash and parent IDs.
  • Synthetic data generated from licensed data. Prompts seeded with licensed records, or a teacher model fine-tuned on them, produce outputs whose rights trace back. Some model licenses already treat this chain explicitly: Gemma's terms define Model Derivatives to include models trained to imitate Gemma through its outputs, including distillation and synthetic data [7]. Check whether your data licenses use similar derivative language.
  • Eval and holdout splits. An eval set cut from licensed data is still licensed data, and benchmark results computed on it may be shared with partners.
  • Embeddings and vector indexes. A RAG index built from licensed documents is a derived artifact with its own deletion problem; see embedding and vector index rights.

For rights in transcripts, translations and summaries made from source records, see tracing rights in derived records.

Illustrative lineage record and the termination query

A workable schema needs four tables: licenses, dataset versions, artifacts (datasets, shards, checkpoints, indexes, products) and lineage edges. The example below shows the shape of one edge record and the query that answers a termination or audit request.

Illustrative example: invented to show structure; it does not describe an available dataset.

lineage_edge:
  edge_id: E-2026-118734
  relation: wasDerivedFrom          # PROV-O style: used | wasGeneratedBy | wasDerivedFrom
  child_artifact: ckpt/support-assist-8b/sft-v3/step-12000
  child_type: checkpoint
  parent_artifact: ds/support-tickets/v2026-07-15
  parent_type: dataset_version
  parent_manifest_sha256: 9f2c...e41a
  license_id: LIC-0042
  use_context: sft                  # pre-training | sft | eval | distillation | rag
  run_id: mlflow:4c1e7b0a
  recorded_by: dataloader-hook/1.8
  recorded_at: 2026-08-02T14:11:09Z
-- Everything downstream of one license, any depth
WITH RECURSIVE downstream AS (
  SELECT child_artifact, child_type FROM lineage_edge
  WHERE license_id = 'LIC-0042'
  UNION
  SELECT e.child_artifact, e.child_type
  FROM lineage_edge e JOIN downstream d ON e.parent_artifact = d.child_artifact
)
SELECT child_type, count(*) FROM downstream GROUP BY child_type;
Question you will be askedRecord that answers itTypical failure mode
Which checkpoints used license X?Recursive walk of lineage edgesWarm-start parents not logged
Which products serve those checkpoints?Registry deployment recordsServing aliases point to retired versions
Which copies must be destroyed?Artifact table with storage_uriScratch copies and notebook caches untracked
Was the use inside allowed_uses?use_context vs allowed_usesEval split reused for SFT
What goes in a training-content summary?Dataset versions per released modelSynthetic data not attributed to sources

When the query returns a list, pair it with a certificate of data destruction workflow so deletion evidence references the same artifact IDs.

Feeding regulatory summaries and documentation from the same records

The same lineage graph is the input for regulatory training-data disclosures. Under Article 53 of the EU AI Act, providers of general-purpose AI models must keep technical documentation, maintain a copyright policy and publish a sufficiently detailed summary of training content [1]; the Commission's mandatory template for that summary was published in July 2025 [2]. As of October 2026, these obligations have applied since 2 August 2025, and the AI Office's enforcement powers apply from 2 August 2026 [1], so per-model lists of data sources should come from lineage, not memory.

High-risk system providers also face documentation-keeping duties for a fixed period after placing a system on the market [3], which means lineage records often need to outlive the licensed data itself. That tension between retention duties and license deletion duties is covered in training data retention requirements, and whether a model trained on personal data can count as anonymous is covered in EDPB Opinion 28/2024 for data buyers. Keep the lineage records (IDs, digests, edges) separate from the data bytes so you can delete one and retain the other.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What to ask a data supplier so lineage works on day one

Lineage is cheaper when the delivery arrives with stable identifiers. Ask suppliers for a manifest with per-file checksums, a version label that changes whenever the content changes, a data dictionary, and release notes listing additions, removals and corrections between versions. Ask, too, for a stated record key so that a later correction or withdrawal request can be mapped to specific shards rather than the whole corpus.

On the buyer side, read the license with the lineage table open: every permitted use should map to a use_context value, and every deletion or termination obligation should map to a query. The general contract vocabulary is explained in AI data license terms, and common buyer practice is discussed in do AI labs delete data after training and how long buyers keep licensed data. For wider delivery mechanics, start at the delivery and schemas hub or the AI data hub.

If you are sourcing operational data from US companies, SourceX handles the commercial process: each dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery, with per-dataset diligence materials covering source, rights, preparation and allowed use. Those materials are a natural seed for your license and dataset version tables; you can describe the data you need on the buyers page.

Sourcing licensed data you can trace from delivery to model

SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records and documents, and manages licensing and ongoing purchases. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Describe the data, uses and delivery you need at sourcex.si/buyers.

Sources

  1. European Commission AI Act Service Desk, "Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  2. WilmerHale, "European Commission Releases Mandatory Template for Public Disclosure of AI Training Data" (2025). https://wilmerhale.com/en/insights/blogs/wilmerhale-privacy-and-cybersecurity-law/european-commission-releases-mandatory-template-for-public-disclosure-of-ai-training-data
  3. European Commission AI Act Service Desk, "Article 18: Documentation keeping". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-18
  4. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  5. Jain et al., MLCommons, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  6. Hugging Face, "Dataset Cards". https://huggingface.co/docs/hub/en/datasets-cards
  7. Google AI for Developers, "Gemma Terms of Use (archived 1 April 2024 version)" (2024). https://ai.google.dev/gemma/terms-archive/terms_04_01_24
  8. Pruthi et al., "Estimating Training Data Influence by Tracing Gradient Descent" (2020). https://arxiv.org/pdf/2002.08484

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data