Skip to content

Provenance, rights and permitted use

Datasets Without Provenance: Remediate, Re-License, Quarantine or Retire

Quick answer

Training data without provenance needs a recorded decision, not a shrug. For each flagged dataset, answer four questions in order: can the missing provenance be recovered, can rights be obtained from whoever holds them, can the data be replaced with licensed records, and how exposed are the models already trained on it. The answers map to five outcomes: remediate the documentation, re-license, replace, quarantine, or retire. Log every outcome and its rationale in your training data register.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why provenance gaps are the normal audit result

Expect a long list of gaps, because missing and wrong license metadata is the default state of public datasets. The Data Provenance Initiative traced more than 1,800 text datasets and reported license omission above 70% and license error rates above 50% on popular hosting sites [1]. A corpus assembled from hub downloads, internal exports and vendor deliveries will reproduce those error rates unless someone has checked each entry.

The gap types differ, and so do the fixes. Separate them before deciding anything:

  • Missing license: no license file, or a hub card that says "other" or "unknown".
  • Misattributed license: the aggregator's license (often a permissive tag on the wrapper) differs from the license of the underlying sources.
  • Broken chain of title: a broker or reseller cannot show where it got the data or what it was allowed to pass on.
  • Missing consent or notice records: personal data with no evidence of the notice or contract terms that covered it.
  • Unknown generation: synthetic or model-labeled records with no record of the generator, its terms or the seed data.

If you have not yet produced this inventory, start with the provenance audit of an existing training corpus; this guide picks up where the audit's findings list ends.

The four questions that decide each dataset

The decision runs dataset by dataset, and it must also be made at the model and application layers, because license risk carries forward from a dataset into every checkpoint trained on it and every product built on that checkpoint [2]. Ask these four questions in order and stop at the first answer that resolves the gap.

  1. Can provenance be recovered? Look for the original paper, the upstream repository's commit history, archived dataset cards, vendor statements of work, purchase orders and data processing agreements. If the gap was a documentation failure, not a rights failure, you remediate.
  2. Can rights be obtained? Identify who holds the rights today: the original creator, the company whose operational records these are, or the platform. If that party will grant a license for your actual use, you re-license.
  3. Can the data be replaced? Ask whether the capability the dataset provides (a domain, a task format, a language, an eval slice) can be rebuilt from licensed sources at comparable quality. If yes, you replace.
  4. What is the model exposure? Determine which checkpoints, fine-tunes and deployed systems consumed the dataset. Exposure decides whether you quarantine (stop future use, keep everything for audit) or retire (stop use and plan model-level action).

Illustrative example: invented to show structure; it does not describe an available dataset.

FindingProvenance recoverable?Rights obtainable?Replaceable?Model exposureOutcome
Instruction set, hub tag "apache-2.0", sources unlistedYes, paper lists sourcesn/an/aTwo SFT checkpointsRemediate; record source licenses
Customer-support transcripts from an acquired companyPartly; DPA silent on trainingPossibly, via current controllerYesOne production fine-tuneRe-license or replace; quarantine meanwhile
Scraped forum dump, origin unknownNoNoYesPre-training mix, 0.4% of tokensRetire from future mixes; document exposure
Eval set of contract clauses, author unknownNoNoYesUsed only for scoringRetire and rebuild eval
Synthetic dialogues, generator unrecordedNoUnknown termsYesNone yetQuarantine; regenerate with recorded settings

Remediate when the gap is documentation, not rights

Remediation is the right call when the rights existed all along and only the record is missing. Typical cases are datasets whose source papers enumerate their inputs, internal exports with a traceable system of origin, or vendor deliveries whose contracts sit in procurement rather than in the ML repo.

Write the recovered facts into a machine-readable record rather than a wiki note. Croissant expresses dataset metadata, file resources and record structure as schema.org JSON-LD, which lets license and source fields travel with the files into loaders and catalogs [5]. Pair it with a datasheet that answers the motivation, composition, collection process and intended-use questions Gebru and colleagues proposed [6]. The Data Provenance Standards explained page lists the metadata categories worth filling.

Remediation fails in a predictable way: the team "recovers" provenance by copying the aggregator's license tag. Verify the license of each underlying source, since aggregated collections are where the audit found many licenses miscategorized or omitted [1].

Re-license when a rights holder exists and will agree

Re-licensing applies when you can name the party that holds the rights and that party is willing to grant terms covering training. For an open dataset released under a noncommercial license such as CC BY-NC, the dependable path to commercial training use is a separate grant from the licensor, not a reinterpretation of the public license.

Scope the new license to what the gap actually was. A useful re-license names the exact records (by manifest or hash list), the permitted uses (pre-training, SFT, evaluation, retrieval), whether past training is covered, the term, and what happens to derived models at expiry. Retroactive coverage of models already trained is the clause most often missing, and without it the re-license only fixes future runs.

Copyright analysis of training remains unsettled. The U.S. Copyright Office's Part 3 report on generative AI training is, as of October 2026, still a pre-publication version [3], so do not treat any single fair-use argument as a substitute for a license where a willing licensor exists. Collect the evidence the chain of title documents page describes before relying on any grant.

Replace when the capability matters more than the records

Replacement is often cheaper than disputing provenance when the dataset's value is a capability, not the specific records. Define the replacement by the capability: task format, domain, language, label schema, volume, time range and the eval slices it must reproduce.

Replacement data from operating businesses can be stronger than the original. Support ticket histories with resolution codes, engineering incident records, or finance and legal workflow documents carry real distributions that scraped approximations miss. When you request licensed operational data, describe the data and use rather than a named supplier; SourceX buyer sourcing works this way, looking for US businesses that hold the described data, with every release approved by the supplying company. Sourcing is on request, and a request does not ensure a match.

Run the replacement through the same tests you will apply to any new supplier. The provenance claim sample testing page covers how to check a supplier's statements on a sample before signing, and the licensed vs synthetic vs scraped comparison helps decide which source type fits the capability.

Quarantine: excluded from training, retained for audit

Quarantine means the dataset is blocked from every new training, fine-tuning and evaluation run, but the files, manifests, hashes and lineage records are kept intact. You keep them because a regulator, auditor, licensor or court may later ask what a specific checkpoint was trained on, and deleted evidence cannot answer.

Enforce quarantine in the pipeline, not by memo. Practical controls include a deny-list of dataset IDs and content hashes checked by the data loader, read-only storage with access limited to governance and legal, a flag in the training data register, and a block in the mix configuration tooling. Quarantine is a holding state with an owner and a review date, not a permanent answer.

Quarantine evaluation data with extra care. A benchmark whose items have leaked into training corpora stops measuring generalization; OpenAI stopped reporting SWE-bench Verified scores after concluding it had become contaminated [10]. An eval set of unknown origin carries both risks: you cannot show you may use it, and you cannot show it is clean.

Retire, and decide what happens to trained models

Retirement is the outcome when provenance cannot be recovered, rights cannot be obtained, and keeping the data in circulation is not defensible. Retiring the dataset stops future use; the harder decision is what to do with checkpoints already trained on it.

Model-level options, from lightest to heaviest:

  • Document and accept: record exposure (token share, epochs, which checkpoints) and the rationale for continued use. Fits low-share, low-sensitivity cases.
  • Restrict deployment: keep the checkpoint for research or internal use only, outside products that face end users.
  • Fine-tune forward: stop branching new work from the affected base, and retrain downstream adapters on a clean base.
  • Retrain: rebuild from a clean mix. SISA-style sharded training reduces the cost of removing data later, but it must be designed in before training, not retrofitted [7].

For general-purpose models placed on the EU market, the training-content summary under AI Act Article 53(1)(d) uses the Commission's July 2025 template [8], so the retirement record should show what you disclosed and when the data left the mix.

Recording the decision in your register

Every outcome needs an entry that someone outside the ML team can audit. NIST's AI RMF puts this work under GOVERN and MANAGE: assigned accountability, documented risk decisions and tracked responses [4]. ISO/IEC 5259-4 provides a process framework for data quality in ML that the same records can feed [9].

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset_id: ds-0417-support-transcripts
finding: missing_consent_records
gap_detail: "Acquired-company DPA silent on model training"
decision: quarantine
interim_until: 2027-01-31
next_step: re_license_or_replace
models_exposed: [ft-support-v3, ft-support-v3.1]
exposure: "Fine-tune only; 18% of SFT examples"
loader_block: [dataset_id, sha256_manifest]
retained_artifacts: [manifest.jsonl, croissant.json, dpa_v2.pdf]
owner: ai-governance
approved_by: [privacy_counsel, head_of_ml]
rationale: "Rights path exists via current controller; replacement scoped in parallel"

Keep the register linked to the AI training data use register so that each dataset's permitted uses and its remediation history live in one place.

Get licensed replacements for datasets you retire

SourceX sources operational datasets from US companies on request, rights-reviews each one for ownership and consents, and delivers it under a license that defines the records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded, and nothing is contracted until a supplier agrees. Describe the capability you need to replace at SourceX for AI data buyers.

For the wider sourcing process, see how to procure enterprise training data and the provenance buyer's guide hub.

Frequently asked questions

Can we keep using a model trained on data we later retired?

Sometimes, as a documented risk decision rather than a default. Record the dataset's share of training, the checkpoints affected and why continued use is acceptable, and revisit it if a rights holder raises a claim.

Does quarantine mean deleting the dataset?

No. Quarantine blocks new use but keeps the files and lineage so you can answer later questions about what a model saw. Deletion belongs to retirement, and only after counsel confirms no retention duty applies.

Can a noncommercial open dataset be made commercial?

Reliably, only through a separate grant from the rights holder. The public license itself does not change, and contributors' own rights may limit what the licensor can grant.

Sources

  1. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  2. SoftwareSeni, "How AI licence risk compounds across your dataset, model and application stack". https://www.softwareseni.com/how-ai-licence-risk-compounds-across-your-dataset-model-and-application-stack/
  3. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  4. NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
  5. Akhtar et al., MLCommons (arXiv), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  6. Gebru et al. (arXiv), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
  7. Bourtoule et al., IEEE S&P 2021 (arXiv), "Machine Unlearning" (2019). https://arxiv.org/abs/1912.03817v2
  8. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  9. ISO, "ISO/IEC 5259-4:2024 Data quality for analytics and machine learning, Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  10. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data