Skip to content

Provenance, rights and permitted use

Data Provenance for AI Training Data: A Buyer's Guide to Source, Rights and Permitted Use

Quick answer

Data provenance for AI training is the documented evidence of where each dataset and record came from, who had the right to license it, which consents and opt-outs apply, how it was processed, and what the buyer may do with it. Set it as a requirement for every licensed dataset and verify it with documents, not assurances: a source inventory, chain-of-title records, notice and consent versions, opt-out logs, license texts, synthetic-generation records and a permitted-use register.

By SourceX Editorial · Updated

Five provenance questions and the evidence that answers each

Every provenance requirement reduces to five questions, each answered by specific documents. CASRAI defines training-data provenance as documentation of sources, collection methods, licensing, consent basis, time range and processing steps [1]; buyers also need opt-out and permitted-use evidence. For the basic definition, see what data provenance is and why buyers care and the data provenance glossary entry.

QuestionEvidence to requestTypical failureGo deeper
Where did it come from?Source inventory: system of record, export method, date range, every prior holder"Aggregated from partners" with no named sourceChain-of-title documents; resold and brokered data
Who agreed?Notice and terms versions at collection, consent records, customer contracts and DPAs, staff and contractor agreementsTerms changed after collection; a DPA limits use to service deliveryConsent and notice records; customer contracts and DPAs
Who opted out?Dated robots.txt, TDM reservation and content-credential checks; a withdrawal listSignals checked once, long after collectionEU TDM opt-outs; opt-out evidence log
What is it?Copyright status per source, original license texts, synthetic-generation and annotation recordsDataset-card license copied without reading the upstream licenseOpen dataset license audit; synthetic data provenance records
What may it be used for?Allowed and prohibited uses per dataset or record; register entries; training-run logsOne license for a corpus whose sources carry different termsPermitted-use metadata schema; training data use register

Within the AI data buyer's map, this hub covers the evidence. Contract terms that allocate risk when evidence fails, such as warranties and indemnities, belong to the AI training data licensing guide; de-identification methods to the de-identified data hub; statute-by-statute duties to the AI training data compliance hub.

Why a license field or a supplier's word is not provenance

Labels and assurances fail often enough to be treated as claims to test. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission rates above 70% and error rates above 50% on popular dataset hosting sites [2]. On the Hugging Face Hub, a dataset's license comes from a YAML metadata block in the dataset card that the uploader writes [3], so it records what someone declared, not what the original source permits.

Permission also changes over time. The "Consent in Crisis" audit of 14,000 web domains found that between 2023 and 2024 roughly 5% of tokens in the C4 corpus, and more than 28% of its most actively maintained critical sources, became fully restricted, and that about 45% of C4 was restricted once terms of service were counted [4]. A provenance file therefore needs dates: when each record was collected, and which terms and signals applied on that date.

Where it came from: systems of record, custody and chain of title

Origin evidence names the system that created each record and every party that held it before the buyer. For operational data, that means the system of record (a helpdesk, CRM, ERP, issue tracker or document management system), the export method, the collection window and the legal entity that controlled the system. Each extra hop adds a document: client authorization when a service provider holds a client's data, SaaS platform terms for exported data, contractor IP assignments and policies covering employee-authored records.

How the supplier acquired the data matters as well as who holds it. In the Bartz v. Anthropic class action, which involved author claims over training data, the case settled with final approval granted in July 2026 [5]. Buyers screening for that risk can start with pirated and shadow-library source screening.

Who agreed: notices, consents and customer contracts

Consent evidence shows that the people and organizations behind the data were told about, or agreed to, a use that covers AI training. Ask for the privacy notice, terms of service and customer contract versions in force when each record was collected, not the current ones. FTC staff warned in February 2024 that adopting more permissive data practices, such as AI training, through a surreptitious, retroactive change to terms or privacy policies may be unfair or deceptive [6].

The consequences can reach the model: the FTC's 2021 final order against Everalbum required deletion of models and algorithms developed using users' uploaded photos and videos [7]. Matching each record to the notice and terms in force at collection shows which records a notice covers. For personal information, add the de-identification evidence package.

Who opted out: TDM reservations, robots.txt and content credentials

Opt-out evidence matters most for web-derived text and published media, and it must be dated. Under Article 53(1)(c) of the EU AI Act, providers of general-purpose AI (GPAI) models must keep a copyright policy that identifies and complies with rights reservations made under Article 4(3) of the Digital Single Market (DSM) Directive [8]. The GPAI Code of Practice copyright chapter asks signatories to crawl only lawfully accessible content and to honor machine-readable reservations, including robots.txt as specified in IETF RFC 9309 [9].

Signals vary by medium. The TDM Reservation Protocol (TDMRep) is a W3C Community Group specification, not a W3C Standard, that lets a site declare a reservation and point to licensing terms [10]. For media files, C2PA stated in January 2026 that its core Content Credentials specification has no standard TDM assertion [11]; the Creator Assertions Working Group's cawg.training-mining extension marks training as allowed, constrained or not allowed [12]. As of late 2025, courts had not settled what counts as machine-readable: the Hamburg Higher Regional Court's December 2025 ruling in Kneschke v. LAION required an opt-out that machines can interpret, while Dutch and Danish courts took different approaches [13].

SourceX does not source scraped public web content; for operational records licensed from the company that created them, contracts and notices carry most of the provenance weight. For web-derived components, compare robots.txt, ai.txt, TDMRep and other AI usage signals.

Content evidence classifies each source by copyright status and origin, because the same license wording means different things for owned records, third-party works and generated text. Sort sources into owned, licensed, public domain or unknown, and quarantine the unknown class until resolved (copyright status classification). Open datasets need the original license text, including non-commercial and share-alike terms inherited from upstream sources in aggregated datasets.

Synthetic data needs its own record, created at generation time: generator model and version, prompts or templates, seed data and its license, and the provider terms in force at generation. Provider terms can restrict training on outputs; as of October 2026, Anthropic's help center, for example, says its terms do not allow outputs to be used to train models that compete with its own [14]. Annotated and preference data needs annotator agreements and disclosure of any AI assistance (human annotation provenance); licensed vs synthetic vs scraped data compares the source types.

What it may be used for: permitted-use metadata at record level

Permitted-use evidence turns the license into machine-readable fields that travel with the data, so pipelines can filter records by allowed use. The Data & Trust Alliance's Data Provenance Standards, released in January 2024, cover a data entry's source, legal rights and privacy protections, a timestamp, how data was generated, data type, and intended uses and restrictions [15]; one of its leads has said the standards track a dataset's origin and creation method [16]. When a corpus mixes sources with different terms, buyers need the same fields per record (when dataset-level documentation isn't enough).

Datasheets for Datasets frame the human-readable questions on motivation, composition, collection and recommended uses [17]. Croissant-RAI is a machine-readable format for that documentation [18], and NeurIPS 2026 requires responsible-AI metadata for its Evaluations and Datasets Track [19].

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "provenance_id": "prv-2025-0412",
  "dataset_id": "ds-support-tickets-v3",
  "record_id": "tkt-00418273",
  "source": {
    "supplier": "Supplier A (US logistics firm)",
    "system_of_record": "helpdesk ticketing system",
    "export_method": "admin CSV export, query ref Q-17",
    "collected_at": "2023-06-14T09:42:00Z",
    "custody": ["Supplier A"]
  },
  "agreement": {
    "customer_terms_version": "ToS v4.2, effective 2022-11-01",
    "privacy_notice_version": "PN-2023-01",
    "contract_refs": ["MSA-077 s.9.3"]
  },
  "opt_out_checks": "not applicable: non-web source",
  "content": {
    "copyright_status": "owned",
    "origin": "human",
    "third_party_material": "attachments removed"
  },
  "processing": [
    {"step": "pii_replacement", "method": "NER + pattern rules, surrogate values", "date": "2025-02-10"},
    {"step": "dedup", "method": "exact hash", "date": "2025-02-11"}
  ],
  "permitted_use": {
    "license_ref": "LIC-2025-014",
    "allowed": ["pre-training", "fine-tuning", "evaluation"],
    "prohibited": ["retrieval display to end users", "resale"],
    "license_end": "2028-02-28",
    "status": "active"
  }
}

The permitted_use block feeds the AI training data register. Training-run input logs that record which licensed records each model saw let a team trace which models trained on a licensed dataset and act on record-level takedown and withdrawal obligations.

Which laws and frameworks now ask for provenance records

Several AI laws require, or will soon require, developers to describe training data in ways that depend on supplier provenance records.

RequirementWho and when (as of October 2026)Provenance fields it draws onDeep dive
EU AI Act Art. 53(1)(c) copyright policyGPAI model providers; obligations are subject to enforcement by the AI Office and national competent authorities [8]Opt-out checks, acquisition method, licensesArticle 53 training data obligations
EU AI Act Art. 53(1)(d) public summary of training contentGPAI model providers, using the Commission template published on 24 July 2025 [20]Source categories, data types, collection periodsCompleting the summary template
EU AI Act Art. 10(2) data governanceHigh-risk AI systems; covers collection processes and the origin of data [21]; Regulation (EU) 2026/1744 reportedly moved Annex III application to 2 December 2027 [22]Origin, collection method, preparation stepsArticle 10 data governance
California AB 2013Generative AI developers serving Californians; due 1 January 2026 and before each substantial modification; covers sources or owners, copyrighted or licensed data, personal information, collection periods, synthetic data [23]Source inventory, copyright status, synthetic flagAB 2013 supplier records
Colorado SB26-189Developers of automated decision-making technology used in consequential decisions; documentation to deployers, including training data categories, from 1 January 2027 [24]Data categories, intended useUS state AI laws and training data

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

How to verify a supplier's provenance claims

Verification means tracing a sample of records to evidence that existed before the sale, such as system exports, dated contracts and signed agreements. Labels such as "rights-cleared" are conclusions, not evidence (provenance red flags). A workable sequence:

  1. Get the file first. Request the source inventory, chain-of-title documents, notice versions, license texts, opt-out logs and synthetic-generation records before the sample; a due diligence questionnaire structures the request.
  2. Trace a random sample. Draw record IDs stratified by source system and collection year, and ask to see each in the system-of-record export with its timestamp and the notice or contract in force that day (sample testing).
  3. Read licenses at the source. Open each open or third-party license where the original publisher posted it, not a hosting site's metadata field [2].
  4. Re-run opt-out checks. For web-derived components, re-check signals for a sample of URLs and compare results and dates with the supplier's log.
  5. Reconcile the manifest. Record counts, file hashes and provenance IDs should agree across sample, file and delivery; a dataset bill of materials can package them.
  6. Bind evidence to the contract. Get a signed data rights attestation that references the file, and use data warranties for residual risk; a warranty allocates loss but does not prove title.

The training data due diligence checklist covers pre-signing items, and the data source documentation check lists documents to expect with licensed business data. For data you already hold, run a corpus provenance audit and remediate, re-license, quarantine or retire each gap.

If you source operational data through SourceX, it looks for US businesses that hold the data you describe, and the supplying company approves every release. Each dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and diligence materials on source, rights, preparation and allowed use are prepared per dataset for your review. You can state the provenance evidence your team requires in your request.

Provenance priorities by training stage

The five questions apply at every stage, but the evidence that decides a purchase depends on how the data will be used; procurement by training stage covers the commercial side.

StageProvenance questions that dominateStart with
Pre-trainingAcquisition path, opt-outs and copyright status across many sources; inputs to the GPAI training-content summaryPre-training rights grant
Fine-tuning and post-trainingAnnotator agreements, AI-assisted labels, synthetic origin and output terms, employee-authored contentFine-tuning datasets; training on other models' outputs
EvaluationWhether items appeared in public sources or training corpora, and who has seen the holdoutEvaluation datasets; quality and contamination
Retrieval (RAG)Rights to store, display and quote at inference time; takedown and update handlingGrounding license vs training license; RAG content licensing

Need to know where your training data came from?

Describe the operational data you need and the provenance evidence your governance process requires. SourceX looks for US businesses that hold that data, checks the data and the supplier's licensing permissions, and manages a license that defines which records are included, what they can be used for and how delivery happens; nothing is contracted until a supplier agrees. See how sourcing and rights review work.

Guides in this section

Sources

  1. CASRAI, "Training data provenance". https://casrai.org/dictionary/term/training-data-provenance
  2. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  3. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  4. Longpre et al., "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
  5. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  6. Federal Trade Commission, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  7. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  8. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  9. European Commission, "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  10. W3C TDM Reservation Protocol Community Group, "TDM Reservation Protocol (TDMRep)". https://www.w3.org/2022/tdmrep
  11. C2PA, "C2PA clarification to C2PA TDM assertions reference" (2026). https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
  12. IPTC Metawatch, "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
  13. Kluwer Copyright Blog, "LAION Round 2: Machine-Readable but Still Not Actionable: The Lack of Progress on TDM Opt-Outs (Part 1)". https://legalblogs.wolterskluwer.com/copyright-blog/laion-round-2-machine-readable-but-still-not-actionable-the-lack-of-progress-on-tdm-opt-outs-part-1/
  14. Anthropic Help Center, "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  15. IAPP, "Leading corporations' proposed data provenance standards aim to enhance quality of AI training data" (2023). https://iapp.org/news/a/leading-corporations-proposed-data-provenance-standards-aims-to-enhance-quality-of-ai-training-data
  16. Help Net Security, "Cross-industry standards for data provenance in AI" (2024). https://www.helpnetsecurity.com/2024/07/22/saira-jesani-data-trust-alliance-data-provenance-standards/
  17. Gebru et al., "Datasheets for Datasets" (2018; CACM 2021). https://arxiv.org/pdf/1803.09010
  18. Jain et al., "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  19. NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  20. European Commission, "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  21. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  22. European Parliament and Council, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  23. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  24. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data