Provenance, rights and permitted use
Data Provenance for AI Training Data: A Buyer's Guide to Source, Rights and Permitted Use
Quick answer
Data provenance for AI training is the documented evidence of where each dataset and record came from, who had the right to license it, which consents and opt-outs apply, how it was processed, and what the buyer may do with it. Set it as a requirement for every licensed dataset and verify it with documents, not assurances: a source inventory, chain-of-title records, notice and consent versions, opt-out logs, license texts, synthetic-generation records and a permitted-use register.
By SourceX Editorial · Updated
Five provenance questions and the evidence that answers each
Every provenance requirement reduces to five questions, each answered by specific documents. CASRAI defines training-data provenance as documentation of sources, collection methods, licensing, consent basis, time range and processing steps [1]; buyers also need opt-out and permitted-use evidence. For the basic definition, see what data provenance is and why buyers care and the data provenance glossary entry.
| Question | Evidence to request | Typical failure | Go deeper |
|---|---|---|---|
| Where did it come from? | Source inventory: system of record, export method, date range, every prior holder | "Aggregated from partners" with no named source | Chain-of-title documents; resold and brokered data |
| Who agreed? | Notice and terms versions at collection, consent records, customer contracts and DPAs, staff and contractor agreements | Terms changed after collection; a DPA limits use to service delivery | Consent and notice records; customer contracts and DPAs |
| Who opted out? | Dated robots.txt, TDM reservation and content-credential checks; a withdrawal list | Signals checked once, long after collection | EU TDM opt-outs; opt-out evidence log |
| What is it? | Copyright status per source, original license texts, synthetic-generation and annotation records | Dataset-card license copied without reading the upstream license | Open dataset license audit; synthetic data provenance records |
| What may it be used for? | Allowed and prohibited uses per dataset or record; register entries; training-run logs | One license for a corpus whose sources carry different terms | Permitted-use metadata schema; training data use register |
Within the AI data buyer's map, this hub covers the evidence. Contract terms that allocate risk when evidence fails, such as warranties and indemnities, belong to the AI training data licensing guide; de-identification methods to the de-identified data hub; statute-by-statute duties to the AI training data compliance hub.
Why a license field or a supplier's word is not provenance
Labels and assurances fail often enough to be treated as claims to test. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission rates above 70% and error rates above 50% on popular dataset hosting sites [2]. On the Hugging Face Hub, a dataset's license comes from a YAML metadata block in the dataset card that the uploader writes [3], so it records what someone declared, not what the original source permits.
Permission also changes over time. The "Consent in Crisis" audit of 14,000 web domains found that between 2023 and 2024 roughly 5% of tokens in the C4 corpus, and more than 28% of its most actively maintained critical sources, became fully restricted, and that about 45% of C4 was restricted once terms of service were counted [4]. A provenance file therefore needs dates: when each record was collected, and which terms and signals applied on that date.
Where it came from: systems of record, custody and chain of title
Origin evidence names the system that created each record and every party that held it before the buyer. For operational data, that means the system of record (a helpdesk, CRM, ERP, issue tracker or document management system), the export method, the collection window and the legal entity that controlled the system. Each extra hop adds a document: client authorization when a service provider holds a client's data, SaaS platform terms for exported data, contractor IP assignments and policies covering employee-authored records.
How the supplier acquired the data matters as well as who holds it. In the Bartz v. Anthropic class action, which involved author claims over training data, the case settled with final approval granted in July 2026 [5]. Buyers screening for that risk can start with pirated and shadow-library source screening.
Who agreed: notices, consents and customer contracts
Consent evidence shows that the people and organizations behind the data were told about, or agreed to, a use that covers AI training. Ask for the privacy notice, terms of service and customer contract versions in force when each record was collected, not the current ones. FTC staff warned in February 2024 that adopting more permissive data practices, such as AI training, through a surreptitious, retroactive change to terms or privacy policies may be unfair or deceptive [6].
The consequences can reach the model: the FTC's 2021 final order against Everalbum required deletion of models and algorithms developed using users' uploaded photos and videos [7]. Matching each record to the notice and terms in force at collection shows which records a notice covers. For personal information, add the de-identification evidence package.
Who opted out: TDM reservations, robots.txt and content credentials
Opt-out evidence matters most for web-derived text and published media, and it must be dated. Under Article 53(1)(c) of the EU AI Act, providers of general-purpose AI (GPAI) models must keep a copyright policy that identifies and complies with rights reservations made under Article 4(3) of the Digital Single Market (DSM) Directive [8]. The GPAI Code of Practice copyright chapter asks signatories to crawl only lawfully accessible content and to honor machine-readable reservations, including robots.txt as specified in IETF RFC 9309 [9].
Signals vary by medium. The TDM Reservation Protocol (TDMRep) is a W3C Community Group specification, not a W3C Standard, that lets a site declare a reservation and point to licensing terms [10]. For media files, C2PA stated in January 2026 that its core Content Credentials specification has no standard TDM assertion [11]; the Creator Assertions Working Group's cawg.training-mining extension marks training as allowed, constrained or not allowed [12]. As of late 2025, courts had not settled what counts as machine-readable: the Hamburg Higher Regional Court's December 2025 ruling in Kneschke v. LAION required an opt-out that machines can interpret, while Dutch and Danish courts took different approaches [13].
SourceX does not source scraped public web content; for operational records licensed from the company that created them, contracts and notices carry most of the provenance weight. For web-derived components, compare robots.txt, ai.txt, TDMRep and other AI usage signals.
What it is: copyright status, open licenses and synthetic origin
Content evidence classifies each source by copyright status and origin, because the same license wording means different things for owned records, third-party works and generated text. Sort sources into owned, licensed, public domain or unknown, and quarantine the unknown class until resolved (copyright status classification). Open datasets need the original license text, including non-commercial and share-alike terms inherited from upstream sources in aggregated datasets.
Synthetic data needs its own record, created at generation time: generator model and version, prompts or templates, seed data and its license, and the provider terms in force at generation. Provider terms can restrict training on outputs; as of October 2026, Anthropic's help center, for example, says its terms do not allow outputs to be used to train models that compete with its own [14]. Annotated and preference data needs annotator agreements and disclosure of any AI assistance (human annotation provenance); licensed vs synthetic vs scraped data compares the source types.
What it may be used for: permitted-use metadata at record level
Permitted-use evidence turns the license into machine-readable fields that travel with the data, so pipelines can filter records by allowed use. The Data & Trust Alliance's Data Provenance Standards, released in January 2024, cover a data entry's source, legal rights and privacy protections, a timestamp, how data was generated, data type, and intended uses and restrictions [15]; one of its leads has said the standards track a dataset's origin and creation method [16]. When a corpus mixes sources with different terms, buyers need the same fields per record (when dataset-level documentation isn't enough).
Datasheets for Datasets frame the human-readable questions on motivation, composition, collection and recommended uses [17]. Croissant-RAI is a machine-readable format for that documentation [18], and NeurIPS 2026 requires responsible-AI metadata for its Evaluations and Datasets Track [19].
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"provenance_id": "prv-2025-0412",
"dataset_id": "ds-support-tickets-v3",
"record_id": "tkt-00418273",
"source": {
"supplier": "Supplier A (US logistics firm)",
"system_of_record": "helpdesk ticketing system",
"export_method": "admin CSV export, query ref Q-17",
"collected_at": "2023-06-14T09:42:00Z",
"custody": ["Supplier A"]
},
"agreement": {
"customer_terms_version": "ToS v4.2, effective 2022-11-01",
"privacy_notice_version": "PN-2023-01",
"contract_refs": ["MSA-077 s.9.3"]
},
"opt_out_checks": "not applicable: non-web source",
"content": {
"copyright_status": "owned",
"origin": "human",
"third_party_material": "attachments removed"
},
"processing": [
{"step": "pii_replacement", "method": "NER + pattern rules, surrogate values", "date": "2025-02-10"},
{"step": "dedup", "method": "exact hash", "date": "2025-02-11"}
],
"permitted_use": {
"license_ref": "LIC-2025-014",
"allowed": ["pre-training", "fine-tuning", "evaluation"],
"prohibited": ["retrieval display to end users", "resale"],
"license_end": "2028-02-28",
"status": "active"
}
}
The permitted_use block feeds the AI training data register. Training-run input logs that record which licensed records each model saw let a team trace which models trained on a licensed dataset and act on record-level takedown and withdrawal obligations.
Which laws and frameworks now ask for provenance records
Several AI laws require, or will soon require, developers to describe training data in ways that depend on supplier provenance records.
| Requirement | Who and when (as of October 2026) | Provenance fields it draws on | Deep dive |
|---|---|---|---|
| EU AI Act Art. 53(1)(c) copyright policy | GPAI model providers; obligations are subject to enforcement by the AI Office and national competent authorities [8] | Opt-out checks, acquisition method, licenses | Article 53 training data obligations |
| EU AI Act Art. 53(1)(d) public summary of training content | GPAI model providers, using the Commission template published on 24 July 2025 [20] | Source categories, data types, collection periods | Completing the summary template |
| EU AI Act Art. 10(2) data governance | High-risk AI systems; covers collection processes and the origin of data [21]; Regulation (EU) 2026/1744 reportedly moved Annex III application to 2 December 2027 [22] | Origin, collection method, preparation steps | Article 10 data governance |
| California AB 2013 | Generative AI developers serving Californians; due 1 January 2026 and before each substantial modification; covers sources or owners, copyrighted or licensed data, personal information, collection periods, synthetic data [23] | Source inventory, copyright status, synthetic flag | AB 2013 supplier records |
| Colorado SB26-189 | Developers of automated decision-making technology used in consequential decisions; documentation to deployers, including training data categories, from 1 January 2027 [24] | Data categories, intended use | US state AI laws and training data |
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How to verify a supplier's provenance claims
Verification means tracing a sample of records to evidence that existed before the sale, such as system exports, dated contracts and signed agreements. Labels such as "rights-cleared" are conclusions, not evidence (provenance red flags). A workable sequence:
- Get the file first. Request the source inventory, chain-of-title documents, notice versions, license texts, opt-out logs and synthetic-generation records before the sample; a due diligence questionnaire structures the request.
- Trace a random sample. Draw record IDs stratified by source system and collection year, and ask to see each in the system-of-record export with its timestamp and the notice or contract in force that day (sample testing).
- Read licenses at the source. Open each open or third-party license where the original publisher posted it, not a hosting site's metadata field [2].
- Re-run opt-out checks. For web-derived components, re-check signals for a sample of URLs and compare results and dates with the supplier's log.
- Reconcile the manifest. Record counts, file hashes and provenance IDs should agree across sample, file and delivery; a dataset bill of materials can package them.
- Bind evidence to the contract. Get a signed data rights attestation that references the file, and use data warranties for residual risk; a warranty allocates loss but does not prove title.
The training data due diligence checklist covers pre-signing items, and the data source documentation check lists documents to expect with licensed business data. For data you already hold, run a corpus provenance audit and remediate, re-license, quarantine or retire each gap.
If you source operational data through SourceX, it looks for US businesses that hold the data you describe, and the supplying company approves every release. Each dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and diligence materials on source, rights, preparation and allowed use are prepared per dataset for your review. You can state the provenance evidence your team requires in your request.
Provenance priorities by training stage
The five questions apply at every stage, but the evidence that decides a purchase depends on how the data will be used; procurement by training stage covers the commercial side.
| Stage | Provenance questions that dominate | Start with |
|---|---|---|
| Pre-training | Acquisition path, opt-outs and copyright status across many sources; inputs to the GPAI training-content summary | Pre-training rights grant |
| Fine-tuning and post-training | Annotator agreements, AI-assisted labels, synthetic origin and output terms, employee-authored content | Fine-tuning datasets; training on other models' outputs |
| Evaluation | Whether items appeared in public sources or training corpora, and who has seen the holdout | Evaluation datasets; quality and contamination |
| Retrieval (RAG) | Rights to store, display and quote at inference time; takedown and update handling | Grounding license vs training license; RAG content licensing |
Need to know where your training data came from?
Describe the operational data you need and the provenance evidence your governance process requires. SourceX looks for US businesses that hold that data, checks the data and the supplier's licensing permissions, and manages a license that defines which records are included, what they can be used for and how delivery happens; nothing is contracted until a supplier agrees. See how sourcing and rights review work.
Guides in this section
- AI Training Data Register: Rights, Uses and Model LineageBuild an AI training data inventory that records each dataset's source, rights basis, permitted uses, restrictions, expiry and the models it trained.
- Chain of Title for AI Training Data: Documents by RouteWhich documents prove a supplier can license training data: route-by-route chain of title for originated, assigned, acquired, client and contractor data.
- Client Data Held by Service Providers: AI Licensing TestCan a BPO, agency or MSP license client data for AI training? The processor rules, the client authorization fields and the exclusion checks buyers need.
- Consent and Notice Records for AI Training DataWhich consent and notice records to request for AI training data, how to trace sampled records to them, and when consent is not the right legal basis.
- Customer Contracts, DPAs and AI Training Use: Buyer CheckHow AI data buyers check a B2B supplier's MSAs, DPAs and order forms before licensing customer-related records for model training, evals or agents.
- Data Provenance Standards: 8 Categories for Data BuyersHow the cross-industry Data Provenance Standards work, what each metadata category asks for, evidence to request from suppliers, and where the gaps are.
- Data Rights Attestation Template for AI Training DataCopyable data rights attestation template for AI training data suppliers: source, rights basis, consent, opt-out checks, exclusions and permitted uses.
- Employee-Authored Data for AI Training: Rights ChecksHow buyers verify that a supplier owns employee-authored emails, chats and documents and gave notices that cover licensing them for AI training.
- EU TDM Opt-Outs (DSM Article 4): Checks for AI Data BuyersWhat counts as a valid Article 4 text and data mining opt-out, how courts read machine-readable reservations, and what buyers must verify in acquired data.
- Open Dataset License Audit for Commercial AI TrainingA repeatable audit for open datasets in a training mix: trace each source, read the license text, classify use, flag unspecified terms, record decisions.
- Synthetic Data Provenance: Generator, Prompts, Seeds, TermsWhat a provenance record for synthetic training data should capture: generator and version, output terms, prompts, seed data, filters and human edits.
- Training Data Provenance Audit: Datasets Already in UseRun a provenance audit of training, fine-tuning, eval and RAG data already in use: inventory, trace, verify licenses and consent, risk-rate and remediate.
- AI Training Data Collection Consent Form: Required TermsWhat a consent form for commissioned AI data collection must say: training use, licensees, sharing, withdrawal, biometrics, minors and per-record proof.
- ai.txt vs robots.txt: AI Usage Signals ComparedCompare robots.txt, ai.txt, TDMRep, C2PA/CAWG content credentials and IETF AI preferences, and set a policy for which opt-out signals your datasets honor.
- Annotation Data Provenance: Rights, Guidelines, AI UseAnnotation data provenance for buyers: annotator IP assignment, guideline version tracking, per-label metadata and checks for LLM use in human labels.
- Are Business Records Copyrighted? Training-Data RightsWhich parts of tickets, logs and documents carry copyright, why contracts and trade secrets govern the rest, and what AI data buyers should license.
- Brokered Data Provenance: Verifying Every Sublicense HopHow AI data buyers trace a resold or brokered dataset back to its original holder: the documents to request at each hop and where sublicense chains break.
- C2PA Content Credentials as Training Data ProvenanceHow ML teams extract, validate and record C2PA manifests in licensed images, video and audio, and what Content Credentials can and cannot prove.
- Can a Platform License Its Users' Posts for AI Training?How to test whether a forum, review or Q&A site's user terms let it sublicense posts for AI training: grant scope, terms versions, opt-outs, embeds.
- CAWG Training-Mining Assertion: Reading Do-Not-Train FlagsHow to read cawg.training-mining flags in C2PA Content Credentials, map allowed, constrained and notAllowed to ingestion rules, and log decisions.
- Contractor IP Assignment Checks for AI Training DataHow AI data buyers verify that contractor and freelancer work was assigned to the supplier: agreements to request, clause tests, and exclusion rules.
- Copyright Status of Training Data: A Six-Class TaxonomyClassify every training-data source as owned, licensed, open-licensed, public domain, unknown or flagged, and set handling rules for each copyright class.
- Dataset Bill of Materials: SPDX 3.0 and CycloneDX for AIHow to record training datasets, sources and licenses in an AI bill of materials using SPDX 3.0 Dataset fields or CycloneDX data components.
- Permitted-Use Metadata Schema for Training Data RecordsField-by-field schema for recording permitted uses, prohibitions, expiry and territory on each training record, with ODRL mapping and fail-closed filters.
- Pirated Books in AI Training Data: A Screening WorkflowHow AI labs screen book and article corpora for shadow-library copies: manifests, hash and metadata matching, acquisition records and removal steps.
- Privacy Policy Version at Collection: Mapping Training DataMap each training record to the privacy notice and terms version in force when it was collected, and handle data gathered before AI training was disclosed.
- Provenance Red Flags When Buying AI Training DataWarning signs that a data supplier's provenance claims will not hold up, and how to escalate, re-scope or walk away before you license AI training data.
- Record-Level Data Provenance for Mixed-Rights DatasetsWhen dataset-level provenance breaks down: how to track consent, notice version and permitted use per record or token, and what fine-grained lineage costs.
- SaaS Terms and Exported Data: Can It Be Licensed for AI?How AI data buyers check SaaS platform terms before licensing exported CRM, helpdesk and call data, and which vendor-generated fields to strip or verify.
- TDM Opt-Out Evidence Log: Fields, Rules and RetentionDesign an evidence log that proves which TDM opt-out signals were checked, when, with what result and what was excluded from an AI training corpus.
- Testing a Supplier's Provenance Claims on a SampleHow to verify a data supplier's provenance claims on sampled records: stratify the sample, trace each record to evidence, and set pass/fail rules.
- Third-Party Content in Licensed Training CorporaHow AI data buyers find attachments, quoted text, stock media and syndicated articles inside licensed corpora, and decide whether to exclude or clear them.
- Third-Party Rights in Screen Recordings and Agent DataWhose rights attach to software UIs, documents and customer data captured in screen recordings and agent trajectories, and what buyers should verify first.
- Trade Secrets in AI Training Data: NDA and Secret ScreeningHow AI buyers keep other companies' trade secrets and NDA-covered material out of licensed corpora: DTSA exposure, screening signals and exclusion rules.
- Training Data Without Provenance: Five Remediation PathsWhat to do with training data that has no provenance: a decision guide to remediate, re-license, replace, quarantine or retire datasets after an audit.
- Training on LLM Outputs: Checking Provider Output TermsCan you train a model on LLM outputs? How to trace model-generated records in a dataset, read provider output terms, and test the competing-model clause.
- Upstream License Tracing in Aggregated AI DatasetsHow to trace every component of an instruction-tuning collection to its source license, propagate the strictest terms and record it in a provenance card.
- Web-Scraped Dataset Provenance: Records to DemandWhich crawl-time records to require from web-data vendors: crawl dates, user agents, robots.txt and terms snapshots, TDM opt-outs, takedown propagation.
Sources
- CASRAI, "Training data provenance". https://casrai.org/dictionary/term/training-data-provenance
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
- Longpre et al., "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- Federal Trade Commission, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission, "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- W3C TDM Reservation Protocol Community Group, "TDM Reservation Protocol (TDMRep)". https://www.w3.org/2022/tdmrep
- C2PA, "C2PA clarification to C2PA TDM assertions reference" (2026). https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
- IPTC Metawatch, "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
- Kluwer Copyright Blog, "LAION Round 2: Machine-Readable but Still Not Actionable: The Lack of Progress on TDM Opt-Outs (Part 1)". https://legalblogs.wolterskluwer.com/copyright-blog/laion-round-2-machine-readable-but-still-not-actionable-the-lack-of-progress-on-tdm-opt-outs-part-1/
- Anthropic Help Center, "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- IAPP, "Leading corporations' proposed data provenance standards aim to enhance quality of AI training data" (2023). https://iapp.org/news/a/leading-corporations-proposed-data-provenance-standards-aims-to-enhance-quality-of-ai-training-data
- Help Net Security, "Cross-industry standards for data provenance in AI" (2024). https://www.helpnetsecurity.com/2024/07/22/saira-jesani-data-trust-alliance-data-provenance-standards/
- Gebru et al., "Datasheets for Datasets" (2018; CACM 2021). https://arxiv.org/pdf/1803.09010
- Jain et al., "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- European Commission, "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.