Provenance, rights and permitted use
Data Rights Attestation Template for AI Training Data Suppliers
Quick answer
A data rights attestation is a short statement that a supplier signs for each dataset or delivery. It records where the records came from, the legal basis for licensing them, what consent and notice existed, which opt-out and exclusion checks were run, and which uses are allowed. Below is a copyable template built on the Data Provenance Standards categories [1]. It is tied to a provenance ID and version [2] and comes with guidance on signatories, refreshes and how it relates to the license.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What a supplier attestation proves, and what it does not
An attestation is signed evidence of facts as of a date. It is not a promise to pay if those facts turn out to be wrong. It gives your reviewers a single document stating the supplier's position on source, rights, privacy and restrictions, so diligence can test those statements against the underlying records. Warranties, indemnities, remedies and the rights grant itself belong in the license. Draft those using the AI training rights grant clause guide and the data license term sheet.
Keeping the two separate stops a common failure mode. Suppliers resist an attestation that reads like an open-ended warranty, and procurement then drops it entirely. When the attestation sticks to facts ("these records were exported from our Zendesk instance between these dates under these customer terms"), suppliers are more likely to sign it. The license can then say that the attestation is accurate and that an inaccurate one is a breach.
| Document | Who produces it | What it carries | Where liability sits |
|---|---|---|---|
| Rights attestation (this page) | Supplier, per dataset or delivery | Factual statements on source, basis, consent, opt-outs, exclusions, de-identification | Referenced by the license |
| Chain-of-title pack | Supplier | Contracts, policies, assignments that prove the statements | Evidence only |
| Due diligence checklist | Buyer reviewers | Questions and pass/fail tests | Internal |
| License | Both parties | Grant, permitted uses, term, warranties, indemnity, audit | Contractual |
For the evidence behind each statement, see chain of title for AI training data. For the reviewer-side questions, use the training data due diligence checklist.
Fields the attestation should carry, mapped to the provenance standards
Build the attestation around the categories buyers already use for provenance metadata. The Data & Trust Alliance standards cover source, legal rights, privacy and protection, generation date, data type, generation method, intended use and restrictions [1]. The standards operate at the dataset level, and a unique provenance identifier ties the categories together [2]. Each attestation should therefore name one dataset ID and one version, not "the data".
These fields also feed downstream disclosures. As of October 2026, providers of general-purpose AI models in the EU publish a training-content summary using the Commission's template under AI Act Article 53(1)(d) [4]. California AB 2013 requires generative AI developers to post documentation about their training datasets, including their sources and whether they contain copyrighted material or personal information [5]. For high-risk systems, AI Act Article 10 expects data governance covering collection processes and data origin [6]. Regulation (EU) 2026/1744 amends Article 10 [11] and reportedly moved the high-risk dates to 2 December 2027 (Annex III) and 2 August 2028 (Annex I). Ask for the fields now, because collecting them after training is much harder.
The Data Provenance Standards explainer covers each category in depth. For record-level encoding, see the permitted-use metadata schema.
The template: copyable fields for a per-dataset attestation
The template below is a set of fields to paste into your supplier onboarding pack. It is a structure for counsel to adapt, not finished contract language. Keep each answer factual and point to an evidence location rather than pasting evidence inline.
Illustrative example: invented to show structure; it does not describe an available dataset.
DATA RIGHTS ATTESTATION
Attestation ID: ATT-____ Supersedes: ATT-____ / none
Dataset provenance ID: ________ Version / delivery number: ____
Manifest hash (SHA-256): ________ Record count in manifest: ____
1. SOURCE
1.1 Supplying entity (legal name, jurisdiction): ________
1.2 Source systems (e.g., Zendesk, Salesforce, Jira, SharePoint, call recorder): ________
1.3 Collection window (first and last record timestamps, UTC): ________
1.4 Generation method: [ ] system export [ ] new recording [ ] human-authored
[ ] machine-generated/synthetic [ ] derived (transcript, translation, summary)
1.5 Data type: [ ] structured [ ] unstructured [ ] mixed; formats: ________
1.6 Third-party or acquired data included? [ ] no [ ] yes -> list origin and license
2. RIGHTS BASIS
2.1 Basis for licensing: [ ] owned (created by employees in course of work)
[ ] held under customer/vendor contract permitting this use
[ ] licensed in from third party with sublicense right [ ] other: ____
2.2 Contracts, terms of service or DPAs that govern the records: ________
2.3 Any contract that restricts AI training or onward licensing? [ ] no [ ] yes -> excluded? ____
2.4 Records created under government contracts or grants? [ ] no [ ] yes -> clause: ____
2.5 Evidence location (data room path): ________
3. PERSONAL DATA, CONSENT AND NOTICE
3.1 Personal data present before preparation? [ ] no [ ] yes -> categories: ____
3.2 Notice or consent in effect at collection (policy name, version, date): ________
3.3 Legal basis relied on, per jurisdiction (e.g., GDPR Art. 6 basis): ________
3.4 Recording-consent method for audio/video (two-party notices, signed releases): ____
3.5 De-identification method applied: ________ Tool/version: ________
3.6 Health data: [ ] none [ ] HIPAA Safe Harbor [ ] HIPAA Expert Determination (report ref ____)
3.7 Residual risks known to the supplier: ________
4. OPT-OUT AND EXCLUSION CHECKS
4.1 Individual opt-outs or deletion requests honored as of (date): ________
4.2 Customers or counterparties who declined AI use, and how they were removed: ________
4.3 Machine-readable reservations checked (TDM reservation, robots.txt,
Content Credentials "do not train"): [ ] not applicable [ ] checked -> result ____
4.4 Excluded categories (e.g., children's data, biometrics, privileged legal files,
payment card data, export-controlled technical data): ________
4.5 Sampling or scan used to confirm exclusions (method, sample size, date): ________
5. INTENDED USE AND RESTRICTIONS
5.1 Uses the supplier understands are permitted: [ ] pre-training [ ] fine-tuning/SFT
[ ] evaluation [ ] retrieval (RAG) [ ] other: ____
5.2 Restrictions known to the supplier (field, territory, model type, output use): ____
5.3 Cross-reference: governing license ID ________ (the license controls if in conflict)
6. ONGOING DUTIES
6.1 Supplier will notify the buyer within ____ days if any statement becomes untrue.
6.2 Re-attestation required for each refresh, new delivery or schema change.
7. SIGNATURE
Name / title: ________ Authority basis (role, board or delegated authority): ________
Date: ________ Statements are true to the signatory's knowledge after reasonable inquiry.
Before signing, run the attestation fields past the provenance documentation check to see which supporting documents a process typically requests. Match the manifest hash against a delivery manifest like the sample manifest.
How to word the rights basis and consent sections so they can be tested
Ask for statements a reviewer can check against a named document. "We own all rights" cannot be tested. "Records were authored by our employees under our employment agreements and IP assignment policy v4 (data room 2.3)" can be. For employee-authored material, pair section 2.1 with the checks in employee-authored data rights.
Service providers holding client records are a frequent gap. A managed-service firm may hold a ticket archive while its customers' master service agreements limit secondary use. Section 2.3 forces that question; for the follow-up review, see customer contracts and DPAs.
For EU personal data, require the supplier to name its legal basis instead of writing "compliant with GDPR". EDPB Opinion 28/2024 sets out how legitimate interest may apply to AI model development and warns that unlawful processing during development can affect later use of the model [8]. Collect the notice text that was live during the collection window via consent and notice records.
For US health data, the attestation should state which HIPAA method was used. Safe Harbor removes 18 identifier types; Expert Determination relies on a qualified expert's documented analysis. Neither method removes all re-identification risk [7].
Opt-out checks: what to attest and where evidence lives
Section 4 covers signals that operational exports often miss. The GPAI Code of Practice copyright chapter commits signatories, when crawling the web, to reproduce and extract only lawfully accessible content and to respect machine-readable rights reservations [3]. Content Credentials can also carry a creator's "do not train" preference [9]. Most first-party business records carry neither signal, so "not applicable, records are internal system exports" is a valid answer. But the supplier must say so explicitly.
Where files were sourced from public or partner sites, require the date of the robots.txt and TDM check and where the logs are kept. The EU TDM opt-out guide explains how Article 4 reservations work. Exclusions in section 4.4 should name the scan or sample that confirmed them. "No children's data" with no supporting check is a statement a reviewer cannot rely on.
Signatories, versioning and re-attestation for refreshes
The signatory must hold authority to bind the supplier on the facts attested. Typically that is a data owner or general counsel, with the authority basis written down. A data engineer who ran the export can co-sign as to the technical facts in sections 1 and 4.5 but should not sign alone.
Treat attestations like dataset versions. Each delivery, refresh or schema change gets a new attestation ID that names the one it supersedes, and the manifest hash binds the statement to exact files. Log each attestation in your training data use register so that a later model card or Article 53 summary can trace back to signed statements [4]. Dataset documentation practice such as Data Cards treats upstream sources and intended use as core fields, and the attestation is the signed version of those entries [10].
When an attestation and the delivered data disagree, quarantine the affected records and follow provenance gap remediation. Do not quietly patch the paperwork. The data provenance buyer's guide shows where attestations fit among the other provenance records. Teams that would rather describe the data they need and receive per-dataset diligence materials can start on the SourceX buyer page.
Sourcing operational data with attestations ready for review
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases. Every dataset is reviewed for ownership and consents, prepared with diligence materials covering source, rights, preparation and allowed use, and delivered under a license that defines records, uses, term and delivery. Every release is approved by the supplying company, and a request does not guarantee a match. Describe the data you need at the SourceX buyer page.
Frequently asked questions
Is a data origin attestation letter the same as a provenance certificate?
In practice, buyers use both terms for the same supplier-signed statement. No regulator issues a "provenance certificate" for licensed training data. Its value comes from who signs it, what evidence it cites, and whether the license makes inaccuracy a breach.
Should we ask for one attestation per supplier or per dataset?
Ask for one per dataset version or delivery. Rights, consent terms and exclusions differ by source system and collection window. A supplier-wide letter cannot be matched to a manifest.
Can the attestation replace our own diligence?
No. It narrows what reviewers need to test, and it creates a record of the supplier's statements. Reviewers should still sample the records and read the cited contracts.
Sources
- IAPP, "Leading corporations' proposed data provenance standards aim to enhance quality of AI training data". https://iapp.org/news/a/leading-corporations-proposed-data-provenance-standards-aims-to-enhance-quality-of-ai-training-data
- Help Net Security, "Cross-industry standards for data provenance in AI" (2024). https://www.helpnetsecurity.com/2024/07/22/saira-jesani-data-trust-alliance-data-provenance-standards/
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (bill text)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
- MIT Technology Review, "Adobe wants to make it easier for artists to blacklist their work from AI scraping" (2024). https://www.technologyreview.com/2024/10/08/1105234/adobe-wants-to-make-it-easier-for-artists-to-blacklist-their-work-from-ai-scraping
- Pushkarna, Zaldivar, Kjartansson (Google Research), FAccT 2022, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- European Parliament and Council of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.