Skip to content

Provenance, rights and permitted use

Testing a Supplier's Provenance Claims on a Sample

Quick answer

To verify data provenance claims, treat every statement in the supplier's datasheet as a hypothesis and test it on records you choose. Draw a stratified sample across source systems, date ranges, customers and record types, then ask the supplier to trace each sampled record to its native system ID, creation date, the notice or terms version in force, any consent artifact, and the result of its exclusion checks. Agree pass/fail thresholds and consequences in the pilot terms before evidence arrives.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why provenance statements fail without record-level testing

Provenance statements fail because they are usually written at dataset level by the party selling the data, and dataset-level summaries hide record-level exceptions. The Data Provenance Initiative audited more than 1,800 text datasets and found license information omitted in over 70% of cases and license errors above 50% on popular hosting sites [1]. Vendors now market provenance as a product feature [2], which is useful, but a marketing explainer or a datasheet is still self-reported until a sample backs it.

The common failure modes are specific. A supplier describes "our support tickets" but the export includes tickets from a reseller's tenant; a "2019 onward" date range includes migrated legacy records with reset creation dates; a privacy notice that permitted product improvement was replaced mid-period by one that did not; or an exclusion list (opt-outs, minors, regulated accounts) was applied to the CRM but not to the attachment store. None of these show up in a dataset card. They show up when you pull a record and ask where it came from.

Testing also protects you downstream. As of October 2026, California AB 2013 requires developers of generative AI systems offered to Californians to post documentation about their training data, including its sources [9], and the FTC has warned that companies can be liable when they use customer data in ways that break their own privacy promises [6]. A buyer who cannot trace a sample to evidence has little to put in that documentation. For the broader framework, start with the data provenance guide for AI training data.

Designing the sample: stratify by where errors hide

A provenance sample should be stratified by the dimensions along which the supplier's evidence changes, not drawn uniformly at random. Errors cluster at boundaries: a second source system, a system migration, a notice change, a large customer with its own contract, or a record type with attachments. A uniform draw from a ticket corpus where 90% of rows come from one system will barely touch the other 10%, and that 10% is often where the problem is.

Build strata from the supplier's own description, then add strata for known change points:

  • Source system: each system named in the manifest (for example Zendesk, Salesforce Service Cloud, Jira, a document management system, a file share).
  • Date range: split at every notice or terms change, system migration and acquisition date the supplier discloses.
  • Customer or tenant: top accounts by record volume sampled individually, with the long tail pooled.
  • Record type: structured rows, free-text fields, attachments, call recordings and transcripts tested separately.
  • Exclusion-sensitive segments: records near opt-out dates, accounts flagged as minors or regulated, and health or financial categories.

Sample size depends on the error rate you need to rule out; the sample size guide for estimating a dataset's error rate covers the arithmetic. As a rule, you choose the record IDs, not the supplier. If you have not yet negotiated sample access, the mechanics are in how to request a training data sample from a supplier.

Tracing each record to source-system evidence

Each sampled record should trace to an identifier and timestamp that exist in the system it came from, not only in the supplier's export. Most business systems keep a stable native key: a Zendesk ticket ID, a Jira issue key such as ENG-4812, a Salesforce record ID, or an External ID field that links a Salesforce record to its originating system [3]. Ask the supplier to show the native key, the system's own created timestamp and the export job that produced the row.

Watch for three trace failures. First, keys generated at export time (UUIDs with no mapping table) cannot be traced back at all. Second, created dates that cluster on a single day usually mark a migration, which means the true origin is an earlier system the supplier has not disclosed. Third, attachments and embedded files often carry their own provenance, such as a customer's contract uploaded into a ticket, that the parent record's rights do not cover.

The record-level provenance guide explains when this depth is warranted across a whole corpus rather than a sample.

Integrity matters as well as origin. NIST SP 800-218A extends secure development practice to the integrity of training, testing and fine-tuning data [7], which in a pilot means you should get a hash manifest for the sample files and confirm the full delivery later matches the same pipeline.

Checking notices, consents and exclusions on sampled records

For each sampled record, the supplier should be able to name the notice, terms or consent version that applied when the record was created and show that the use you plan falls within it. A dataset-level statement that "customers agreed to our terms" is not evidence; the evidence is a versioned document plus a mapping from record dates to versions. Where consent is individual (call recording consent, research consent, biometric consent), ask for the artifact or log entry tied to that record or session.

Exclusions need their own test. Ask the supplier to run its exclusion logic on your sampled IDs and return the result per record: opt-out lists, deletion requests, accounts under legal hold, minors, and contractually excluded customers. Then spot-check by naming a few records you expect to be excluded and confirming they are absent from the full extract.

For health data, check whether the supplier claims HIPAA de-identification or a limited data set; a limited data set still contains PHI, excludes 16 categories of direct identifiers and is disclosed under a data use agreement [5], so the evidence you need is different. The consent and notice records guide lists the document types to request.

Provenance sample test worksheet

The worksheet below is the core pilot artifact: one row per sampled record, one column per claim being tested. Fill it from evidence the supplier produces, not from their summary.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldWhat the supplier must showPass conditionExample entry
sample_idYour assigned sample IDMatches your draw listS-0142
stratumSource system / date band / record typeMatches the sampling frameService desk / 2021-H2 / ticket + attachment
native_record_idKey in the originating systemResolves in a live system screen or audit exportTicket 88213
created_at_sourceCreation timestamp from the source systemWithin the claimed date range; not a migration date2021-09-14T16:02Z
export_job_idPipeline run that produced the rowSame pipeline as full deliveryexp-2026-08-run-07
notice_versionTerms or privacy notice in force at creationVersion text permits the planned useCustomer terms v4.2 (2021-03)
consent_artifactRecord-level consent, if the claim relies on oneArtifact present and dated before capturen/a (contract basis)
exclusion_checkResult of opt-out, deletion, minor, hold and excluded-customer checksAll clear, with check dateClear, run 2026-08-20
third_party_contentEmbedded content owned by others (attachments, quoted email)Flagged and covered or removedVendor invoice PDF removed
redaction_methodHow personal fields were removed or replacedMethod named; residual check passedNamed-entity replacement; manual review
file_hashSHA-256 of the delivered fileMatches the manifest3f9a...c21
findingPass, minor, or majorPer pilot termsMinor: attachment rights unclear

Record the reviewer and date on each row. A worksheet like this doubles as the start of a training data use register and supports the data statement or datasheet your team publishes internally [8].

Setting pass/fail thresholds in the pilot terms

Thresholds and their consequences belong in the pilot terms, written before the supplier sends evidence, so results cannot be renegotiated after the fact. Buyer guides on data vendors recommend fixing pass/fail criteria ahead of the order and putting refund or replacement terms in writing [4]; for provenance, the criteria differ by severity rather than by a single error rate.

Illustrative example: invented to show structure; it does not describe an available dataset.

Finding classExampleSuggested consequence
MajorRecord cannot be traced to a source system; notice in force did not permit the use; excluded record presentStratum fails; supplier must explain root cause and re-run exclusions on the full stratum before any purchase
SystemicSame major finding in two or more records of one stratumDrop the stratum or re-scope the license; consider the provenance gap remediation options
MinorMissing export job ID, unflagged quoted email, late hashFix before delivery; does not block
PassAll fields evidencedNo action

Decide in advance who adjudicates disagreements (usually counsel plus the data lead) and what the supplier must do on a failed stratum: remove it, re-prepare it, or produce the missing evidence within the pilot. Map any finding that touches ownership rather than consent to the chain of title documents, because a missing assignment or platform-terms restriction is not fixable by redaction. Common warning signs are collected in provenance red flags when buying AI training data.

Fitting provenance testing into the wider pilot

Provenance testing runs alongside, not instead of, quality and fitness testing. The SourceX guide to evaluating data supplier quality covers completeness, label accuracy and duplication, and how to run a data pilot with a supplier covers scoping, timelines and access. This page covers only the question those guides leave open: whether sampled records are what the supplier says they are, from where it says, under terms that allow your use.

Sequence matters. Run provenance tests first on a small stratified draw, because a failed source or notice test can remove a whole stratum and change the quality sample you would otherwise build. Then run quality tests on what survives, and only then size the full purchase.

SourceX prepares diligence materials per dataset covering source, rights, preparation and allowed use, and every dataset it delivers is rights-reviewed for ownership and consents and supplied under a license that defines records, uses, term and delivery. Buyers who want to review those materials for a dataset they need can start at the SourceX buyer page.

Verifying provenance claims with SourceX

SourceX sources operational datasets from US companies on request, and every release is approved by the supplying company, so you describe the data and the provenance evidence you need rather than a named business. Personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement. Describe the data you need provenance evidence for.

Sources

  1. Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  2. Zyte, "What is AI data provenance?". https://dev.zyte.com/learn/what-is-ai-data-provenance/
  3. Salesforce, "upsert() (SOAP API Developer Guide)". https://developer.salesforce.com/docs/atlas.en-us.api.meta/object_ref/sforce_api_calls_upsert.htm
  4. CloudPano, "Selecting the Best AI Training Data Provider: A Practical Buyer's Guide". https://www.cloudpano.com/blog/selecting-best-ai-training-data-provider-buyers-guide
  5. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  6. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  7. National Institute of Standards and Technology, "Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile (NIST SP 800-218A)" (2024). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-218A.pdf
  8. Bender and Friedman, Transactions of the ACL, "Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science" (2018). https://aclanthology.org/Q18-1041/
  9. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data