Skip to content

Privacy and preparation

Data Provenance Standards explained for data suppliers

By SourceX Editorial · Updated

Short answer

The Data Provenance Standards are an open metadata specification from the Data & Trust Alliance that groups dataset information into Source, Provenance and Use, so AI developers can select training data with known origins and rights. For a supplier, filling them in means documenting where records came from, how they were prepared and what the license permits.

Key takeaways

  • Version 1.0.0 of the Data Provenance Standards organizes dataset metadata into three groups: Source, Provenance and Use.
  • The Use group carries the rights and privacy fields buyers check first, including license to use, consent documentation location and privacy-enhancing technologies applied.
  • The specification is published under a Creative Commons license, so suppliers can adopt its structure freely.
  • Most answers come from records a supplier already keeps: system inventories, contracts, notices and preparation logs.

What are the Data Provenance Standards?#

The Data Provenance Standards are a metadata specification published by the Data & Trust Alliance, a not-for-profit consortium that the specification says was established in September 2020. Version 1.0.0 groups dataset metadata into Source, Provenance and Use, and states that this metadata is needed 'to enable proper dataset selection for AI Model Training.'

The standards are released under the Creative Commons CC-BY-SA-4.0 license, with code and supporting artifacts under Apache 2.0, so a supplier can use the structure without asking permission. They are voluntary: no law requires them, but they give buyers and suppliers a shared vocabulary for questions every licensing deal asks anyway.

For a company licensing operating records, the value is practical. A buyer's diligence questionnaire and a supplier's documentation can follow the same field list instead of being reconciled line by line over email.

What does each group ask a supplier to document?#

Each group in the Data Provenance Standards asks a supplier a different question. Source identifies the dataset and who stands behind it, Provenance describes where the data came from and how it came to exist, and Use sets out the rights, privacy treatment and permitted purposes.

The GitHub version of the specification lists six elements under Source, nine under Provenance and twelve under Use. The table summarizes the question each group answers, where a supplier usually finds the answers, and which part of the SourceX Evidence Packet holds the same information.

Release authorization, the fifth part of the Evidence Packet, has no direct counterpart among the three groups. It records who approved the release and under which contract, which buyers usually request separately.

What does each group ask a supplier to document?
GroupQuestion it answersSupplier records that answer itSourceX Evidence Packet section
SourceWhat is this dataset and who issued it?Dataset name and description, issuing entity, contact, versionProvenance
ProvenanceWhere did the data come from, and when and how was it generated?System inventory, date ranges, collection method, formats, preparation lineageProvenance and privacy record
UseWhat may be done with it, under which rights and protections?License terms, notices and consents, confidentiality class, preparation methods, geography limitsLicensing rights, permitted use and privacy record

The Use group, element by element#

The Use group is the part a buyer's counsel reads first, because it carries rights and privacy information. Its elements include confidentiality classification, consent documentation location, privacy-enhancing technologies applied, data processing geography included or excluded, data storage geography allowed or forbidden, license to use, intended data use, and copyright, patent and trademark status.

The specification's code list for privacy-enhancing tools includes redaction, masking, minimization, pseudonymization, tokenization, encryption, k-anonymity, l-diversity, t-closeness and differential privacy, among others. Using those terms in your own preparation logs means the Use group largely fills itself in.

The Use group, element by element
Use elementWhat a supplier fills inWhere the answer comes from
Confidentiality classificationThe sensitivity level of the delivered dataYour data classification policy
Consent documentation locationWhere notices, consents or other legal bases are recordedPrivacy notices, employee notices, customer terms
Privacy-enhancing technologiesMethods applied, such as redaction, masking or pseudonymizationPreparation log and review results
Processing geographyWhere the data may or may not be processedContract terms and customer commitments
Storage geographyWhere the data may or may not be storedContract terms and security policy
License to useThe license under which the data is providedThe signed data license
Intended data usePermitted purposes, such as training or evaluationThe permitted-use clause
Copyright, patent and trademark statusWhether protected material is present and who owns itRights review notes and carve-outs for client material

Why buyers now ask for provenance metadata#

Buyers now ask for provenance metadata because they must explain their own training data choices to customers, auditors and regulators. California's AB 2013, signed on September 28, 2024, requires developers of generative AI systems to post training-data documentation that states, among other things, whether datasets were purchased or licensed and whether they include personal information. In the EU, Article 53(1)(d) of the AI Act requires general-purpose model providers to publish a summary of training content using an AI Office template, which the European Commission released on July 24, 2025.

Those duties sit with the AI developer, not the supplier, but they flow back through diligence. A buyer that has to say whether data was licensed, and whether it held personal information, will ask its suppliers for the facts in a form it can file.

The research community moved the same way. The Data Provenance Initiative audited 44 data collections spanning more than 1,800 fine-tuning text datasets, documenting their sources, licenses and creators, and its tools generate a Data Provenance Card for any filtered subset. Earlier work such as Datasheets for Datasets, published in Communications of the ACM in December 2021, and MLCommons' Croissant format with its Responsible AI extension set similar expectations.

For a supplier, the lesson is simple. Provenance answers prepared in a recognized structure are easier for a buyer to review, while ad hoc answers invite rounds of follow-up questions.

How to produce provenance metadata for a licensed dataset#

Producing provenance metadata for a licensed dataset is mostly a matter of collecting answers the licensing process already generates, then recording them per dataset version.

Keep the metadata with the dataset card and the file manifest. Together they let a buyer confirm that the records delivered match the records described.

  • Start from the system inventory: each source system, its owner, date range and export method.
  • Record the collection context: which notices, terms or consents applied when the records were created.
  • Log every preparation step with the method, tool and reviewer, using the standard privacy-enhancing tool terms.
  • Attach the rights review: contracts checked, carve-outs made and who owns any protected material.
  • Fill the Use elements from the signed license: permitted purposes, geography limits and confidentiality class.
  • Version the metadata with the dataset, so a refreshed delivery gets a new record instead of an edited old one.

Illustrative: an engineering firm answers a provenance request#

Illustrative: a fictional mechanical and electrical engineering firm is preparing project records, RFIs, submittal reviews and internal design review comments from Deltek and Procore for license. The buyer's diligence team asks for metadata aligned with the Data Provenance Standards.

The firm's operations lead fills Source and Provenance from the data inventory: systems, project years, export methods and preparation lineage. For Use, the general counsel records the confidentiality class, where employee notices and client contract reviews are kept, the privacy-enhancing tools applied, storage limited to the United States, the license itself, and training and evaluation as the permitted purposes.

The copyright element changed the scope. Drawings and specifications delivered to clients were marked as client-controlled and excluded, while the firm's internal review comments and RFI histories stayed in. The metadata shipped with the dataset card and the manifest.

How SourceX maps provenance to the Evidence Packet#

SourceX maps provenance to the SourceX Evidence Packet, which records provenance, licensing rights, permitted use, the privacy record and release authorization for each dataset. Those sections cover the questions the Source, Provenance and Use groups ask, and release authorization adds the supplier's sign-off.

The inputs come from the SourceX five-step transaction: Supply documents systems and history, Rights documents contracts and carve-outs, Preparation documents privacy methods, and Approval and Delivery document what was released, to whom and when.

Frequently asked questions

Are the Data Provenance Standards mandatory?

No. They are a voluntary industry specification. Buyers may ask for metadata in their structure, and using it can make diligence smoother, but no law requires a supplier to adopt them. A supplier can also map the same answers onto a buyer's own questionnaire.

Who fills in the metadata, the supplier or the buyer?

Mainly the supplier, because it holds the facts about origin, collection and rights. The buyer may add details about its own intended use, and both sides should agree on the final record before delivery. Keep the agreed version with the release record.

How do the standards relate to C2PA content credentials?

They address different problems. C2PA defines content credentials, cryptographically bound records of an individual media asset's provenance. The Data Provenance Standards describe a whole dataset's source, provenance and permitted use. A dataset of images might use both; a dataset of tickets or emails mainly needs the second.

Does provenance metadata expose confidential information?

It can, so treat it as part of the licensed materials. Describe systems and methods without including credentials, internal hostnames or customer names, and share the metadata under the same confidentiality terms as the data unless you choose to publish a summary.

Do we need new metadata for every delivery?

Yes, for every dataset version. A refresh with new records, a change in preparation method or a new permitted use changes the answers. Version the metadata with the dataset and keep earlier versions with their release records.

Sources

  • The Data Provenance Standards version 1.0.0 specification defines dataset metadata in three groups, Source, Provenance and Use, and says this metadata is needed to enable proper dataset selection for AI model training; the GitHub version lists six Source, nine Provenance and twelve Use elements. Source
  • The Use group includes confidentiality classification, consent documentation location, privacy-enhancing technologies applied, processing and storage geographies, license to use, intended data use, and copyright, patent and trademark status; Annex 7.2 lists privacy-enhancing tool codes including redaction, masking, pseudonymization, tokenization, k-anonymity, l-diversity, t-closeness and differential privacy. Source
  • The specification states that the Data & Trust Alliance was established in September 2020 as a not-for-profit consortium. Source
  • The Data Provenance Standards are released under Creative Commons CC-BY-SA-4.0, with code and supporting artifacts under Apache 2.0. Source
  • The Data Provenance Initiative released an audit covering 44 data collections spanning more than 1,800 fine-tuning text-to-text datasets, documenting sources, licenses and creators, and its tools generate a Data Provenance Card for any filtered subset. Source
  • Datasheets for Datasets by Gebru and co-authors was published in Communications of the ACM, vol. 64, no. 12 (December 2021). Source
  • MLCommons' Croissant is a metadata standard for machine-learning datasets built on schema.org's Dataset vocabulary, with a Responsible AI extension. Source
  • The C2PA Explainer defines a Content Credential, or C2PA Manifest, as a cryptographically bound structure that records an asset's provenance. Source
  • California AB 2013, signed September 28, 2024, requires developers of generative AI systems to post training-data documentation stating, among other things, whether datasets were purchased or licensed and whether they include copyrighted material or personal information. Source
  • Article 53(1)(d) of the EU AI Act requires providers of general-purpose AI models to publish a sufficiently detailed summary of training content following an AI Office template, which the European Commission published on July 24, 2025. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify