Skip to content

Provenance, rights and permitted use

C2PA Content Credentials in Training Data: Reading Provenance Manifests on Licensed Media

Quick answer

C2PA Content Credentials are signed manifests embedded in, or linked from, an image, video or audio file that record who produced it, with what tool, and which edits followed [1]. In a training dataset they are useful provenance evidence: a valid signature tells you the claims are unaltered since signing and which certificate signed them. They do not prove licensing rights, do not prove human origin, and their absence proves nothing. Extract them before preprocessing, validate them, record the results per file, and treat identity data as personal data.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What a C2PA manifest actually contains

A manifest is a signed bundle of assertions about one asset, plus a claim that binds those assertions to the file's bytes [1]. For dataset work, four parts matter.

  • Assertions. Typed statements such as c2pa.actions (created, opened, edited, cropped, converted), thumbnails, ingredient references and, in many files, creator or AI-related metadata [1][4]. The actions assertion can carry a digital source type, for example a value indicating the asset came from a trained algorithmic model rather than a camera capture [4].
  • Ingredients. References to parent assets, each of which may have its own manifest. A composited stock image or an edited clip can carry a chain of manifests back to the original capture.
  • Hard binding. A hash over the asset content (for JPEG and PNG, a data hash excluding the manifest box; for MP4 and other ISO BMFF files, a BMFF hash) that ties the manifest to these exact bytes.
  • Claim signature. A COSE signature using an X.509 certificate. The validator checks the signature, the certificate chain and, optionally, whether the signer appears on a trust list you choose.

Manifests are usually embedded (in JPEG as JUMBF boxes in APP11 segments; in MP4 as a dedicated box), but they can also sit in a .c2pa sidecar or at a remote URL referenced from the file. Your extraction pipeline has to handle all three, or you will record false "no manifest" results.

Why Content Credentials are provenance evidence, not rights evidence

A validated manifest proves who signed a statement about a file and that the file has not changed since; it does not prove the signer owned the rights or licensed them to your supplier. A camera manufacturer's signature on a capture tells you the device that produced the pixels, not who holds copyright, whether the people in frame consented, or what the photographer agreed with an agency. Rights still come from contracts, releases and the supplier's chain of title, covered in chain of title for AI training data.

The same caution applies to AI-origin questions. A manifest that says "created by camera" is a signed claim from that device's certificate holder; a manifest that says "generated by model X" is a claim from that tool. Neither replaces content inspection, and detection of AI-generated media is a separate problem. Many files in a typical licensed archive will have no manifest at all, because the capture device, editing software or distribution platform never wrote one or stripped it.

Absence of a manifest proves nothing. Many CMSs, CDNs, social platforms and image libraries re-encode files and drop metadata boxes. Record "no manifest found" as a neutral state, never as a negative signal about origin.

Training permission flags are an extension, not core C2PA

The core C2PA specification does not define a standard training-and-data-mining assertion; C2PA has clarified that such assertions are provided by third parties extending the spec [2]. The main one in circulation is the Creator Assertions Working Group's cawg.training-mining assertion, which states for categories such as AI training, generative AI training, AI inference and data mining whether use is allowed, not allowed or constrained [3][1]. Older files and tools may still use the earlier c2pa.training-mining label, so parsers should read both.

Treat these flags as a rights-holder preference signal to log and honor, not as a license. How they interact with robots.txt, TDMRep and other signals is covered in AI usage preference signals compared, and the assertion itself in the CAWG training-and-data-mining assertion. Metadata-based controls only work when they survive distribution and when downstream users read them, a limit noted across the research literature on content controls [6].

As of October 2026, for EU-facing general-purpose model providers, Article 53(1)(c) requires a copyright policy that identifies and complies with rights reservations under Article 4(3) of the DSM Directive, including through state-of-the-art technologies [7]. Whether a given manifest flag counts as such a reservation is a question for counsel, but a pipeline that discards the flag cannot show it was considered.

How to extract and validate manifests at dataset scale

Run extraction on the files exactly as delivered, before any resize, transcode, crop or format conversion. Practical steps:

  1. Hash first. Compute SHA-256 of each delivered file and store it as the provenance key. Any later manifest result must reference this hash.
  2. Extract with a maintained SDK. The open-source c2patool CLI and the c2pa-rs Rust library (with Python and JavaScript bindings) read embedded, sidecar and remote manifests and emit the manifest store as JSON plus a validation report. Pin the SDK version, because validation codes change between spec versions.
  3. Separate three outcomes. Record whether the manifest is well-formed, whether the hard binding still matches the bytes, and whether the signing certificate chains to a trust list you configured. A file can be well-formed with a broken binding (edited after signing) or validly signed by an untrusted certificate.
  4. Walk ingredients. Extract the full ingredient tree and record depth. A training-mining flag or source type on a parent ingredient may matter even when the active manifest is silent.
  5. Capture flags in normalized fields. Map cawg.training-mining and legacy labels to one column per category, preserving the raw JSON.
  6. Persist before you strip. Most training pipelines decode to tensors and discard metadata. Store the extracted manifest JSON and the validation report in a sidecar table keyed by the original file hash, then strip if you must.

For video and audio, check each delivered rendition separately. A manifest on the mezzanine MP4 does not carry over to the frames you extract or the WAV you derive, so record the derivation in your own lineage, as described in record-level provenance.

Illustrative per-file provenance record

The output of extraction should be a row per delivered file that a reviewer can read without re-running the SDK.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
file_sha2569f2c…e41aJoins manifest evidence to the exact delivered bytes
media_typeimage/jpegDetermines binding type and extraction path
manifest_locationembedded / sidecar / remote / noneSeparates "absent" from "not looked for"
manifest_wellformedtrueParse succeeded
binding_validfalseBytes changed after signing; treat claims as historical
signer_cn / signer_trustedExample Camera Co. / trueWho signed, and whether your trust list accepts them
digital_source_typedigitalCaptureSigned claim about origin, not proof
ingredient_depth2How many parent assets are referenced
tdm_ai_training / tdm_genai_trainingnotAllowed / notAllowedPreference flags to honor or escalate
identity_assertion_presenttrueTriggers privacy handling
sdk_version / validated_atc2patool x.y.z / 2026-10-09T14:02ZReproducibility of the result

Store this alongside the dataset's use register, such as the one described in the AI training data register, so permitted-use decisions and manifest evidence sit together.

Identity assertions carry personal data

Identity assertions can name the creator and link verified accounts or credentials, which makes them personal data in most privacy regimes [5]. A dataset that preserves full manifests may therefore hold names, social handles or organizational identities of photographers and editors that never appear in the pixels.

Decide per dataset whether you need identity fields at all. Common practice is to keep a hashed signer identifier and the trust-list result for provenance, and move named identity content into an access-restricted store governed like any other personal data; see privacy and de-identification for AI training data. Manifests can also carry location metadata and thumbnails, so include those fields in your privacy review rather than assuming the manifest is purely technical.

Using manifests in supplier diligence

Content Credentials are most useful when you ask suppliers to preserve them and to explain gaps. Before acceptance, ask the supplier for:

  • The share of files delivered with manifests, and the reason others lack them (capture devices, platform stripping, supplier pipeline).
  • Confirmation that files were not re-encoded between their archive and your delivery, or a list of transforms applied.
  • Any training-mining or do-not-train flags they found and how they handled those files.
  • Whether they removed identity assertions, and if so, a record of the removal method.

Then test a sample yourself rather than accepting the summary, following testing a supplier's provenance claims on a sample. A delivery where every manifest has a broken binding, or where flags marked "notAllowed" appear in the training split, is a red flag worth escalating. NIST's generative AI profile treats content provenance as part of managing generative AI risk, which supports documenting these checks in your risk records [8]. For the broader framework, start at the provenance buyer's guide or the data provenance glossary entry.

If you need licensed operational media rather than web-sourced files, describe the data you need to SourceX. SourceX sources on request from US businesses that hold the described data, does not source scraped web content or generic CCTV or photos, and rights-reviews each dataset for ownership and consents.

Sourcing licensed media with documented provenance

SourceX finds US companies that hold the media and operational data you describe, assesses data and licensing permissions, and delivers under a license that defines records, uses, term and delivery once the supplier agrees. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Start a request on the SourceX buyers page.

Frequently asked questions

Does stripping metadata during preprocessing break provenance?

It breaks the in-file evidence, not your record of it. If you hashed the delivered file and stored the extracted manifest and validation report first, the provenance trail survives in your lineage store even though training tensors carry no metadata.

Can a valid manifest be on an AI-generated image?

Yes. A generative tool can sign a manifest that truthfully declares algorithmic origin, and the signature will validate. Validity speaks to integrity and signer identity, not to whether the content was captured by a camera.

Should files with a "notAllowed" training flag be excluded automatically?

Most teams quarantine them pending review rather than silently training on them. Whether a license from the rights holder overrides a flag written earlier, or by a different party, is a contractual and legal question to settle with counsel.

Sources

  1. C2PA / Creator Assertions Working Group, "Training and Data Mining Assertion (Version 1.0)" (2024). https://github.com/decentralized-identity/cawg-training-and-data-mining-assertion/blob/v1.0/docs/modules/ROOT/pages/index.adoc
  2. Coalition for Content Provenance and Authenticity (C2PA), "C2PA clarification to C2PA TDM assertions reference" (2024). https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
  3. IPTC (metawatch registry), "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
  4. Content Authenticity Initiative open-source documentation, "Actions assertions". https://opensource.contentauthenticity.org/docs/manifest/assertions-actions
  5. C2PA / Creator Assertions Working Group, "CAWG Identity Assertion (Version 1.0)" (2024). https://c2pa.org/specifications/specifications/1.0/specs/CAWG_Identity_Assertion.html
  6. arXiv, "A Survey of Web Content Control for Generative AI" (2024). https://arxiv.org/pdf/2404.02309
  7. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  8. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data