Provenance, rights and permitted use
Do-Not-Train Flags in Media Metadata: The CAWG Training-and-Data-Mining Assertion
Quick answer
The "do not train" flag in Content Credentials is the cawg.training-mining assertion, a Creator Assertions Working Group (CAWG) extension carried inside a C2PA manifest [1]. It holds separate entries for data mining, AI training, generative AI training and AI inference, each set to allowed, constrained or notAllowed [1][2]. The flag does not technically block use, but it is evidence that a reservation was discoverable [3]. A defensible ingestion filter excludes notAllowed, routes constrained to legal review, and logs every decision.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Where the training-and-mining assertion lives in a manifest
The assertion is one labeled entry in the assertion store of a signed C2PA manifest, so you read it only after you parse and validate that manifest. Content Credentials are embedded in the asset (a JUMBF box in JPEG, PNG, MP4, WAV and other supported containers) or referenced from a remote manifest store. The manifest holds a claim, a signature and assertions such as c2pa.actions, ingredient references and, when the creator set one, cawg.training-mining [1][2].
The label is a CAWG extension, not a core C2PA label [1]. Older tooling and manifests written against C2PA 1.x specifications may carry the earlier c2pa.training-mining label with c2pa.-prefixed entry keys, so a robust parser matches both the cawg. and legacy prefixes and records which one it saw. For the general mechanics of reading manifests, ingredients and validation states, see reading C2PA Content Credentials in training data; this page covers only the permission vocabulary.
In practice you extract manifests with the open-source C2PA SDKs (c2pa-rs, c2pa-python, c2pa-node) or the c2patool command-line utility. Each returns a JSON report listing manifests, the active_manifest identifier, assertions by label, and a validation status array. Your filter should read the training-mining assertion from the active manifest first, then walk ingredient manifests.
The four permission keys and what each one covers
The assertion separates four uses so a creator can, for example, allow search indexing while refusing generative training [1][2]. The CAI SDK documentation describes keys for data mining, machine-learning training, generative AI training and inference [2]. Treat each key independently; one allowed value never implies the others.
| Entry key | Use it covers | Typical ingestion consequence |
|---|---|---|
cawg.data_mining | Automated analysis, indexing, extraction of patterns or statistics | Governs dedup indexing, captioning pipelines and corpus analytics |
cawg.ai_training | Training or fine-tuning any ML model, including classifiers and CV models | Governs detection, segmentation and embedding models |
cawg.ai_generative_training | Training models that generate content similar to the input | Governs diffusion, video-generation, TTS and music models |
cawg.ai_inference | Using the asset as model input at inference time (for example RAG or conditioning) | Governs retrieval corpora and reference-image prompting |
Each entry carries a use value. When use is constrained, the entry can include a free-text constraint_info field that states the conditions, for example a contact address or a pointer to license terms [1][2]. That text is not machine-actionable, which is why constrained needs a person.
Values: allowed, constrained and notAllowed
The three values form a closed vocabulary, and your parser should reject anything else as malformed rather than guessing [1]. allowed states the creator does not object to that use. notAllowed states a reservation. constrained states the use is permitted only under conditions described elsewhere.
Absence is the value most pipelines get wrong. A missing assertion, a missing entry key or a stripped manifest says nothing about permission; it is not allowed. Many social platforms and image CDNs strip embedded metadata on upload, so a file without credentials may simply have lost them. Your rights basis for such files must come from the license or chain of title, not from silence.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"label": "cawg.training-mining",
"data": {
"entries": {
"cawg.data_mining": { "use": "allowed" },
"cawg.ai_inference": { "use": "allowed" },
"cawg.ai_training": {
"use": "constrained",
"constraint_info": "Non-generative training permitted under license ref LIC-0042; contact rights desk"
},
"cawg.ai_generative_training": { "use": "notAllowed" }
}
}
}
In this record, a team training an object detector would route the asset to legal review (constrained, with a license reference to check), and a team training an image generator would exclude it outright.
A decision table for ingestion filters
The safest default maps notAllowed to exclude, constrained to review, and only an explicit allowed from a validly signed manifest to the automated path. This is a policy choice, not a rule written in the specification; adjust it with counsel. It reflects a simple asymmetry: a wrongly excluded asset costs a little recall, while a wrongly included reserved asset can contaminate a checkpoint you cannot cheaply retrain.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Manifest state | Entry for your use | Signature / validation | Filter action | Log fields |
|---|---|---|---|---|
| Present | notAllowed | Any | Exclude, even if license says otherwise, until counsel resolves | asset hash, manifest ID, signer, entry key, value |
| Present | constrained | Valid | Hold for legal review; read constraint_info | plus constraint_info text, reviewer, decision |
| Present | allowed | Valid, trusted signer | Ingest under the governing license | plus signer cert chain, trust list version |
| Present | allowed | Invalid or untrusted | Treat as absent | plus validation status codes |
| Present | Key missing | Any | Treat as absent for that use | entry keys present |
| Absent or stripped | n/a | n/a | Rely on license and chain of title only | source batch, license reference |
| Conflicting (active vs ingredient) | notAllowed anywhere in lineage | Any | Exclude pending review | full ingredient path |
Two rows deserve emphasis. An allowed flag in a manifest whose signature fails validation proves nothing, because anyone can write JSON. And a reservation on an ingredient (for example, a stock photo composited into a derived image) should propagate to the derivative unless your license expressly covers the ingredient. For the broader problem of embedded third-party works, see third-party content inside licensed corpora.
Why a declarative flag still matters legally
The flag is declarative and cannot stop copying, but it creates a verifiable record that a developer could have identified the reserved rights [3]. That evidentiary weight is the reason to honor it even where you believe another legal basis applies.
In the EU, Article 53(1)(c) of the AI Act requires providers of general-purpose AI models to maintain a copyright policy that identifies and complies with rights reservations under Article 4(3) of the DSM Directive, including through state-of-the-art technologies [4]. As of October 2026, those obligations have applied since 2 August 2025, with AI Office enforcement for new models reported to apply from August 2026. The GPAI Code of Practice Copyright chapter describes how signatories can show compliance through a maintained copyright policy [5]. Whether a given metadata flag counts as an effective machine-readable reservation is a question for counsel; see EU text and data mining opt-outs under DSM Article 4 and text and data mining exceptions by country.
Creator tools have made the flag easier to set. Reporting in 2024 described Adobe's effort to let artists attach a do-not-train preference to their Content Credentials, while noting it only works if AI developers choose to respect it [6].
How the flag fits with other opt-out signals
The CAWG assertion is asset-level and travels with the file, while robots.txt, ai.txt and TDMRep operate at site or path level and are lost the moment a file is copied off the origin [7]. That makes the assertion the most useful signal for licensed media deliveries, where you rarely see the original hosting context. A survey of web content control mechanisms for generative AI catalogs these approaches and their limits, including reliance on voluntary compliance [7].
Other embedded fields can conflict with it. IPTC photo metadata has its own data-mining property, and XMP packets may carry rights statements written by a DAM. When signals disagree, apply the most restrictive one for the relevant use and record the conflict. The full comparison is in AI usage signals compared; the SourceX glossary defines an opt-out.
Carrying the decision into your training data register
The filter's output should become a per-asset permitted-use record, not a one-time pass/fail. Store the assertion snapshot alongside the asset so a later audit can show what the flag said at ingestion time, even if the source re-signs or strips it.
Illustrative example: invented to show structure; it does not describe an available dataset.
Checklist for an ingestion job that honors training-and-mining flags:
- Parse every asset with a C2PA SDK; store the raw manifest JSON and validation status codes.
- Match labels
cawg.training-miningand legacy prefixed variants; record which matched. - Resolve the entry for your intended use (
ai_trainingfor a CV classifier,ai_generative_trainingfor a generator,ai_inferencefor retrieval). - Apply the decision table; never treat absence or an unsigned
allowedas permission. - Walk ingredient manifests and propagate any
notAllowed. - Queue
constraineditems with theirconstraint_infotext for counsel. - Write asset hash, manifest ID, signer, value, action and policy version to your register.
- Re-run the filter when your intended use changes; a generative fine-tune is a new use.
Machine-readable dataset documentation such as Croissant-RAI can carry usage conditions at dataset level, which lets downstream teams see the filter policy that produced a shard [8]. Map the per-asset result into the fields described in encoding permitted uses per record and track it in your AI training data use register. If you are auditing an existing corpus rather than a new delivery, the training corpus provenance audit covers retroactive checks.
Failure modes seen in media ingestion pipelines
Most failures come from preprocessing that destroys or ignores the manifest before the filter runs. Transcoding with ffmpeg, resizing with Pillow or re-encoding to WebP usually drops the JUMBF box, so extract credentials from the original bytes, before any transform, and key results by content hash.
Other common failures:
- Packing before parsing. Writing assets into tar shards before reading manifests loses the link between file and flag; extract first, then pack (see WebDataset tar shards).
- Single-key logic. Checking only
ai_generative_trainingand ignoringai_traininglets reserved assets into "non-generative" CV models that are later reused for generation. - Frame-level video. Sampling frames from a credentialed MP4 produces images with no credentials; propagate the parent decision to every derived frame and clip. Video carries other rights layers too, covered in rights layers in a video clip.
- Trust lists. Accepting any signer certificate turns the flag into an unauthenticated claim; pin the trust list version you validated against.
Licensing media where permitted use is documented
When you license media rather than collect it, the flag is one input; the license should state permitted uses directly. SourceX sources operational datasets from US companies on request, including new recordings of hands-on work, and every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Nothing is held in stock and a request does not guarantee a match; you can describe the media you need to SourceX.
For the wider framework, start at the data provenance buyer's guide or the AI data hub.
Request licensed media with documented permitted uses
SourceX looks for US businesses that hold the data you describe, assesses data and licensing permissions, and agrees allowed uses in a license before anything is transacted. Every release is approved by the supplying company. Tell SourceX what media and uses you need.
Sources
- IPTC (Metawatch), "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
- Content Authenticity Initiative (open-source SDK docs), "Assertions and actions (manifest documentation)". https://opensource.contentauthenticity.org/docs/manifest/assertions-actions
- ECIJA, "Cinco programas para protegerse de los entrenadores de IA". https://www.ecija.com/en/news-and-insights/cinco-programas-para-protegerse-de-los-entrenadores-de-ia/
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- MIT Technology Review, "Adobe wants to make it easier for artists to blacklist their work from AI scraping" (2024). https://www.technologyreview.com/2024/10/08/1105234/adobe-wants-to-make-it-easier-for-artists-to-blacklist-their-work-from-ai-scraping
- arXiv, "A Survey of Web Content Control for Generative AI" (2024). https://arxiv.org/pdf/2404.02309
- arXiv (Jain et al., MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.