Provenance, rights and permitted use
Encoding Permitted Uses per Record: A Rights Metadata Schema for Training Data
Quick answer
Permitted-use metadata is a small set of structured fields attached to every training record, or to the license segment it belongs to, that states which uses are allowed (pre-training, fine-tuning, evaluation, retrieval), which are prohibited, until when, where and for whom. Use controlled vocabularies, link each record to a license ID and notice version, and make pipelines fail closed: a record without resolvable rights fields is excluded, not assumed permitted.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why dataset-level license fields are not enough
A single license string on a dataset card cannot represent the mixed permissions that real operational corpora carry. One delivery can combine tickets collected under three privacy-notice versions, records from customers whose contracts exclude AI training, and a subset licensed for evaluation only. When that nuance lives in a PDF, the training job sees none of it.
The public record shows how often dataset-level labels fail even in their simplest form. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission above 70% and license error rates above 50% on popular hosting sites [2]. If the top-level label is wrong that often, an engineer cannot rely on it to gate per-record decisions. The record-level provenance guide covers when record granularity is worth its cost; this page covers the fields themselves.
The intent matches established provenance frameworks. The Data & Trust Alliance standards treat intended use and restrictions as distinct metadata categories alongside source, lineage and legal rights [6], which the provenance standards explainer breaks down. For high-risk systems in the EU, Article 10 of the AI Act also expects training, validation and test sets to sit under documented data governance practices [8], and a queryable rights layer is one practical way to show that.
The core fields: a rights block every record should carry
The minimum viable rights block has about a dozen fields, most of which point to a license segment rather than repeating terms on every row. Keep terms normalized in a rights_policy table keyed by policy_id; store only the pointer, the notice reference and record-specific exclusion flags on the record itself. That keeps a 200-million-row corpus cheap to re-filter when a policy changes.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Type | Controlled values or format | Purpose |
|---|---|---|---|
record_id | string | stable, supplier-scoped ID | Join key for audits and deletions |
policy_id | string | ID in rights_policy table | Points to the license segment that governs this record |
license_id | string | contract or schedule reference | Traces back to the executed agreement |
permitted_uses | array of enum | pretrain, finetune_sft, finetune_pref, eval_only, rag_index, synthetic_seed | What the pipeline may do |
prohibited_uses | array of enum | same vocabulary plus redistribute, public_release, model_output_resale | Explicit denials that override defaults |
valid_from / valid_until | ISO 8601 date | 2026-01-01 | Term window for new training runs |
territory | array | ISO 3166-1 alpha-2 codes or WORLD | Where processing is allowed |
licensee_scope | enum | licensee_only, licensee_and_affiliates, named_entities | Who may use the record |
notice_version | string | supplier notice ID plus effective date | Links to the terms in force at collection |
exclusion_flags | array of enum | opt_out, deletion_request, legal_hold, contract_excluded | Record-level overrides |
deid_method | enum | safe_harbor, expert_determination, pseudonymized, none | Privacy treatment applied |
rights_asserted_by | string | attestation document ID | Who vouched for the rights |
The notice_version field deserves special care, because a record collected in 2022 is governed by the 2022 notice, not today's. The notice-version matching guide explains how to derive it from collection timestamps. For contractual meaning of terms such as field of use, see the field of use glossary entry.
Designing the use vocabulary so filters stay unambiguous
Every permitted-use value must map to one pipeline stage and one license phrase, or the filter becomes a guess. "Training" alone is too coarse: many licenses treat supervised fine-tuning, preference tuning, evaluation and retrieval indexing differently, and a record admissible for an eval harness may be barred from gradient updates.
Three rules keep the vocabulary stable:
- Hierarchy, not synonyms. Define
finetuneas a parent offinetune_sftandfinetune_pref, and decide in writing whether a grant of the parent implies the children. Do not letsft,instruction_tuningandchat_tuningcoexist as peers. - Prohibitions win. If a value appears in both
permitted_usesandprohibited_uses, the record is treated as prohibited for that use. ODRL lets a policy declare a conflict strategy for this case; your schema should pick one and document it. - Version the vocabulary. Store
vocab_versionon the policy so a later split (for example, separatingrag_indexfromrag_cache) does not silently reinterpret old grants.
Evaluation deserves its own value. An eval_only record that leaks into a fine-tuning mix both breaches terms and contaminates the benchmark, so tag eval holdouts at ingestion and block them in the training loader by default.
Expressing the same rules in ODRL for interchange
ODRL is the W3C policy language for exactly this job: a policy holds permissions, prohibitions and duties over an asset, narrowed by constraints such as time, place or purpose. It is a good interchange format between a supplier's rights system and a buyer's catalog, even if your pipeline reads a flattened table at run time.
ODRL is already used for AI-relevant rights. The TDMRep vocabulary, published by a W3C Community Group rather than as a W3C Standard, defines its text-and-data-mining policies as an ODRL profile [1]. On the dataset-documentation side, MLCommons Croissant describes datasets in schema.org-based JSON-LD, and its Croissant-RAI extension makes responsible-AI documentation machine-readable [3][4], so a rights block can travel next to that description. Because core ODRL has no pretrain or finetune action, define them in a small profile of your own rather than overloading the generic use action.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"@context": "http://www.w3.org/ns/odrl.jsonld",
"@type": "Agreement",
"uid": "urn:policy:lic-0042-seg-b",
"profile": "https://example.org/odrl-ai-uses/v1",
"assigner": "urn:party:supplier-a",
"assignee": "urn:party:licensee-x",
"target": "urn:dataset:support-tickets-2021-2025:seg-b",
"permission": [{
"action": ["aiu:finetune_sft", "aiu:eval_only"],
"constraint": [
{ "leftOperand": "dateTime", "operator": "lteq", "rightOperand": { "@value": "2027-12-31", "@type": "xsd:date" } },
{ "leftOperand": "spatial", "operator": "isAnyOf", "rightOperand": ["US", "CA", "GB"] }
]
}],
"prohibition": [{
"action": ["aiu:pretrain", "distribute", "aiu:rag_index"]
}]
}
At ingestion, a resolver expands this policy into the flat permitted_uses, prohibited_uses, valid_until and territory columns, writes the ODRL uid into policy_id, and stores the source JSON-LD unchanged for audit. Use W3C PROV-O relations such as prov:wasDerivedFrom to connect derived records, for example transcripts or summaries, to the source record whose policy they inherit [5].
Fail-closed filtering in the training and retrieval pipeline
The filter should exclude any record whose rights cannot be resolved, rather than treat missing fields as permissive. Silent defaults are the most common failure mode: a join on policy_id returns null, the loader's WHERE clause evaluates to unknown, and the record slips through, or a new shard arrives without the rights sidecar and is ingested anyway.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Condition at job time | Decision | Logged reason |
|---|---|---|
policy_id missing or not found | Exclude | rights_unresolved |
Requested use not in permitted_uses | Exclude | use_not_granted |
Requested use in prohibited_uses | Exclude | use_prohibited |
Job start date after valid_until | Exclude | term_expired |
Compute region not in territory | Exclude | territory_mismatch |
Job owner outside licensee_scope | Exclude | scope_mismatch |
Any exclusion_flags present | Exclude | flag name |
| All checks pass | Include | granted:<policy_id> |
Implement the checks as a single function that returns a decision and a reason, call it in the dataloader for training and in the indexer for RAG, and write per-run counts by reason to the run manifest. Validate the rights block itself at delivery with a schema contract; the JSON Schema delivery validation guide shows how to make policy_id and notice_version required and enum-checked so gaps are caught before ingestion. Re-run the filter when a deletion request or policy change arrives, and record which model checkpoints consumed records that are now excluded.
Retrieval needs one extra control. A RAG index built under a valid license can keep serving chunks after the term ends, so store policy_id in the vector store's metadata and apply the same decision function at query time, not only at index build.
Carrying signals that arrive embedded in media
Media files can bring their own use signals, and the schema should ingest them rather than ignore them. Content credentials can carry a training-and-data-mining assertion that states whether AI training, generative training, inference and data mining are allowed; it originated in the C2PA specification and, as of October 2026, is maintained by the Creator Assertions Working Group as cawg.training-mining [7]; the CAWG training-and-data-mining assertion guide covers how those assertions are evolving.
Map an embedded "not allowed" to a record-level exclusion_flags value such as embedded_tdm_notallowed, and never let an embedded "allowed" widen what the license grants. The license is the ceiling; embedded signals can only narrow it.
What to ask suppliers for so the schema can be populated
Buyers can only fill these fields if suppliers deliver the inputs in machine-readable form. Ask for a rights sidecar file per shard (Parquet or JSONL keyed by record_id), a policy table or ODRL documents for every license segment, the notice versions and their effective dates, a list of excluded customers or contracts expressed as record IDs, and a named attestation for who asserts the rights. The data rights attestation template shows how to frame that last item.
Keep the dataset-level view too. A training data use register records, per dataset and license, what each corpus may be used for, while the record-level schema enforces it inside jobs. For a plain-language view of what licensed buyers may do with data, see what buyers are allowed to do with licensed data and AI data license terms explained.
When you source operational data through SourceX for buyers, every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, allowed uses, term and delivery, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Those are the inputs a schema like this one needs. The wider provenance hub and the AI data licensing guide cover the surrounding diligence.
Source operational data with permitted uses you can encode
SourceX sources operational datasets from US companies on request and manages licensing, with every release approved by the supplying company and allowed uses set in a license. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Describe the data and the uses you need at sourcex.si/buyers.
Sources
- W3C, "TDMRep vocabulary". https://www.w3.org/ns/tdmrep
- Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Akhtar et al., MLCommons (arXiv), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- Jain et al., MLCommons (arXiv), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- W3C, "PROV-O: The PROV Ontology". https://w3.org/TR/prov-o/
- VKTR, "AI Transparency: Unpacking the New Data Provenance Standards". https://www.vktr.com/digital-marketing/ai-transparency-unpacking-the-new-data-provenance-standards/
- C2PA (Linux Foundation project), "C2PA clarification to C2PA TDM assertions reference" (22 January 2026). https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
- EUR-Lex, "Regulation (EU) 2024/1689 of the European Parliament and of the Council (Consolidated)". https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02024R1689-20260727
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.