Provenance, rights and permitted use
Documenting Opt-Out Checks: An Evidence Log for TDM Reservations
Quick answer
To document TDM opt-out compliance, keep an append-only log with one row per asset per check: the URL or asset ID, which signal you looked for (robots.txt, TDMRep, HTTP header, HTML meta, embedded credentials, terms of use), the raw value returned, the fetch timestamp, the crawler and parser versions, the decision, and the exclusion ID that removed the asset. Tie every row to a dataset version so you can prove what a given training run actually contained.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why an opt-out log has to record effort, not just outcomes
An auditor or court asks what you checked and how, not only what you kept, so the log must show the scanning effort itself. Commentary on the Hamburg decision in Kneschke v. LAION reads the court as expecting developers to document their efforts to scan for rights reservations [1], and the case background shows that whether a reservation was expressed in machine-readable form became a central question, even though the judgment itself rested on the scientific-research exception [2]. A final corpus manifest proves only what survived; it cannot prove that a reserved work was looked for and dropped.
For providers of general-purpose AI models in the EU, Article 53(1)(c) of the AI Act requires a policy to comply with Union copyright law, including identifying and complying with reservations of rights under Article 4(3) of the DSM Directive, using state-of-the-art technologies [3]. The GPAI Code of Practice copyright chapter turns that into commitments: one up-to-date copyright policy, crawlers that read and follow robots.txt and other appropriate machine-readable protocols, and a point of contact for complaints [4]. The Code is voluntary and signing it does not by itself equal compliance [5]. An evidence log is how you demonstrate that the policy ran.
The legal rule itself is covered in EU text and data mining opt-outs under DSM Article 4; this page is the record design.
Which opt-out signals the log must cover
Log every signal type your copyright policy says you honor, because an unlogged signal is indistinguishable from an ignored one. As of October 2026, the practical set for web-published works is:
- robots.txt under RFC 9309: record the user-agent group that matched, the specific rule (Allow or Disallow path), and the HTTP status of the robots.txt fetch [7]. RFC 9309 treats an unreachable file (5xx or network error) as a full disallow and a 4xx as no restriction, and says crawlers should not rely on a cached copy for more than 24 hours, so status code and cache age are evidence, not noise [7].
- TDMRep: the W3C community group protocol expresses
tdm-reservation(0 or 1) and an optionaltdm-policyURL through/.well-known/tdmrep.json, HTTP response headers, or HTML<meta>elements [8]. Log which location produced the value and whether locations disagreed. - Embedded content credentials: images, audio and video may carry a training-and-data-mining assertion in a C2PA manifest; that assertion is now defined by the Creator Assertions Working Group as
cawg.training-miningrather than as a core C2PA label [9]. Log whether a manifest was present, whether its signature validated, and the assertion value. - Natural-language reservations in site terms or imprint pages: whether these count as machine-readable was a live question in Kneschke [1][2]. Log the URL, a hash of the captured text, and who classified it.
- Rightsholder notices and complaints received after collection, which the Code's complaint mechanism contemplates [4]. These are opt-outs too, just late ones.
US law has no statutory TDM opt-out mechanism; the Copyright Office's Part 3 report (still a pre-publication version as of October 2026) analyzes training under fair use instead [10]. Teams that train globally still log EU-style signals, because one corpus usually feeds models placed on the EU market.
Opt-out evidence log schema
The core of the log is a per-check record that a reviewer can replay without re-crawling. Keep raw responses, not just parsed booleans, because parsers change and disputes turn on what the server actually returned.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
check_id | chk_2026-09-14_000418273 | Unique, immutable key for this check |
asset_id | sha256:9f2c...e1 (content hash) | Survives URL changes and dedup |
source_url | https://example-news.test/2026/03/report.html | What was requested |
signal_type | robots_txt / tdmrep_wellknown / tdmrep_header / tdmrep_meta / c2pa_assertion / tos_text / notice | Which mechanism was checked |
signal_location | https://example-news.test/robots.txt | Exact resource consulted |
http_status | 200 | Drives RFC 9309 error handling |
raw_value_ref | s3://evidence/raw/chk_...gz + sha256 | Replayable proof of the response |
parsed_value | Disallow: /2026/ for User-agent: ExampleBot | What the parser concluded |
fetched_at | 2026-09-14T08:12:44Z | When the signal was observed |
crawler_ua / tool_version | ExampleBot/3.2, optout-parser 1.8.0 | Reproducibility after upgrades |
decision | exclude / include / hold_for_review | Outcome of the policy rule |
rule_id | POL-CR-4.2 | Links decision to copyright policy clause |
exclusion_id | exc_77120 | Proves removal happened downstream |
dataset_version | webtext-v14.2 | Ties evidence to a specific training input |
reviewer | null or analyst ID | Human sign-off for ambiguous cases |
Two companion tables complete the design. An exclusions table maps exclusion_id to the asset hashes removed and the pipeline step that removed them (crawl-time skip, post-crawl filter, or retroactive purge). A dataset version table records the manifest hash of each corpus version and the time window of checks it relied on, which is the link an auditor follows from a model card back to raw evidence. The fields overlap with record-level provenance and the training data use register; reuse their identifiers instead of inventing parallel ones.
Rules that make the log credible under scrutiny
A log persuades only if it is complete, tamper-evident and reproducible, so write those properties down as engineering rules. The following practices are practitioner recommendations rather than requirements from any single regulation:
- Append-only storage. Write to object storage with object lock or a WORM retention mode, and chain daily batches with a hash of the previous batch. Corrections are new rows, never edits.
- Check before fetch, and log both. Record the robots.txt and TDMRep lookups that preceded each content request, including lookups that blocked it. A log that contains only fetched pages cannot show what was skipped.
- Record negatives. "No tdmrep.json (404), no header, no meta" is evidence. Store it explicitly rather than inferring it from a missing row.
- Version the parser. When
tool_versionchanges, re-run the new parser over stored raw values and log deltas. Parser bugs on wildcard or longest-match rules are a common silent failure. - Re-scan on refresh. Before each new dataset version, re-check signals for retained assets and record changes; a site that adds a reservation after your crawl creates a deletion decision, not a historical footnote.
- Make exclusions verifiable. Sample
exclusion_idvalues each release and confirm the hashes are absent from the shipped shards and from tokenized caches. - Retain for the model's life. Keep evidence at least as long as any model trained on that dataset version is on the market, since questions arrive after release.
The EU's public-summary template for Article 53(1)(d) asks providers to describe crawled sources and the measures taken to respect reservations [6]. Generate that narrative from the log's aggregates, not from memory, so the published summary and the internal evidence never diverge.
Failure modes that break the evidence chain
Most opt-out evidence fails at the joins between systems rather than in the scanner. Watch for these:
- Cached robots.txt older than 24 hours with no recorded fetch time, contrary to RFC 9309 guidance [7].
- User-agent mismatch: the crawler announced one token while the log evaluated rules for another, or a wildcard group was applied when a specific group existed.
- Redirect drift: robots.txt evaluated for the original host while content came from a redirected host with different rules.
- Dedup erasing evidence: near-duplicate removal kept a copy from a permissive mirror while the canonical source reserved rights. Log the provenance of the surviving copy.
- Third-party corpora with no logs: an open dataset or vendor delivery arrives as shards with no check history. Treat it as unscanned; see auditing an existing training corpus.
- Credentials stripped in preprocessing: image resizing or transcoding removed the C2PA manifest before the assertion was read [9].
What to require from suppliers of published content
When you buy rather than crawl, contract for the supplier's opt-out evidence in the same structure you keep yourself. Ask for the per-check log or an export of it, the copyright policy clause each decision maps to, parser and crawler versions, the collection window, and a commitment to pass on rightsholder notices received after delivery. Then test it: pick 50 to 100 asset IDs, fetch their current signals, and compare against the supplier's recorded values, as described in testing a supplier's provenance claims on a sample.
Opt-out logs are one part of a wider evidence pack; the AI training data audit readiness checklist shows where they sit next to licenses, de-identification records and use registers. For the contractual side of record keeping, see SourceX's guide to keeping a record of what you licensed, the opt-out glossary entry, and how SourceX approaches data governance. The broader framework lives in the provenance hub.
Licensed operational data avoids much of this problem because rights come from an agreement with the holder rather than from the absence of a reservation. SourceX does not source scraped web content; it sources operational datasets from US companies on request, and every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Teams that want that kind of data can describe what they need to SourceX.
Source licensed training data with documented rights
If your corpus needs data whose rights rest on a license rather than on opt-out scans, SourceX looks for US businesses that hold the data you describe and manages the licensing agreement. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and every release is approved by the supplying company. Tell SourceX what data you need.
Sources
- Kluwer Copyright Blog (Wolters Kluwer), "Kneschke vs LAION: landmark ruling on TDM exceptions for AI training data (Part 2)" (2024). https://legalblogs.wolterskluwer.com/copyright-blog/kneschke-vs-laion-landmark-ruling-on-tdm-exceptions-for-ai-training-data-part-2/
- Bristows (Inquisitive Minds), "First court decision on text and data mining copyright exceptions: Kneschke v LAION" (2024). https://inquisitiveminds.bristows.com/post/102jmd3/first-court-decision-on-text-and-data-mining-copyright-exceptions-kneschke-v-lai
- EUR-Lex, Publications Office of the European Union, "Regulation (EU) 2024/1689 (Artificial Intelligence Act)" (2024). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- European Commission (Shaping Europe's digital future), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
- European Commission, "Explanatory notice and template for the public summary of training content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- IETF / RFC Editor, "RFC 9309: Robots Exclusion Protocol" (2022). https://www.rfc-editor.org/rfc/rfc9309
- W3C TDM Reservation Protocol Community Group, "TDM Reservation Protocol (TDMRep)". https://www.w3.org/2022/tdmrep
- Coalition for Content Provenance and Authenticity (C2PA), "C2PA clarification to C2PA TDM assertions reference". https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.