Skip to content

Provenance, rights and permitted use

AI Usage Signals Compared: robots.txt, ai.txt, TDMRep, Content Credentials and IETF Preferences

Quick answer

robots.txt controls crawler access to URLs; ai.txt is a proposed, non-standard file for stating AI usage terms; TDMRep reserves EU text-and-data-mining rights per site or resource; Content Credentials carry a training-and-mining preference inside the file itself; and IETF AIPREF is drafting a shared vocabulary that rides on robots.txt and HTTP headers. As of October 2026, only robots.txt is a published IETF standard. A defensible policy checks all five, logs each at collection time, and lets the most restrictive signal win.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What each signal actually controls

Each mechanism answers a different question, which is why they conflict in practice. A 2024 survey of web content controls for generative AI groups them by carrier and scope: crawler directives at site level, policy files at site level, page-level meta tags and headers, and metadata embedded in individual files [1]. The distinction that matters most for a governance lead is whether the signal governs access (may you fetch this URL) or use (may you train on what you fetched).

  • robots.txt (RFC 9309). Standardized in September 2022, it tells crawlers which paths they may fetch, matched by user-agent token. The standard frames its rules as requests crawlers are expected to honor, not access authorization [1][2]. It says nothing about use once content is fetched, so AI-specific tokens such as GPTBot, CCBot or Google-Extended are vendor conventions layered on top, not part of the standard.
  • ai.txt. The name covers several unrelated proposals, including vendor-published file formats and individual proposals for a site-wide file placed at the site root or under /.well-known [1]. No version is a standard, and crawlers that read one may ignore another.
  • TDMRep. A W3C Community Group specification, not a W3C Recommendation, built around EU DSM Article 4. It expresses a tdm-reservation flag and an optional tdm-policy URL through a /.well-known/tdmrep.json file, HTTP response headers, HTML meta elements or EPUB metadata [4].
  • Content Credentials (C2PA with CAWG). C2PA has clarified that its core specification has no TDM assertion; preferences come through extensions [5]. The Creator Assertions Working Group defines a training-and-data-mining assertion with entries for AI training, generative AI training, AI inference and data mining, each set to allowed, not allowed or constrained [6].
  • IETF AIPREF. Following the IAB AI-CONTROL workshop [3], the AIPREF working group is drafting a vocabulary for how automated systems may use content (draft-ietf-aipref-vocab) and an attachment that adds a Content-Usage rule to robots.txt groups and HTTP headers (draft-ietf-aipref-attach). As of October 2026, both are Internet-Drafts, not RFCs, so their terms can still change.

For deeper treatment of single mechanisms, see reading C2PA manifests on licensed media, the CAWG do-not-train assertion and EU Article 4 opt-out checks.

ai.txt vs robots.txt: the practical difference

robots.txt is the only signal almost every crawler already parses, but it governs fetching rather than training use. That creates two failure modes. A publisher who wants search indexing but not model training must block a training-specific user agent and hope every relevant crawler identifies itself; a publisher who blocks all crawlers also blocks search. ai.txt-style files try to separate use from access, but because there is no single specification, a dataset builder cannot assume any given crawler read or honored one [1].

The IETF attachment draft resolves part of this by putting usage preferences inside robots.txt itself, with the rule that paths a crawler cannot fetch carry no usage preference at all. In other words, a disallow is still the strongest robots-layer signal, and Content-Usage refines what happens to content that is allowed. The IAB AI-CONTROL workshop that preceded the working group framed the same problem: extend a format crawlers already read instead of inventing a parallel file [3].

How granularity and persistence change what you can verify

Site-level signals describe a place at a moment; file-level signals travel with the asset. A robots.txt or tdmrep.json file can change the day after collection, and the record of what it said is lost unless the crawler captured it. Content Credentials, by contrast, stay attached to an image, video or document as it moves between platforms, which is why C2PA and CAWG assertions matter for media that reaches you through aggregators rather than direct crawls [5][6].

The weakness of embedded metadata is stripping. Many upload pipelines and image CDNs remove XMP and JUMBF boxes, so absence of a manifest proves nothing. Treat a present "not allowed" assertion as binding and a missing manifest as unknown, never as permission.

Signals also drift. The Consent in Crisis audit of domains feeding major training corpora found restrictions in robots.txt rising sharply within a single year, and frequent mismatches between robots.txt and the same site's terms of service [7]. A corpus assembled in 2023 may contain content whose owners now object, which is a reason to record the signal state at collection time and re-check before reuse.

Article 53(1)(c) of the AI Act requires general-purpose AI model providers to put in place a copyright policy that identifies and complies with Article 4(3) reservations, including through state-of-the-art technologies [8]. Those duties have applied since 2 August 2025, with AI Office enforcement powers for new models reported to apply from August 2026.

The GPAI Code of Practice Copyright chapter gives signatories a route to show compliance; it centers on following robots.txt as standardized in RFC 9309 and on other appropriate machine-readable protocols [9]. Which other protocols count is not settled, which is why a conservative policy honors TDMRep and recognized AI-preference tokens alongside robots.txt. Outside the EU, these signals generally have no statutory force on their own, but ignoring a clear, documented objection can still weigh in contract, unfair-practices or copyright disputes. Counsel should decide how your organization treats them.

Comparison table: carrier, granularity, weight and logging

The table below is a working summary for policy drafting, not a legal ranking. Adoption notes are qualitative because no reliable cross-signal census exists.

Illustrative example: invented to show structure; it does not describe an available dataset.

SignalCarrierGranularityStatus (Oct 2026)EU weightWhat to log at collection
robots.txt/robots.txtSite, path, user agentIETF Proposed Standard (RFC 9309) [1]Named in GPAI Code [9]Raw file, fetch timestamp, HTTP status, matched group and token
AIPREF Content-Usagerobots.txt rule, HTTP headerSite, path, responseInternet-Draft [3]Plausible "machine-readable means"Parsed vocabulary terms, draft version parsed against
ai.txtSite-root or /.well-known fileSiteCompeting non-standard proposals [1]UncertainRaw file, which spec your parser assumed
TDMReptdmrep.json, header, meta, EPUBSite, path, resourceW3C CG Report [4]Designed for Art. 4tdm-reservation value, tdm-policy URL, carrier
C2PA/CAWG assertionManifest in file (JUMBF) or sidecarIndividual assetExtension spec [5][6]Plausible "machine-readable means"Manifest hash, signer, assertion values, validation result
HTML meta (noai, noimageai)Page headPageInformal convention [1]UncertainTag value and page URL

Setting the organization's signal policy

A usable policy names the signals honored, the date each became mandatory, the precedence rule and the evidence kept. The most restrictive applicable signal should win: a CAWG "not allowed" for AI training overrides a permissive robots.txt, and a TDMRep reservation overrides the absence of any AI-specific token. Apply the policy to every acquired corpus, not just your own crawls, and require crawl vendors to deliver the per-URL signal log described in provenance records to demand from crawl vendors.

Illustrative example: invented to show structure; it does not describe an available dataset.

signal_policy:
  version: 2026-10
  honored_signals:
    - { id: robots_txt_rfc9309, mandatory_from: 2022-09-01 }
    - { id: aipref_content_usage, mandatory_from: 2026-01-01, parser_draft: draft-ietf-aipref-vocab-NN }
    - { id: tdmrep, mandatory_from: 2024-01-01 }
    - { id: cawg_training_mining, mandatory_from: 2025-01-01 }
    - { id: ai_txt, mode: record_and_review }
  precedence: most_restrictive_wins
  missing_signal: unknown_not_permission
  recheck_before_reuse_days: 180
per_record_log:
  url: "https://example.com/post/123"
  fetched_at: "2026-03-14T09:12:00Z"
  robots_group_matched: "User-agent: *"
  content_usage: "train-ai=n"
  tdm_reservation: 1
  tdm_policy_url: "https://example.com/tdm-policy.json"
  cawg_ai_training: null
  decision: excluded
  reason: "tdm_reservation=1 and content_usage train-ai=n"

The content_usage value above is illustrative syntax; pin your parser to a specific draft version, since the vocabulary terms may still change. Store the decision in your training data use register so downstream teams can see why records were dropped.

Checklist before a web-derived corpus enters training or a RAG index:

  • Raw robots.txt and tdmrep.json snapshots exist for every source domain, with timestamps.
  • Crawl user-agent strings are recorded, so you can show which group applied.
  • Content Credentials were validated, not just detected, and stripped-metadata rates are reported.
  • Conflicts were resolved by the documented precedence rule, with counts excluded per signal.
  • Retrieval use is evaluated separately, since inference and grounding preferences can differ from training preferences; see grounding license vs training license.

When opt-out signals are the wrong control

Opt-out signals only describe what a publisher objected to; they never grant a license. Content that carries no reservation is not cleared for training, and terms of service, paywalls and third-party material inside a page can still restrict use. For high-value or sensitive data, a negotiated license is the stronger control because it defines records, uses and term directly; the provenance hub and the glossary entry on opt-out cover where signals end and contracts begin.

SourceX does not source scraped web content. It sources operational datasets from US companies, where each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, so buyers comparing signal-based web data with licensed alternatives can describe the data they need. More buyer guides are in the AI data hub.

Request licensed data with documented rights

If opt-out signals leave too much uncertainty for a training or RAG use case, SourceX can look for US businesses that hold the operational data you describe and manage the licensing process. Every release is approved by the supplying company, and a request does not guarantee a match. Start at sourcex.si/buyers.

Sources

  1. arXiv, "A Survey of Web Content Control for Generative AI" (2024). https://arxiv.org/pdf/2404.02309
  2. IETF Datatracker (IAB AI-CONTROL workshop slides), "AI, Robots.txt (Jimenez, Arkko)" (2024). https://datatracker.ietf.org/doc/slides-aicontrolws-ai-robotstxt/
  3. IETF Datatracker, "IAB Workshop on AI-CONTROL (aicontrolws) materials" (2024). https://datatracker.ietf.org/group/aicontrolws/materials/
  4. W3C TDMRep Community Group, "TDM Reservation Protocol (TDMRep), Final Community Group Report". https://www.w3.org/community/reports/tdmrep/CG-FINAL-tdmrep-20220216/
  5. Coalition for Content Provenance and Authenticity (C2PA), "C2PA clarification to C2PA TDM assertions reference". https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
  6. IPTC Metawatch, "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
  7. arXiv (Longpre et al.; NeurIPS 2024 Datasets and Benchmarks), "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
  8. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  9. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data