Provenance, rights and permitted use
EU Text and Data Mining Opt-Outs (DSM Article 4): What AI Data Buyers Must Check
Quick answer
Article 4 of the EU DSM Directive lets anyone text-and-data-mine lawfully accessed works, including for commercial AI training, unless the rightholder has reserved that use "in an appropriate manner," such as machine-readable means for content published online [1][2]. For buyers, the opt-out question attaches to every work inside an acquired corpus: you need evidence of which reservation signals were checked, when, by whom, and what was excluded, because general-purpose model providers must identify and honor these reservations under AI Act Article 53(1)(c) [5].
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What Article 4 actually permits and reserves
Article 4 is a conditional exception: it covers reproductions and extractions of lawfully accessible works for text and data mining, and copies may be kept as long as necessary for that purpose [1]. Article 4(3) makes the exception unavailable where the rightholder has expressly reserved the use in an appropriate manner, with machine-readable means named as the example for content made publicly available online [1][2]. Unlike Article 3, which serves research organizations and cultural heritage institutions, Article 4 is the exception commercial AI developers rely on, so a valid reservation removes the Article 4 basis and leaves a license from the rightholder as the main route.
Three conditions therefore decide whether a work in your corpus was minable under Article 4:
- Lawful access. Content obtained by bypassing paywalls or logins, or taken from infringing sources, fails before any opt-out analysis begins.
- No valid reservation at the relevant time. The reservation must be "express" and "appropriate" for the medium.
- Purpose limited to mining. Retention and later use beyond what the mining requires need their own basis.
For a country-by-country view of where this exception differs or does not exist, see text and data mining exceptions by country. This page stays on the operational test for Article 4 reservations in data you acquire; the opt-out glossary entry defines the general term.
What counts as a machine-readable reservation in 2026
As of October 2026, no single signal is settled as sufficient, and national courts read "machine-readable" differently [3]. The safest buyer position is to treat every recognized signal as potentially valid and require suppliers to show they checked all of them.
The signals that matter in practice:
- robots.txt directives for named AI crawlers (for example, rules targeting CCBot or other training user-agents). Common Crawl says it follows robots.txt and avoids paywalled and login-protected content, which makes the crawl date and robots state at that date central evidence for any web-derived corpus [10].
- TDMRep, the W3C Community Group protocol, which lets a site declare a
tdm-reservationvalue and atdm-policypointer to licensing terms, via a/.well-known/tdmrep.jsonfile, HTTP headers or HTML meta tags [7]. It is a Community Group specification, not a W3C Standard [7]. - Embedded media metadata, such as do-not-train assertions in C2PA manifests or the CAWG training-and-data-mining assertion; see do-not-train flags in media metadata and reading C2PA Content Credentials.
- Natural-language terms of service stating that mining or AI training is prohibited. Whether these qualify is the contested edge case.
The GPAI Code of Practice copyright chapter gives signatories concrete commitments, including honoring robots.txt and other appropriate machine-readable protocols when crawling and not circumventing access controls [6]. The Code is voluntary, but it is the most concrete published benchmark for what "state-of-the-art" reservation compliance can look like [5][6].
Kneschke v. LAION and the divergence buyers must plan around
The Hamburg litigation over the LAION-5B image-text dataset is the most-cited case on Article 4 reservations, and it does not end the debate. The first-instance court decided the case on the scientific-research exception, so its Article 4 remarks were not the basis of the outcome. Commentary on that decision ties the Article 4 analysis directly to the AI Act duty to identify reservations and stresses that providers should document their opt-out scanning [4].
The December 2025 appeal ruling, as analyzed on the Kluwer Copyright Blog, held that a reservation must be interpretable by machines, not merely detectable by them, and commentators note that other national courts take different views [3]. That gap matters for buyers: a natural-language clause in website terms may be treated as valid in one member state and ineffective in another.
An EU Commission consultation launched in December 2025 to identify state-of-the-art, technically implementable opt-out protocols, closing in January 2026, confirms the question remains open as of October 2026 [8]. Do not build procurement terms that assume one reading wins. Plan for the strictest plausible standard on any corpus you intend to train on in the EU.
Opt-outs added after collection
Reservations added after a crawl or collection date are the hardest practical issue, and there is no settled rule as of October 2026. A conservative working approach is to record the collection date per item, treat reservations present at collection as binding, and re-check before each new training run because a later reservation may affect fresh mining even if earlier copies were lawful.
Three facts change the analysis:
- Retention. Article 4(2) allows copies only as long as necessary for mining purposes [1]. A corpus held indefinitely and re-mined for each model generation is a weaker position than a single documented run.
- New mining events. Each new training or filtering pass on retained copies may be a separate act; check reservations again at that point.
- Evidence of state at collection. Without archived robots.txt files, TDMRep responses and HTTP headers captured at crawl time, you cannot show a reservation was absent.
Build this into an evidence log for TDM reservation checks rather than relying on a supplier's narrative.
Why buyers carry the risk under the AI Act
The AI Act assigns opt-out compliance to the model provider, not the data vendor. Article 53(1)(c) requires general-purpose AI model providers to maintain a copyright policy that identifies and complies with Article 4(3) reservations, including through state-of-the-art technologies [5].
Article 53(1)(d) also requires a public summary of training content, using the template the Commission published on 24 July 2025 [9]. If you cannot describe how acquired data was screened for reservations, that summary and your copyright policy will rest on assumptions. The companion guide on training data summaries buyers need from suppliers covers the summary fields; this page covers the opt-out evidence behind them.
Supplier questionnaire for Article 4 opt-out evidence
Send the questions below before any license is signed and require document answers, not yes/no responses. The decision column reflects a conservative buyer posture.
Illustrative example: invented to show structure; it does not describe an available dataset.
| # | Question to supplier | Acceptable evidence | Red flag | Buyer decision if red flag |
|---|---|---|---|---|
| 1 | How was each item accessed, and was access lawful? | Access method per source (license, API terms, public page without login) | "Scraped," paywall bypass, shadow-library origin | Exclude source |
| 2 | Which reservation signals were checked? | robots.txt, TDMRep, HTTP headers, HTML meta, embedded metadata, site terms | robots.txt only | Re-screen or exclude |
| 3 | When were they checked? | Per-item collected_at and optout_checked_at timestamps | Single corpus-level date | Treat as unverified |
| 4 | Were raw signal responses archived? | Stored robots.txt, tdmrep.json, headers with hashes | Summary report only | Request archives |
| 5 | How were natural-language terms handled? | Documented rule (excluded, reviewed, or ignored) | "Not machine-readable, so ignored" | Apply strictest reading for EU training |
| 6 | Were later reservations re-checked? | Re-check log before delivery | None since collection | Re-screen before training |
| 7 | What was removed, and can removal be audited? | Exclusion list with domains, URLs, item IDs | No exclusion list | Hold acceptance |
| 8 | Does the corpus contain embedded third-party works? | Policy for quoted text, attachments, stock media | Unknown | Scope a review |
An illustrative per-item record a supplier might deliver alongside the corpus:
{
"item_id": "doc-000184",
"source_url": "https://example.org/articles/1234",
"access_basis": "public_page_no_login",
"collected_at": "2025-11-03T14:22:10Z",
"optout_checked_at": "2025-11-03T14:22:09Z",
"signals": {
"robots_txt": {"status": "allow", "archive_sha256": "…"},
"tdmrep": {"tdm_reservation": 0, "source": "/.well-known/tdmrep.json"},
"http_headers": "none",
"embedded_metadata": "none",
"site_terms_ai_clause": "not_found"
},
"decision": "include",
"recheck_due_before": "next_training_run"
}
Contract terms that make opt-out evidence enforceable
Your license should turn the questionnaire into obligations. Ask for a representation describing the reservation screening method and its date range, delivery of the per-item evidence, a duty to notify you of reservations discovered later, and a cooperation obligation to remove affected items. Pair this with a chain-of-title review for licensed sources and a scoped check for third-party content inside licensed corpora, since embedded works carry their own rightholders.
Opt-out screening applies most directly to content published online. Datasets licensed directly from the organization that created them, such as internal support tickets or engineering records, rest primarily on that licensor's ownership and consents rather than on web signals, but embedded third-party material still needs review. For the broader framework, start at the provenance and permitted-use hub.
How SourceX approaches rights for buyers training in the EU
SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents and finance and legal workflows, and does not source scraped web content. Every dataset is rights-reviewed for ownership and consents, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which you can test against the questionnaire above. AI teams anywhere can describe the data they need on the buyer page; requests do not guarantee a match.
Sourcing licensed training data with documented rights
SourceX serves AI teams wherever they are based and manages the commercial process from finding a supplier through assessing data and licensing permissions to an agreed license defining records, uses, term and delivery. Nothing is contracted until a supplier agrees, and every release is approved by the supplying company. Describe the dataset you need at sourcex.si/buyers.
Sources
- Law Society of Ireland Gazette, "Clash of the titans" (2024). https://www.lawsociety.ie/gazette/in-depth/2024/december/clash-of-the-titans/
- Hannes Snellman, "Text and data mining for AI training". https://www.hannessnellman.com/news-and-views/blog/text-and-data-mining-for-ai-training/
- Kluwer Copyright Blog (Wolters Kluwer), "LAION Round 2: Machine-Readable but Still Not Actionable: The Lack of Progress on TDM Opt-Outs (Part 1)" (2026). https://legalblogs.wolterskluwer.com/copyright-blog/laion-round-2-machine-readable-but-still-not-actionable-the-lack-of-progress-on-tdm-opt-outs-part-1/
- Kluwer Copyright Blog (Wolters Kluwer), "Kneschke vs. LAION: Landmark Ruling on TDM Exceptions for AI Training Data (Part 2)". https://legalblogs.wolterskluwer.com/copyright-blog/kneschke-vs-laion-landmark-ruling-on-tdm-exceptions-for-ai-training-data-part-2/
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- W3C TDM Reservation Protocol Community Group, "TDM Reservation Protocol (TDMRep), Final Community Group Report". https://www.w3.org/community/reports/tdmrep/CG-FINAL-tdmrep-20220216/
- European Commission, Digital Strategy, "Commission launches consultation on protocols for reserving rights from text and data mining under the AI Act and the GPAI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/consultations/commission-launches-consultation-protocols-reserving-rights-text-and-data-mining-under-ai-act-and
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- Common Crawl Foundation, "Common Crawl's Submission to the UK's Copyright and AI Consultation". https://commoncrawl.org/uk-copyright-and-ai-consultation
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.