Skip to content

Regulation and governance for data buyers

Text and Data Mining Exceptions by Country: Where AI Training Still Needs a License

Quick answer

Text and data mining (TDM) exceptions let you copy lawfully accessed works for analysis in some countries, but none is a general license to train. As of October 2026, the EU allows commercial TDM unless rights are reserved, the UK limits its exception to non-commercial research, Japan and Singapore allow broad analysis with conditions, the US relies on case-by-case fair use, and China requires lawful sources. No exception grants access to non-public data, so private operational datasets still need a license.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

The short comparison: six regimes, one common limit

Every regime below starts from lawful access, and none creates a right to obtain data you cannot already reach. That single condition explains why exceptions matter for public web corpora and barely matter for the support tickets, engineering logs or contract archives that sit behind a company's firewall. Use the table to see which question to ask first in each jurisdiction, then read the country sections for the failure modes.

JurisdictionProvisionWho can rely on itCommercial training?Key conditionsWhere it stops
EUDSM Directive (EU) 2019/790, Art. 3Research organizations, cultural heritage institutionsResearch purpose onlyLawful access; scientific researchCommercial labs are outside Art. 3 [1]
EUDSM Directive, Art. 4Anyone with lawful accessYesNo express, appropriate reservation (machine-readable for online content); keep copies only as long as neededReserved works; contracts and paywalls [1][2]
UKCDPA 1988 s29APerson with lawful accessNoNon-commercial research; acknowledgementAny commercial purpose [5][6]
JapanCopyright Act Art. 30-4AnyoneYes, for non-enjoyment usesPurpose must not be enjoyment of the expressionUnreasonable prejudice, e.g. databases sold for analysis [9]
SingaporeCopyright Act 2021 s244Anyone meeting conditionsYesLawful access and other conditionsUnlawfully accessed copies [10]
US17 U.S.C. 107 fair use (no TDM statute)Decided per caseCase by caseFour factorsMarket substitution, pirated sources [11][12]
ChinaGenerative AI Interim MeasuresService providersRegulated, not exemptedLawful sources; no IP infringementUnlawful corpora [15]

For each regime's detail beyond this page, SourceX keeps country summaries at the laws index, including Japan and the United Kingdom.

EU: Article 3 versus Article 4 of the DSM Directive

In the EU, a commercial lab relies on Article 4, not Article 3, and Article 4 fails wherever the rightsholder has reserved its rights. Article 3 covers TDM for scientific research by research organizations and cultural heritage institutions, which is why a university spin-out and its commercial parent can sit on opposite sides of the line [1]. Article 4 covers anyone with lawful access, but only for works whose rightsholders have not expressly reserved the reproduction right in an appropriate manner, which for content made publicly available online means machine-readable means [1].

The AI Act turns that reservation into an operational duty. Article 53(1)(c) requires general-purpose AI model providers to adopt a copyright policy that identifies and complies with Article 4(3) reservations, including through state-of-the-art technologies [2]. The GPAI Code of Practice, published 10 July 2025 as a voluntary tool, gives signatories a way to show this [3]. Its copyright chapter commits signatories to honor machine-readable reservations such as robots.txt when crawling, and it states that it does not affect licensing agreements with rightsholders [4].

Two failure modes recur in EU diligence:

Licensed data sidesteps both problems because the permission comes from the contract, not from the absence of a reservation. How that fits a written copyright policy is covered in the GPAI Code copyright chapter guide.

UK: section 29A stays non-commercial in 2026

In the UK, commercial AI training cannot rely on a TDM exception, because CDPA s29A covers only computational analysis for non-commercial research by a person with lawful access [5]. IPO guidance from when the exception was introduced says contract research for an outside company is unlikely to qualify, and that analysis with a purpose that is not solely non-commercial is very likely infringing [6]. That catches the common pattern of an academic team running analysis that feeds a commercial model.

The government published its statutory Report on Copyright and Artificial Intelligence on 18 March 2026 under the Data (Use and Access) Act 2025 [7]. Commentators read it as dropping the proposed commercial TDM exception with an opt-out as the preferred option, without legislating a replacement [8]. As of October 2026, s29A remains the operative text. The detail is on UK commercial AI training and copyright.

Japan: Article 30-4 and the enjoyment test

Japan's Article 30-4 is the broadest exception in this set, but it does not reach uses aimed at enjoying the expression or uses that unreasonably prejudice rightsholders. The Agency for Cultural Affairs' clarification, as reported by commentators, treats pure model training as non-enjoyment use while flagging that the exception does not apply where a purpose of enjoyment coexists, such as training intended to reproduce the creative expression of specific works [9]. The proviso on unreasonable prejudice is the key buyer issue: commentary on that guidance cites databases compiled and sold for analysis as the kind of work whose market the exception should not displace [9].

Read that proviso as a direct signal that curated, commercially offered datasets are expected to be licensed even in Japan. RAG pipelines that retrieve and surface source text raise the enjoyment question more sharply than pre-training. See Japan Article 30-4 and licensed data.

Singapore: computational data analysis under section 244

Singapore permits computational data analysis, including use of a work to improve a program's performance, under section 244 of the Copyright Act 2021, subject to conditions led by lawful access [10]. Commentary describes the conditions as covering how the copy was obtained and limits on further communication of the copy, so a corpus assembled by circumventing a paywall or from a known infringing source falls outside it [10]. Conditions and their interaction with contracts are covered in Singapore's computational data analysis exception.

United States: fair use, not a statute

The US has no statutory TDM exception, so training on copyrighted works turns on a four-factor fair use analysis decided case by case [12]. The Copyright Office's Part 3 report on generative AI training, still a pre-publication version as of October 2026, analyzes training under those factors, discusses market effects and licensing, and is guidance rather than law [11]. What it means for licensing strategy is covered in the Copyright Office report and licensing.

Recent case law shows how fact-specific this is. On 29 September 2026 the Third Circuit affirmed that ROSS Intelligence's copying of Westlaw headnotes to build a competing, non-generative legal research tool was not fair use [13]. In Bartz v. Anthropic, a class settlement received final approval in July 2026 [14]. Acquisition route matters on its own, as lawful access and pirated sources explains.

China: lawful-source duties instead of an exception

China does not offer a TDM exception for AI training; it regulates training data through the Interim Measures for generative AI services, in force since 15 August 2023, which require lawful data sources and prohibit infringing intellectual property [15]. TC260's basic security requirements add operational controls on corpus sources, including source vetting and records [15]. Treat China as a provenance-documentation regime: the question regulators ask is whether you can show where each corpus came from and that you had the right to use it.

Why exceptions do not reach private operational data

Exceptions permit copying of works you already lawfully access; they never compel a company to hand over its internal records. CRM exports, Zendesk or ServiceNow ticket histories, Jira and Git logs, scanned invoices, and recorded hands-on work are not publicly reachable, so the only route to them is an agreement with the holder. That agreement also carries terms no exception supplies: permitted uses (pre-training, fine-tuning, evaluation, RAG), retention and deletion, personal-data handling, and audit rights.

Exceptions also leave other law untouched. GDPR, HIPAA, state biometric statutes and trade-secret law apply regardless of copyright status, and contract terms such as API or website terms can bind independently. Retention questions are addressed in training data retention requirements.

Jurisdiction routing record for a training corpus

A per-source routing record lets counsel see, for every corpus component, which legal basis it relies on in each market where the model will be trained or placed.

Illustrative example: invented to show structure; it does not describe an available dataset.

source_id: corp-0042
description: "Field-service work orders, 2019-2024, PDF and CSV"
access_route: license            # license | public_web | research_partner
public_availability: none        # none | partial | open_web
legal_basis_by_market:
  EU: license                    # Art. 4 not relied on; no reservation check needed
  UK: license                    # s29A unavailable for commercial purpose
  JP: license                    # Art. 30-4 not relied on; proviso risk avoided
  SG: license
  US: license                    # no fair use argument required
  CN: license_with_source_records
reservation_snapshot: n/a        # required when basis is EU Art. 4
personal_data: removed_or_replaced; method_recorded: true
permitted_uses: [fine_tuning, evaluation, rag]
evidence: [executed_license.pdf, deidentification_report.pdf, data_card.md]

Rows whose basis is an exception need extra fields: crawl date, reservation snapshot (EU), purpose statement (UK, Japan), and access method (Singapore, US).

Counsel checklist before relying on an exception

Use this before any corpus component goes into a training run on the strength of an exception rather than a license.

  1. Do you have lawful access to the material, through public availability or a subscription whose terms permit this use, without circumventing a paywall, login or technical measure? If not, stop: you need a license.
  2. Which markets will the model be trained in and placed on? Apply the strictest relevant rule, usually the UK for commercial use and the EU for reservations.
  3. For EU reliance, do you hold a dated record of robots.txt, TDMRep and terms checks at crawl time?
  4. For Japan, could the use involve enjoyment of the expression, or displace a market for data sold for analysis?
  5. For the US, is the source free of pirated copies, and is the model's output market distinct from the source's?
  6. Does any contract (API terms, platform terms, NDA) restrict the use independently of copyright?
  7. Does the material contain personal data that needs a separate legal basis or de-identification?

How SourceX fits licensed training data across jurisdictions

When the answer to question 1 is "not public," licensing is the only route, and that is where SourceX works. SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows, and manages the licensing agreement and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery; describe what you need on the SourceX buyers page. For the wider regulatory map, start at the compliance hub or the AI data hub.

Licensed training data when no exception applies

If your corpus plan depends on private records that no TDM exception can reach, SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company. Personal details are removed or replaced before delivery, with the method recorded. Describe the data you need to license.

Sources

  1. Hannes Snellman, "Text and data mining for AI training". https://www.hannessnellman.com/news-and-views/blog/text-and-data-mining-for-ai-training/
  2. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  3. European Commission (AI Office), "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
  4. European Commission (reproduced by Edition Multimedia), "Code of Practice for General-Purpose AI Models: Copyright Chapter (10 July 2025)" (2025). https://www.editionmultimedia.fr/wp-content/uploads/2025/07/Code-of-Practice-for-GPAI-Copyright-10-07-25.pdf
  5. UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 - Consolidated (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  6. UK Intellectual Property Office, "Exceptions to copyright: Research" (2014). https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/375954/Research.pdf
  7. UK Government (GOV.UK), "Report on Copyright and Artificial Intelligence" (2026). https://www.gov.uk/government/publications/report-and-impact-assessment-on-copyright-and-artificial-intelligence/report-on-copyright-and-artificial-intelligence
  8. Reed Smith, "UK copyright and AI report: the opt-out is dead, but what comes next?" (2026). https://www.reedsmith.com/articles/uk-copyright-and-ai-report-the-opt-out-is-dead-but-what-comes-next/
  9. Hugh Stephens Blog, "Japan's Text and Data Mining (TDM) Copyright Exception for AI Training: A Needed and Welcome Clarification from the Responsible Agency" (2024). https://hughstephensblog.net/2024/03/10/japans-text-and-data-mining-tdm-copyright-exception-for-ai-training-a-needed-and-welcome-clarification-from-the-responsible-agency/
  10. Rouse, "Artificial Intelligence in Singapore: Copyright Infringement Defence for Artificial Intelligence and Machine Learning" (2024). https://rouse.com/insights/news/2024/artificial-intelligence-in-singapore-copyright-infringement-defence-for-artificial-intelligence-machine-learning
  11. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  12. Jenner & Block, "US Copyright Office Releases Pre-Publication Version of Report on Copyright Issues in Generative AI Training" (2025). https://www.jenner.com/en/news-insights/client-alerts/us-copyright-office-releases-pre-publication-version-of-report-on-copyright-issues-in-generative-ai-training
  13. IPWatchdog, "Third Circuit Affirms Revised Fair Use Ruling Against ROSS' AI Legal Research Platform in Sealed Opinion" (2026). https://ipwatchdog.com/2026/09/30/third-circuit-affirms-revised-fair-use-ruling-against-ross-ai-legal-research-platform-in-sealed-opinion/
  14. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  15. SESEC, "China Issued the Technical Document on Basic Security Requirements for Generative Artificial Intelligence Services" (2024). https://sesec.eu/2024/04/30/china-issued-the-technical-document-on-basic-security-requirements-for-generative-artificial-intelligence-services/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data