Regulation and governance for data buyers
Japan Article 30-4 and AI Training: When a Data License Is Still Required
Quick answer
Article 30-4 of Japan's Copyright Act lets you exploit works for AI training without permission when the purpose is not to enjoy the thoughts or sentiments they express [2]. The Agency for Cultural Affairs' 2024 guidance confirms that ordinary training usually qualifies [1][3]. The exception has two hard edges. It fails where an enjoyment purpose coexists with analysis, such as targeted style imitation or RAG output, and its proviso excludes uses that unreasonably prejudice rights holders, including copying databases offered for analysis [2][4][5]. In those cases you still need a license.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What Article 30-4 permits and why it is broader than EU or UK exceptions
Article 30-4 permits exploitation of a work, to the extent considered necessary, where the purpose is not to enjoy, or let others enjoy, the thoughts or sentiments expressed in it, and it lists data analysis among its examples [2]. Unlike the EU's DSM Article 4, it has no machine-readable opt-out mechanism, and unlike UK CDPA s29A it is not limited to non-commercial research. That breadth is why Japan is often described as permissive for training, and why the limits matter more than the headline.
The statute is short; the working interpretation comes from the Agency for Cultural Affairs (ACA). The Legal Subcommittee of the Copyright Subdivision of the Council for Cultural Affairs adopted the "Approach to AI and Copyright" in March 2024 after a January 2024 draft and public comment, and the ACA's "General Understanding on AI and Copyright in Japan" is its English summary [1][7][8]. The ACA describes it as the subcommittee's view of how the current Act should be read as of publication, not binding law [1]. Courts are not bound by it, so treat it as the best available reading, not a safe harbor.
As of October 2026, the ACA's English copyright page still presents the 2024 General Understanding [1]. A secondhand report of a May 2026 update could not be confirmed in ACA material, so check the ACA site before relying on any reading here.
For a side-by-side view of how Japan compares with the EU, UK, Singapore and others, see text and data mining exceptions by country.
The non-enjoyment test: where training stops qualifying
The non-enjoyment test fails whenever a purpose to enjoy the expression sits alongside the analytical purpose, even if analysis is also present [3]. The guidance separates the development and training stage from the generation and use stage, and asks whether someone intends to perceive or reproduce the creative expression of particular works [3][8]. A mixed purpose defeats the exception for the copies made toward that purpose.
The final guidance and the ACA's July 2024 checklist identify recurring failure patterns [4][7]:
- Output-targeted training. Intentional over-learning, meaning training or additional training meant to make a model output content that shares creative expression with specific works, such as overfitting on one illustrator's catalog [4].
- Style-imitating fine-tunes. Small adapters, such as LoRA, built from a narrow set of works by one creator to reproduce that creator's expression. Style as an abstract idea is not protected, but a fine-tune designed to reproduce protected expression can lose the exemption [7].
- Retrieval for output. Building a database for retrieval-augmented generation so that a model can output the creative expression of the indexed works [4][7].
Large-scale pre-training on diverse text to learn language patterns sits at the safest end. Every step toward a narrow, identifiable set of works and toward output that reflects their expression moves you toward needing permission.
The unreasonable-prejudice proviso and databases offered for analysis
The proviso removes the exception where the exploitation would unreasonably prejudice the copyright owner's interests in light of the nature or purpose of the work or the circumstances of its exploitation [2]. Commentary on the guidance reports a concrete example that matters to buyers: reproducing a database of works that is offered for information analysis, including AI training [5]. If a licensing market for analytical use of that content exists, copying it under Article 30-4 instead of buying the license undercuts that market.
Two related signals point the same way. Commentary reports that the guidance treats circumvention of technical measures that block collection, such as access controls on a paid corpus, as a factor pointing to unreasonable prejudice [5]. It also reports that a robots.txt block or similar measure can support an inference that a database is, or will be, offered for analysis [5]. A short checklist for counsel:
- Is the corpus marketed or licensed for text and data mining, model training or analytics?
- Did you need to bypass a login, paywall, API rate limit or crawler block to collect it?
- Does the holder publish terms that reserve analytical use or offer an API product for it?
- Would your copy substitute for a purchase the holder actually sells?
A "yes" to any of these means Article 30-4 is a weak basis, and a license is the more defensible route. Databases sold for analytical use are the example the proviso is usually read to cover. Buyers who decide to license data rather than rely on the exception can request licensed operational datasets from SourceX.
RAG, Article 47-5 and the generation stage
RAG is the use case where Article 30-4 is weakest, because retrieval exists to put a work's expression in front of a user [4][7]. Article 47-5 separately allows minor exploitation of works incidental to computer-based services that search for works or provide the results of information analysis, to the extent necessary and subject to its own unreasonable-prejudice limit [2]. That provision is narrow: a RAG system that returns long passages, full documents or close paraphrases is unlikely to fit either article.
At the generation stage, ordinary infringement analysis applies: similarity plus reliance on the original. The guidance indicates that if a work was in the training data, reliance is generally presumed for a similar output, even where the user did not know the work [3]. For buyers, this means the dataset record matters after training: you need to know what was in the corpus to answer an output claim.
Practical consequences for retrieval corpora:
- License documents you index for retrieval, even if you could argue the training copies were non-enjoyment uses.
- Keep the license terms aligned with the retrieval pattern: snippet display, full-text display, summarization and citation behave differently.
- For operational corpora such as support tickets linked to knowledge articles, see support ticket retrieval datasets for how the data is usually structured.
Acquisition source, personal data and cross-border exposure
Article 30-4 answers a copyright question only, and it does not cure problems from how data was obtained or what it contains. The January 2024 draft prompted debate over training on pirated material; commentary on that draft reports that knowingly collecting from infringing sources weighs against the developer, especially where outputs could reproduce the pirated works [6]. Confirm the final wording in the ACA text, and see lawful access and pirated sources for how acquisition changes risk.
Personal data in a corpus remains subject to the Act on the Protection of Personal Information whether or not copyright is cleared. See the Japan APPI page and the Japan country overview for the data protection side.
Copyright is territorial. Copies made in Japan may rely on Article 30-4, but a model trained there and placed on the EU market still faces the AI Act's copyright policy and opt-out obligations, covered in EU AI Act Article 53 obligations. US exposure turns on fair use, discussed in license or rely on fair use.
Decision table: Article 30-4 or a license
The deciding variables are purpose, output behavior, the holder's market and how data was collected. Use the table below as a first-pass triage before counsel review.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Scenario | Purpose risk | Proviso risk | Working conclusion |
|---|---|---|---|
| Pre-training on broad, lawfully accessed public text | Low: pattern learning | Low unless corpus is sold for analysis | Article 30-4 is a reasonable basis; document it |
| Fine-tuning on one artist's or author's catalog to match their expression | High: enjoyment purpose | Medium | License, or redesign the dataset |
| LoRA adapter built from a narrow set of works for style output | High | Medium | License |
| Indexing paid news or journal archives for RAG answers | High: output reaches readers | High: analytical market exists | License with retrieval terms |
| Copying a commercial dataset marketed for AI training | Low to medium | High: database proviso | License |
| Collecting behind a login or API limit by bypassing controls | Varies | High: technical-measure factor | License or stop |
| Training on private operational records from a company | Not the main issue | Not the main issue | Contract and privacy govern; license required anyway |
The last row is often overlooked. Proprietary business records, such as support histories or engineering logs, are not publicly accessible, so the practical question is never Article 30-4; it is whether the holder grants the rights by contract.
Records to keep when you rely on Article 30-4
If you rely on the exception, document why each corpus qualifies so you can defend the position later. A minimal per-corpus record:
Illustrative example: invented to show structure; it does not describe an available dataset.
corpus_id: jp-web-text-2026-03
legal_basis: "Japan Copyright Act Art. 30-4 (non-enjoyment)"
purpose_statement: "General language pre-training; no per-author targeting"
output_controls: ["n-gram regurgitation filter", "memorization eval on held-out sample"]
proviso_review:
sold_for_analysis: false
technical_measures_bypassed: false
robots_or_terms_reservation_checked: true
acquisition: "Lawful access; known infringing domains excluded"
personal_data: "APPI review completed; identifiers filtered"
rag_use: false
guidance_version_checked: "ACA General Understanding (2024); checked October 2026"
reviewer: "counsel initials and date"
Keep this alongside your AI training data audit readiness evidence, and recheck the ACA copyright page for updates before each new corpus [1].
Licensed data for training that Article 30-4 does not cover
SourceX sources operational datasets from US companies on request, rights-reviews each one for ownership and consents, and delivers it under a license that defines records, uses, term and delivery. Nothing is contracted until the supplying company agrees, and a request does not guarantee a match; see the compliance hub and AI data hub for related guides. If your Japan analysis lands in the "license" rows above, describe the data you need at SourceX for buyers.
Sources
- Agency for Cultural Affairs, Government of Japan, "Copyright (English policy page, including General Understanding on AI and Copyright in Japan – Overview)" (2024). https://www.bunka.go.jp/english/policy/copyright/index.html
- Ministry of Justice, Japan (Japanese Law Translation), "Copyright Act (Act No. 48 of 1970), English translation". https://www.japaneselawtranslation.go.jp/en/laws/view/3379
- WIPO (presentation by Tatsuhiro Ueno), "General Understanding on AI and Copyright in Japan". https://www.wipo.int/documents/d/office-japan/docs-en-tatsuhiro-ueno_general-understanding-on-ai-and-copyright-in-japan_set.pdf
- Mondaq, "Japanese Government Published Checklist and Guidance Related to AI and Copyrights" (2024). https://www.mondaq.com/japanese-government-published-checklist-and-guidance-related-to-ai-and-copyrights/1510346
- Hugh Stephens Blog, "Japan's Text and Data Mining (TDM) Copyright Exception for AI Training: A Needed and Welcome Clarification from the Responsible Agency" (2024). https://hughstephensblog.net/2024/03/10/japans-text-and-data-mining-tdm-copyright-exception-for-ai-training-a-needed-and-welcome-clarification-from-the-responsible-agency/
- Privacy World, "Japan's New Draft Guidelines on AI and Copyright: Is It Really OK to Train AI Using Pirated Materials?" (2024). https://www.privacyworld.blog/2024/03/japans-new-draft-guidelines-on-ai-and-copyright-is-it-really-ok-to-train-ai-using-pirated-materials/
- International Bar Association, "Japan's emerging framework for responsible AI: legislation, guidelines and guidance". https://www.ibanet.org/japan-emerging-framework-ai-legislation-guidelines
- Copyright Agency (Australia), "Japan Copyright Office on AI and Copyright" (2024). https://www.copyright.com.au/?p=26441
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.