Regulation and governance for data buyers
Lawful Access and Pirated Sources: How the Way You Acquire Training Data Changes Copyright Risk
Quick answer
How you obtained training data is now a legal fact in its own right. In the EU, UK and Singapore, text and data mining exceptions require lawful access, and the EU's GPAI Code of Practice commits signatories to lawfully accessible content without circumventing technical measures [1][5][7]. Japan's guidance also weighed pirated sources [8]. In the US, the Copyright Office and a 2025 district-court ruling treated pirated copies differently from purchased or licensed ones [10][11]. The defense is a per-dataset acquisition record.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why acquisition channel became a copyright variable
Acquisition channel matters because most exceptions and fair-use arguments analyze the copy you made, and the lawfulness of the source travels with that copy. A model trained on a corpus that mixes licensed archives with a shadow-library torrent does not average out the risk: the unlawful slice can carry its own claim, its own class, and its own damages exposure [11][12].
For a data acquisition lead, this shifts the question from "is training permitted?" to "for each dataset, what is our access basis, and can we prove it?" Five channels cover nearly every corpus: a negotiated license, a purchase or subscription, public web access, a private enterprise transfer, and an unauthorized source such as a pirate library or a leaked dump. The cluster hub on training data compliance maps the surrounding obligations; this page covers only the access question.
What "lawful access" means under EU text and data mining rules
Under EU law, lawful access is a precondition for both DSM Directive exceptions: Article 3 for research organizations and cultural heritage institutions, and Article 4 for everyone else, including commercial model developers [4]. Article 4 permits reproductions for text and data mining only where the miner has lawful access and the rightsholder has not expressly reserved the right in an appropriate manner, such as machine-readable means for online content [4].
The directive's framing ties lawful access to content made available through subscriptions, open access or free online availability, which is why practitioners treat a paywall bypass or a credential-sharing scrape as outside the exception [2][4]. Article 53(1)(c) of the AI Act then requires GPAI providers to keep a copyright policy that identifies and complies with Article 4(3) reservations [3].
The GPAI Code of Practice copyright chapter, dated 10 July 2025, turns this into operational commitments for signatories: when crawling, reproduce and extract only lawfully accessible content, and do not circumvent effective technological measures [1]. Commentators read that commitment as covering paywalls and subscription restrictions, not only DRM [2]. The chapter also asks signatories to exclude websites recognized by courts or authorities as persistently infringing at commercial scale [1]. See the GPAI Code copyright chapter guide for how to write the policy itself.
UK, Singapore and Japan: different treatment of the source
Outside the EU, the UK and Singapore make lawful access a statutory condition, while Japan has no such condition but its guidance still looks at where data came from. The practical consequence is that the same downloaded corpus can be usable in one jurisdiction and unusable in another.
- United Kingdom. CDPA s29A allows copies by a person with lawful access for computational analysis, but only for non-commercial research [5]. As of October 2026, the Data (Use and Access) Act 2025 did not widen s29A, so commercial training on UK-protected works still needs a license [6]. Detail is in the UK commercial AI training guide.
- Singapore. The computational data analysis exception in the Copyright Act 2021 covers commercial use, but lawful access to the source material is a condition, and whether breaching a site's terms of use defeats lawful access is fact-specific [7].
- Japan. Article 30-4 is broad, but the 2024 guidance discussion asked directly whether training on pirated material is acceptable, and commentators read the draft as treating knowing collection from piracy sites as a factor against the developer [8]. The Japan Article 30-4 page covers when a license is still needed.
For a cross-country summary of exceptions, use text and data mining exceptions by country.
US fair use: how source lawfulness entered the analysis
US law has no lawful-access precondition written into a statute, but the acquisition channel has entered the fair-use analysis through agency guidance and case law. As of October 2026, the Copyright Office's Part 3 report on generative AI training remains a pre-publication version from May 2025 [9].
The report's discussion, as summarized by practitioners, treats knowing use of pirated or illegally accessed works as weighing against fair use, particularly where the commercial use produces content that competes with the originals [10]. More on its licensing analysis is in what the Copyright Office report means for licensing.
In Bartz v. Anthropic, the June 2025 summary judgment split along acquisition lines: training on books that were purchased or scanned from print was treated as fair use, while copies downloaded from pirate libraries were not protected by that ruling [11]. The parties settled, and the class settlement received final approval in July 2026 [12]. A settlement is not a merits ruling, so treat the summary-judgment reasoning as persuasive, not binding.
Kadrey v. Meta points the other way on its record: in June 2025 the court found fair use for training that included pirated books, largely because the plaintiffs had not shown market harm, and the case continues as of October 2026 [13]. The split means acquisition channel is a live, contested factor rather than a settled rule. Separately, the Third Circuit on 29 September 2026 affirmed that a non-generative legal research tool's use of Westlaw headnotes was not fair use, a reminder that commercial use of a rightsholder's licensed database content without its permission carries independent risk [14].
Acquisition channels compared by risk and evidence
Each channel has a typical access basis and a characteristic failure mode, and the evidence you need follows from the failure mode. The table below is a starting rubric for intake review; score individual datasets in the pre-acquisition risk assessment.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Channel | Access basis | Typical failure mode | Evidence that rebuts it |
|---|---|---|---|
| Negotiated license from the rightsholder | Contract grant | Licensor lacks rights in part of the corpus (contributor content, third-party embeds) | Executed license, licensor rights warranty, chain-of-title notes, scope covering training |
| Purchase or subscription | Terms of sale or subscription terms | Terms prohibit TDM, automated access or AI training | Archived terms version and date, account record, access logs within terms |
| Public web access | Free availability | Ignored robots.txt or other machine-readable reservation; paywall bypass | Crawl date, user agent, robots and TDM-reservation snapshot, no-circumvention log |
| Private enterprise transfer | Holder's permission | Holder's own customer or vendor contracts restrict reuse | Supplier approval, rights review of underlying contracts, de-identification record |
| Shadow library, torrent, leak | None | Infringing copies at source | None available; exclude and document removal |
Private enterprise data deserves a specific note: no third party has lawful access to a company's support tickets, engineering records or internal documents without the holder's permission, so the license is the access basis and the only one. That is why grounding versus training licenses must say explicitly which use is granted.
The acquisition record that rebuts source risk
The strongest defense against a source-lawfulness challenge is a dataset-level record created at intake, not reconstructed during litigation. Store it alongside the dataset in your data register, keyed by an immutable dataset ID and content hash.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset_id: ds-2026-0417
content_sha256: "9f2c...e81a" # manifest hash at receipt
acquisition_channel: negotiated_license # license | purchase | subscription | public_web | enterprise_transfer
access_basis:
instrument: "Data license v3, executed 2026-08-14"
grant_scope: [pretraining, fine_tuning] # not: rag_grounding
territory: worldwide
term_end: 2029-08-13
counterparty:
role: rights_holder # rights_holder | reseller | aggregator
rights_warranty: true
upstream_chain: "Supplier-generated records; no third-party embeds per rights review"
access_method:
delivery: access_controlled_transfer
circumvention: none
tdm_reservation_checked: n/a_private_data
chain_of_custody:
received: 2026-08-20T14:02Z
received_by: data-acq-team
transforms: [pii_redaction_v2, dedup_minhash]
exclusions:
- reason: "matched known shadow-library hash list"
records_removed: 0
reviewer: regulatory-counsel
review_date: 2026-08-22
Three fields do most of the work. acquisition_channel and access_basis answer the lawful-access question directly; chain_of_custody with a content hash proves the training copy is the copy you received. The same record feeds the EU training content summary and US state disclosures without a second data call.
Screening an existing corpus for unlawful sources
Retroactive screening is worth doing even when the corpus is already in training, because removal and documentation reduce forward exposure. Practical steps used by legal and data teams:
- Inventory by origin. Group shards by the acquisition channel recorded at ingest; any shard with no recorded origin is treated as high risk until proven otherwise.
- Match against known infringing collections. Compare file hashes, ISBNs and title-author pairs against lists derived from shadow libraries named in litigation and enforcement actions.
- Re-check terms for purchased and subscription content. Pull the terms version in force on the acquisition date and flag TDM or AI-training prohibitions.
- Re-verify crawl evidence. Confirm robots.txt and other machine-readable reservation snapshots exist for each crawl date, because an EU copyright policy must identify and comply with Article 4(3) reservations [3][4].
- Record exclusions. Log what was removed, when, and from which checkpoints forward.
The AI training data due diligence checklist lists supplier questions, and the comparison of licensed, synthetic and scraped data covers the trade-offs of each channel.
How SourceX handles access basis for private datasets
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases; it does not source scraped web content. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, and every release is approved by the supplying company. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which maps directly to the acquisition record above. The SourceX source policy explains what is and is not sourced, and AI data buyers can describe the data they need.
Getting lawfully licensed training data with a documented access basis
If your copyright policy requires a license as the access basis, start from the data you need rather than the supplier you know. SourceX looks for US businesses that hold the described data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted; nothing is contracted until a supplier agrees, and a request does not guarantee a match. Describe your dataset requirements to SourceX.
Sources
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- Slaughter and May (The Lens), "Mining the copyright chapter of the GPAI Code" (2025). https://thelens.slaughterandmay.com/post/102ktcs/mining-the-copyright-chapter-of-the-gpai-code
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- Hannes Snellman, "Text and data mining for AI training". https://www.hannessnellman.com/news-and-views/blog/text-and-data-mining-for-ai-training/
- UK Intellectual Property Office (GOV.UK), "Copyright, Designs and Patents Act 1988 - Consolidated (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
- Reed Smith, "UK copyright and AI report: the opt-out is dead, but what comes next?" (2026). https://www.reedsmith.com/articles/uk-copyright-and-ai-report-the-opt-out-is-dead-but-what-comes-next/
- Rouse, "Artificial intelligence in Singapore: copyright infringement defence for artificial intelligence and machine learning" (2024). https://rouse.com/insights/news/2024/artificial-intelligence-in-singapore-copyright-infringement-defence-for-artificial-intelligence-machine-learning
- Privacy World, "Japan's new draft guidelines on AI and copyright: is it really OK to train AI using pirated materials?" (2024). https://www.privacyworld.blog/2024/03/japans-new-draft-guidelines-on-ai-and-copyright-is-it-really-ok-to-train-ai-using-pirated-materials/
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- Manatt, "Copyright Office Releases Pre-Publication Report on Copyrighted Works in Generative AI Training" (2025). https://www.manatt.com/insights/insight/copyright-office-releases-pre-publication-report-on-copyrighted-works-in-generative-ai-training
- Kilpatrick Townsend, "Parties Reach a Landmark Settlement in the Bartz v. Anthropic Litigation" (2025). https://ktslaw.com/insights/alert/2025/9/parties-reach-a-landmark-settlement-in-the-bartz-v-anthropic-litigation
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (3d Cir.)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.