Provenance, rights and permitted use
Screening Training Data for Pirated and Shadow-Library Sources
Quick answer
Pirated books in AI training data carry concrete legal exposure: the Bartz v. Anthropic class settlement received final approval in July 2026 [2]. Screening means three things done together: tracing every book or article file to an acquisition event, matching files against known shadow-library collections by hash and metadata, and keeping a per-item record that proves how each lawful copy was obtained. Anything that cannot be traced gets quarantined, not trained on.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why acquisition method matters separately from training use
US courts have started to treat how a copy was acquired as a separate question from whether training on it is fair use. In Bartz, the June 2025 summary judgment ruling treated training on books Anthropic had purchased or scanned from print as transformative, while the pirated library copies were left for trial and then settled [1].
Kadrey v. Meta came out differently on its record: Judge Chhabria granted Meta partial summary judgment on training with books that included pirated copies, largely because plaintiffs did not present evidence of market harm [3]. That case is not over, and in June 2026 plaintiffs sought interlocutory appeal of the ruling as it relates to downloading from shadow libraries [4]. Counsel should read the two decisions as a warning that acquisition facts are discoverable and contested, not as a safe harbor.
Two further signals point the same way. The US Copyright Office's Part 3 report, still a pre-publication version as of October 2026, concludes that many acts in AI training implicate the reproduction right and discusses the relevance of knowingly using pirated or illegally accessed works [8]. The Third Circuit's September 2026 affirmance in Thomson Reuters v. Ross held that copying for a non-generative legal research tool was not fair use, a reminder that fair use is decided case by case [11].
For deeper background on the legal theory, see lawful access and pirated sources and what the Copyright Office report means for licensing.
What the EU expects from model providers
EU obligations make lawful access an explicit documentation item for general-purpose model providers. Article 53(1)(c) of the AI Act requires a copyright policy, and the Copyright chapter of the GPAI Code of Practice, published 10 July 2025, offers signatories a way to show compliance by maintaining that policy and committing to reproduce and extract only lawfully accessible content [5][6]. Commentary on the final chapter reports a further commitment to exclude websites recognized by courts or authorities as persistently infringing; check the measure wording in the Commission text before relying on it [5].
The Article 53(1)(d) public summary of training content, with a template dated 24 July 2025, asks providers to describe their main data sources [7]. A shadow-library dump discovered in a corpus after a summary has been published is both a copyright problem and a disclosure problem.
Where pirated copies enter a corpus
Most pirated text reaches a lab indirectly, through intermediaries that never recorded where their files came from. The Data Provenance Initiative audited more than 1,800 text datasets and found license information omitted in over 70% of cases and wrong in over 50% on popular hosting sites [9]. A dataset card that says "books, public domain and licensed" is a claim, not evidence.
Common entry points include:
- Aggregated open corpora. Bundles that include a books subset assembled by a third party, sometimes renamed or resharded so the original subset name no longer appears.
- Vendor "licensed" book packs. Suppliers who bought files from a reseller and cannot show the publisher agreement or the purchase chain.
- Web crawls. Pages that mirror full-text EPUB or PDF content from shadow libraries, or link-farm sites hosting book text in HTML.
- Internal scraping and research archives. Directories that researchers populated years ago with no acquisition log, often named after the source site.
- Article and paper collections. PDFs pulled from sites that circumvent publisher paywalls, which overlap with scholarly journal licensing questions.
The screening workflow, step by step
A defensible screen combines manifest review, technical matching and documentary proof, because each method misses what the others catch. Run it before training on any new books or long-form text source, and as a retroactive provenance audit on corpora already in use.
1. Build or demand a source manifest. Every file needs a row: original filename, source system or URL, acquisition date, acquiring party and acquisition method. Reject deliveries where the manifest is reconstructed after the fact without supporting records.
2. Hash at the file level before any normalization. Compute SHA-256 and, where the shadow-library ecosystem uses them, MD5 hashes of the original EPUB, PDF, MOBI or DJVU files. Shadow-library catalogs commonly index files by MD5, so original-file hashes are the most direct match key; once text is extracted to JSONL or Parquet, that link is lost.
3. Match metadata fingerprints. Normalize ISBN-10 and ISBN-13, title, author, publisher and edition, then compare against catalog metadata from known collections your counsel has approved for screening use. Look for shadow-library tells inside files: injected filename prefixes, embedded library IDs, uniform OCR artifacts, repackaged EPUB OPF metadata with a non-publisher creator tool, and watermark pages stripped of publisher branding.
4. Run near-duplicate detection on extracted text. MinHash or SimHash over shingled text finds copies that were re-encoded, OCRed again or partially edited. Hash matching alone fails on these, so treat a text-level near-duplicate of a known pirated title as a match unless an acquisition record explains it.
5. Reconcile against acquisition evidence. For each title, require one of: a purchase receipt or invoice tied to the specific edition, a publisher or aggregator license covering training use, a scan log for a physically owned copy, or a public-domain or open-license determination. The chain-of-title documents page covers what each record should contain.
6. Quarantine, decide and log. Unexplained matches move to a quarantine bucket with no read access from training jobs. Counsel decides on removal, replacement with a lawfully acquired copy, or retention with a documented basis, and the decision is logged per item.
Practical artifact: per-item acquisition record and screening decision table
A per-item record makes acquisition provable later, during litigation discovery or an EU training-content summary review. Store it alongside the data, ideally in a machine-readable format such as Croissant-RAI, which builds data life-cycle and provenance fields into dataset documentation [10].
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"item_id": "bk-000417",
"isbn13": "9780000000000",
"title": "Example Title",
"edition": "2nd, paperback",
"original_file": {"format": "EPUB", "sha256": "3f9a...", "md5": "a1c4..."},
"acquisition": {
"method": "publisher_license",
"counterparty": "Example Publisher",
"agreement_ref": "LIC-2026-031",
"acquired_on": "2026-03-12",
"evidence_uri": "vault://contracts/LIC-2026-031.pdf"
},
"screening": {
"hash_match_known_collection": false,
"metadata_match_known_collection": false,
"near_duplicate_max_jaccard": 0.12,
"screened_on": "2026-09-30",
"screen_version": "v4"
},
"decision": "cleared_for_pretraining",
"decided_by": "legal-ops"
}
Illustrative example: invented to show structure; it does not describe an available dataset.
| Screening result | Acquisition evidence | Default decision | Notes |
|---|---|---|---|
| No hash, metadata or near-duplicate match | Receipt, license or scan log | Clear | Keep evidence linked to item_id |
| No match | None | Quarantine | Absence of a match is not proof of lawful source |
| Hash match to known collection | Receipt for same edition | Escalate to counsel | Lawful copy may still be the pirated file; replace with owned file |
| Hash or near-duplicate match | None | Remove | Purge from all shards, caches and derived sets |
| Metadata match only | License covering training | Clear with note | Record the license scope against the title |
Removing infringing copies from a corpus
Removal has to reach every derived artifact, not just the raw file store. Delete the original file, extracted text, tokenized shards, deduplication indexes, evaluation splits built from the same source and any cached embeddings in retrieval stores. Record the removal with item IDs, timestamps and the shard list so the deletion can be shown later.
Removing files from data does not remove what a trained model has learned. Whether to retrain, fine-tune away or document and accept the exposure is a legal and technical decision for counsel; log the checkpoint IDs trained on the affected data either way, and track that status in your training data use register.
What to require from book and article suppliers
Suppliers of long-form text should be able to answer acquisition questions per title, not per dataset. Ask for the source manifest with original-file hashes, the license or purchase chain for each publisher or rightsholder, the method used to digitize owned copies, and confirmation that no files were obtained from shadow libraries or paywall-circumvention sites. A data rights attestation turns those answers into a signed statement, and the due diligence checklist places them in a wider review.
For lawful routes to book content, see book corpus licensing and licensed text corpora for pre-training. The trade-offs between licensed, synthetic and scraped inputs are compared in licensed vs synthetic vs scraped data, and the full cluster sits under data provenance for AI training data.
Operational datasets from businesses are a different risk profile from books. SourceX does not source scraped web content, and every dataset it sources is rights-reviewed for ownership and consents before delivery under a license that defines records, uses, term and delivery; buyers can describe the data they need.
Sourcing operational data without pirated-source risk
If your team needs operational text instead of books, SourceX sources datasets on request from US companies, such as support histories, engineering records, documents and finance or legal workflows, with every release approved by the supplying company. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Start a buyer request at SourceX.
Frequently asked questions
Does a purchase receipt cure a file that matches a shadow-library hash?
Not automatically. The receipt shows you owned a lawful copy, but the file you trained on may still be the downloaded one. Replace it with a file derived from the purchased copy and record both events.
Is screening only relevant to US litigation?
No. The GPAI Code of Practice copyright chapter frames lawful access as a commitment for signatories, and the EU training-content summary asks providers to describe main data sources [5][7].
How often should a corpus be rescreened?
Rescreen whenever the reference collections you match against are updated, whenever a new source is merged, and before each major pre-training run. Version the screen so each item records which screen it passed.
Sources
- Kilpatrick Townsend, "Parties Reach a Landmark Settlement in the Bartz v. Anthropic Litigation" (2025). https://ktslaw.com/insights/alert/2025/9/parties-reach-a-landmark-settlement-in-the-bartz-v-anthropic-litigation
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
- ChatGPT Is Eating the World, "Kadrey asks Judge Chhabria to certify for interlocutory appeal the 2025 summary judgment re Meta's downloading from shadow libraries" (2026). https://chatgptiseatingtheworld.com/2026/06/09/kadrey-asks-judge-chhabria-to-certify-for-interlocutory-appeal-to-9th-circuit-the-judges-2025-summary-judgment-re-metas-downloading-from-shadow-libraries-for-the-purpose-of-ai-training/
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- European Commission, "The General-Purpose AI Code of Practice" (2025). https://digital-strategy.ec.europa.eu/en/policies/gpai-code-practice
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Jain et al., MLCommons, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.