Skip to content

Text and language data

Book Corpora for LLM Training: Licensing Routes and Lawful Acquisition

Quick answer

A lawful books dataset for LLM training comes from one of four routes: a publisher license (per title or whole list), direct author opt-in, a rights-holder's posted standard terms, or physical copies you buy and digitize yourself. Each route clears a different layer of rights. Whichever you use, keep title-level evidence of how each book was acquired and who granted training rights, because acquisition from pirated libraries is the exposure courts and settlements have punished, and the file format determines how much cleanup the text needs.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why books are worth licensing separately from web text

Books give pre-training and long-context runs something crawls rarely do: long, edited, coherent text with stable structure across tens of thousands of tokens. That makes them useful for long-context training and for writing quality, and different in kind from the short-document mix covered in our guide to licensed text corpora for LLM pre-training. Books are also the content type where acquisition history matters most, since the largest copyright settlement to date turned on how books were obtained [4][5].

Treat a book corpus as its own procurement line within the broader text and language data hub. Its rights stack, its evidence trail and its file formats differ from news, forums or scholarly journals.

The four acquisition routes and what each one clears

The safest route is the one where the party granting rights demonstrably holds them for AI training; for trade books that usually means a publisher and the author together. The table below compares the routes buyers use in practice.

RouteWho grants rightsWhat it clearsTypical gaps
Publisher license, per titlePublisher, with author opt-inTraining use of opted-in titles for a set term [1]Only titles where authors agreed; output caps; term expiry
Publisher license, full list or backlistPublisherVolume, if the publisher holds AI rights for each titleOlder contracts may reserve AI-adjacent rights to authors [2]
Posted standard termsRights holder publishing a rate cardFast per-book or whole-corpus deals [6]Small catalogs; terms written for the licensor's use cases
Buy and scan print copiesYou, as owner of the physical copyLawful acquisition of the copy itself [4]Owning a copy is not a license; fair use is fact-specific and unsettled

The per-title route is the best documented. In the 2024 HarperCollins deal, reported to involve Microsoft, the fee was $5,000 per title split evenly between author and publisher, for a three-year term, limited to selected nonfiction, opt-in only, with verbatim output limited to 200 consecutive words or 5% of a book, and with a commitment not to use pirated content [1]. Use those terms as a reference point for structure, not as a market price.

The author layer: why publisher signatures are often not enough

For trade books, a publisher license frequently needs author permission because standard trade contracts grant enumerated rights and reserve the rest to the author [2]. AI training was not an enumerated right in most backlist contracts, so a publisher may lack the authority to license it alone. That is why the HarperCollins deal was structured as author opt-in rather than opt-out [1].

Royalty-based contract structures add a second problem: they were built around unit sales and subsidiary rights, not one-time training fees, so valuing a title and splitting proceeds is not mechanical [3]. Expect negotiations over whether a fee counts against advances and how multi-author, translated or illustrated works are handled. Translations and anthologies carry extra layers (translator, editor, each contributor), which matters for non-English text corpus licensing too.

Acquisition evidence: what Bartz v. Anthropic changed

How you obtained each book now determines a large part of your legal exposure. In Bartz v. Anthropic, the June 2025 summary judgment treated training on books that were purchased or scanned from print as transformative fair use, but did not extend that to downloading copies from pirated libraries [4]. The parties then settled for at least $1.5 billion, roughly $3,000 per work, with terms that included destroying the pirated datasets [4]. Final approval was granted in July 2026 [5].

Read this carefully. A settlement is not a merits ruling, a district-court fair use decision does not bind other courts, and related cases were still ongoing as of October 2026. The practical lesson is narrower and durable: buyers need title-level proof that no file came from a shadow library, and suppliers should be able to show it.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "title_id": "bk-000417",
  "isbn13": "9780000000002",
  "edition": "2nd, revised",
  "pub_year": 2019,
  "rights_holder": "Example Press LLC",
  "author_opt_in": {"status": "granted", "date": "2026-03-14", "instrument": "opt-in form v2"},
  "contract_reserved_rights_check": "publisher holds AI training rights via amendment",
  "acquisition_method": "publisher_supplied_epub",
  "source_file_sha256": "3f9a...c1",
  "third_party_content": ["epigraph excerpt", "2 licensed photos (excluded)"],
  "permitted_uses": ["pre-training", "long-context training"],
  "output_limit": "no more than 200 consecutive words verbatim",
  "term_end": "2029-03-31",
  "takedown_contact": "rights@example.invalid"
}

A record like this lets you answer an audit question about any single title in minutes. The fields that most often go missing in practice are the opt-in instrument, the acquisition method and the file hash that ties a training shard back to a delivered file.

Output limits and memorization controls you will be asked to honor

Book licenses increasingly constrain outputs, not just inputs. A cap such as 200 consecutive words or 5% of a book [1] is a model-behavior commitment, so you need engineering controls behind it. Deduplication is the first one: repeated passages in training data increase verbatim memorization, and removing near-duplicates reduces it [11]. Books are prone to this because the same work arrives as several editions, reprints and excerpts in anthologies.

  • Deduplicate across editions using ISBN groups and work-level IDs, then at the substring level.
  • Run extraction tests: prompt with an opening passage and measure the longest verbatim continuation against the source.
  • Keep a per-title output filter for licensed works where contract caps apply.
  • Log which model checkpoints included which titles, so term expiry and takedowns can be scoped.

Formats and preparation: getting clean text from book files

The best inputs are publisher-supplied EPUB 3 files or XML deliverables (BITS, the book counterpart of JATS, or publisher DTDs), because they carry chapter structure, footnotes and front and back matter as distinct elements. PDFs, especially print-ready PDFs with running heads, hyphenation and two-column layouts, need heavy repair, and scanned print adds OCR error. Specify the format in the license, not after delivery.

Ask for these preparation steps or perform them yourself:

  • Flag front matter (copyright page, dedication, table of contents) and back matter (index, acknowledgments, ads for other titles) so you can drop or down-weight them.
  • Rejoin end-of-line hyphenation and remove running heads, folios and page-break artifacts.
  • Preserve chapter and section boundaries as structural markers for long-context packing.
  • Exclude or separately flag third-party material: permissioned quotations, images with captions, and licensed tables.
  • Carry edition metadata (ISBN-13, edition statement, publication year, language, BISAC subject) through to every shard.

Our guide to metadata fields to require with licensed text corpora covers the full field list; for books, ISBN-to-work grouping and the edition statement are the ones that prevent silent duplication.

Regulatory documentation that touches book corpora

Book provenance feeds directly into disclosure duties. Under the EU AI Act, providers of general-purpose AI models must maintain a copyright compliance policy that respects text-and-data-mining reservations under Article 4(3) of the DSM Directive [7], and must publish a summary of training content using the Commission's template dated 24 July 2025 [8]. These duties have applied since 2 August 2025, with AI Office enforcement powers from 2 August 2026 for new models, as of October 2026 [7].

The GPAI Code of Practice's copyright chapter asks signatories to reproduce and extract only lawfully accessible content when crawling the web [9]; it does not govern licensed or purchased books directly, but the same lawful-access logic underpins the acquisition evidence above. In California, AB 2013 requires developers of generative AI systems to post documentation about their training data, with postings due January 1, 2026 [10]. A licensed book corpus with title-level records makes both disclosures easier to write and defend. The rights grant itself is covered in our guide to pre-training data license rights.

Buyer checklist for a book corpus deal

Before you sign, confirm each of these items in writing.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckEvidence to requestRed flag
Rights chain per titleContract or amendment showing AI rights, or author opt-in record [2]"Publisher holds all rights" with no title-level support
Acquisition methodSource of each file (publisher master, purchased copy scan) [4]Files with no provenance, or matches to known shadow-library hashes
Scope of usePre-training, long-context, fine-tuning, eval named explicitlyUse defined only as "AI purposes"
Output limitsExact verbatim cap and how compliance is measured [1]Caps with no measurement method agreed
Term and post-termWhat happens to trained models and retained copies at term endSilence on checkpoints trained during the term
Format and metadataEPUB or XML masters, edition metadata, matter flagsPDF-only delivery with no structural markup
Pricing basisPer title, per corpus or per token, and whether fees affect author royalties [3]Valuation with no stated basis

Posted rate cards [6] can shorten this list for small catalogs, but read them for scope: terms written by a licensor for its preferred uses may not cover long-context or commercial deployment.

Where operational long-form text fits alongside books

Books are not the only source of long, edited text. Many companies hold manuals, reports, proposals, policies and internal documentation that are long-form, professionally written and absent from web crawls. SourceX sources operational datasets from US companies, including documents, engineering records, and support, sales, finance and legal workflow histories, and manages the licensing process and ongoing purchases. If your goal is long-context or domain writing quality rather than trade books specifically, compare these routes in our guides to long-form professional writing datasets and technical manuals and documentation corpora, or describe the long-form text you need.

Publishers weighing the supply side of these deals can read whether media and publishing companies can sell data to AI companies and the media and publishing industry page.

Source long-form text for LLM training with SourceX

SourceX sources data on request from US businesses that hold it; nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, approved for release by the supplying company, and delivered under a license that defines records, uses, term and delivery. Describe the book-length or long-form text your team needs.

Sources

  1. eMarketer, "Microsoft and HarperCollins sign AI licensing deal, but author opt-in still required" (2024). https://www.emarketer.com/content/microsoft-harpercollins-sign-ai-licensing-deal--author-opt-in-still-required
  2. Authors Guild, "HarperCollins AI licensing deal" (2024). https://authorsguild.org/news/harpercollins-ai-licensing-deal/
  3. Publishers Weekly, "An AI licensing primer for book publishers" (2024). https://publishersweekly.com/pw/print/20241125/96590-an-ai-licensing-primer-for-book-publishers.html
  4. Kilpatrick Townsend, "Parties Reach a Landmark Settlement in the Bartz v. Anthropic Litigation" (2025). https://ktslaw.com/insights/alert/2025/9/parties-reach-a-landmark-settlement-in-the-bartz-v-anthropic-litigation
  5. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  6. Source Library, "Licensing". https://sourcelibrary.org/licensing
  7. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  8. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  9. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  10. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  11. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data