Text and language data
Licensed Text Corpora for LLM Pre-Training: Sources, Rights and Volume Planning
Quick answer
Licensed pre-training text comes from four places: openly licensed corpora, publisher and archive licenses, commercial text vendors, and proprietary business text held by operating companies. Use the open corpora as your free baseline, then license only the domains, languages or date ranges they cannot supply. Before signing, recount tokens with your own tokenizer, confirm how the licensor acquired the text, and get a grant that names training, model distribution and commercial use explicitly.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Where licensed pre-training text actually comes from
Licensed text at pre-training scale falls into four source types, and each one trades volume against rights clarity differently. The cluster hub on text datasets for LLM training covers the whole category. This page focuses on the acquisition decision once you know you need licensed volume, not scraped crawl.
- Openly licensed and public-domain corpora. The Common Pile v0.1 is an 8TB collection drawn from 30 sources, including research papers, code, books and encyclopedias [6]. Language-specific efforts such as the German Commons (154B tokens) [5] and the GPT-NL Public Corpus [4] apply similar filters. See what openly licensed corpora cover and where they stop.
- Publisher, book and archive licenses. These include book backlists, news archives, scholarly journals and trade publications. Rights sit with one counterparty, but third-party content inside the archive (wire copy, freelance pieces, quoted material) may not.
- Commercial text aggregators. Syndication and language-data vendors resell content from many publishers. One markets full-text articles with metadata and claims billions of tokens [8], and a multilingual vendor quotes its collection in words and language pairs [9]. Treat those figures as self-reported until you measure them.
- Proprietary operational text. Support tickets, engineering records, contracts and internal documents never appear in web crawls. See proprietary text beyond web crawls.
How the source types compare on volume and rights
Openly licensed corpora are the cheapest and most transparent route, but per-language volume is much smaller than web-scraped alternatives [5]. Commercial sources fill that gap at the cost of contract complexity. The comparison below is a starting grid for splitting a token budget across sources.
| Source type | Typical scale signal | Rights clarity | Main failure mode |
|---|---|---|---|
| Openly licensed corpora | Published token or byte counts (e.g., 8TB [6], 154B tokens [5]) | High per document, if NC/SA licenses are filtered out [4] | License metadata errors inherited from upstream hosts [10] |
| Single-publisher archive | Titles, articles or years of back issues | High for staff content; mixed for syndicated or freelance work | Grant covers "AI use" vaguely, with no model-distribution right |
| Aggregator or vendor | Vendor-stated tokens or words [8][9] | Depends on each upstream publisher's sublicense | Vendor cannot show the chain of title for each feed |
| Proprietary business text | Record counts, systems, date ranges | High once the holder approves release | Embedded personal data and third-party attachments |
License hygiene in the open-corpus tier is not a given. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission above 70% and license error rates above 50% on popular hosting sites [10]. The GPT-NL team excluded NonCommercial and ShareAlike data specifically so its corpus stays commercially usable [4]. Apply the same filter to any open component you plan to mix with licensed text.
Why acquisition history matters as much as the grant
How the licensor obtained the text is now a diligence item in its own right, not a formality. In Bartz v. Anthropic, the June 2025 ruling treated training on books that were lawfully purchased or scanned from print as transformative fair use, while the claims over copies taken from pirate libraries continued and led to a settlement [1]. That class settlement, totaling at least $1.5 billion, received final approval in July 2026 [2].
Kadrey v. Meta reached a narrower result on a different record: Meta won partial summary judgment on fair use in June 2025 because the plaintiffs had not shown market harm, and the case remains ongoing as of October 2026 [3]. The practical lesson for buyers is not a prediction about fair use. A clean license from a counterparty that cannot show lawful acquisition of the underlying text still leaves you exposed.
Ask every licensor for:
- The acquisition route for each content class (owned, commissioned, purchased, licensed in, scanned from physical copies).
- Whether any part of the corpus was assembled from shadow libraries, bulk scrapes or redistributed third-party datasets.
- The upstream agreements that allow them to sublicense for model training, not just display or syndication.
For a full treatment of book-specific routes, see book corpora and lawful acquisition.
Write the training grant around the model, not the data
The grant has to name training, model distribution and commercial use explicitly, because whether trained weights are a derivative work of the training text is legally unclear [4]. A license that permits "AI and machine learning purposes" without addressing weights, outputs and downstream distribution leaves the hardest question to a future dispute. The rights-grant structure is covered in depth in the rights grant you need for pre-training data.
Clauses to scope explicitly:
- Permitted uses: pre-training, continued pre-training, fine-tuning and evaluation listed separately. A license for retrieval or grounding is not a training license; see grounding license vs training license.
- Rights in derived weights: whether weights trained on the corpus can be released, hosted for third parties or open-weighted.
- Post-term status: what happens to trained checkpoints when the license ends.
- Jurisdictional fit: copying done in the UK cannot rely on CDPA s29A, which covers text and data analysis for non-commercial research only [14].
EU market entry adds its own documentation load. Article 53(1)(c) of the AI Act requires general-purpose model providers to maintain a copyright policy that honors rights reservations under Article 4(3) of the DSM Directive, and these duties have applied since 2 August 2025 [12]. The AI Office template for the public summary of training content, published 24 July 2025, asks providers to describe their data sources [13]. Licensors who can document provenance per feed make both obligations easier to meet.
Volume planning: measure tokens yourself
Vendor token counts are not comparable across sellers, because they depend on the tokenizer, on deduplication and on whether boilerplate was counted. Ask for a representative sample and run it through your own tokenizer and filtering pipeline before pricing anything. Then extrapolate the usable-token yield to the full delivery.
Deduplication changes the arithmetic most. Lee et al. found many near-duplicate examples in common language-modeling datasets, including one sentence repeated more than 60,000 times in C4, and reported that over 1% of unprompted output from models trained on them was copied verbatim [11]. Licensed archives carry the same risk through syndicated wire stories, templated filings and reprinted articles. Also test overlap with the crawl you already hold: see testing licensed data for overlap with public web corpora.
Pricing units vary too. One publisher posts training terms of at least $250 per book on a non-exclusive basis, or $200,000 per year for its full corpus [7]. Treat that as one publisher's posted terms, not a market benchmark. Normalize every quote to cost per usable, deduplicated token in your tokenizer before you compare.
Illustrative example: invented to show structure; it does not describe an available dataset.
Licensed pre-training text: request and volume-planning sheet
| Field | What to specify | Example entry |
|---|---|---|
| Domain and content class | Genre, source systems, document types | B2B trade journals, full text plus headlines and bylines |
| Languages and varieties | ISO 639 codes, script, regional variety | en-US, es-MX |
| Date range | First and last publication dates | 2005-01 to 2026-06 |
| Format | Encoding, container, markup handling | UTF-8 JSONL, one document per line, HTML stripped |
| Required metadata | Per-document fields | doc_id, source, pub_date, language, license_id, rights_holder |
| Volume target | Usable tokens after your dedup, your tokenizer | 20B tokens, measured on a 1% sample |
| Permitted uses | Each use named | Pre-training and continued pre-training; no grounding |
| Rights in weights | Release and hosting terms | Commercial API and private deployment |
| Acquisition evidence | Proof per content class | Publisher-owned staff content; wire copy excluded |
| Refresh cadence | One-time or recurring | Quarterly increments, same schema |
| Pricing unit | How the seller prices | Per usable token, normalized from per-title quote |
For the full field list, use metadata fields to require with licensed text corpora.
Pre-signature verification checklist
Run these checks on a representative sample before the contract is final, because defects found after signing are harder to remedy.
- Token recount: tokenize the sample with your production tokenizer and compare against the vendor's stated count [8].
- Dedup rate: run exact and near-duplicate detection within the sample and against your existing crawl [11].
- License field audit: confirm every document carries a license or rights-holder identifier that matches the contract schedule [10].
- NC/SA screen: drop or segregate any component carrying NonCommercial or ShareAlike terms [4].
- Third-party content scan: flag wire stories, embedded images, quoted letters and attachments; see third-party content inside licensed corpora.
- Acquisition attestation: obtain a written description of how each content class was acquired [1].
- Personal data scan: check for names, emails, phone numbers and account numbers, especially in business text.
- EU documentation: confirm the licensor can supply source descriptions you can reuse in the Art. 53 training-content summary [13].
When operational business text belongs in the mix
Operational text is worth licensing when your model needs language that public and published sources do not contain: terse ticket notes, engineering change records, contract redlines and finance workflows. It rarely matches publisher archives on raw volume, so it usually goes to continued pre-training or domain mixes; see domain corpora for continued pre-training and operational free-text notes.
SourceX sources this kind of text on request from US companies: support and sales histories, engineering records, documents, and finance and legal workflows. It does not hold stock, and categories are not inventory, so a request does not guarantee a match. SourceX does not source scraped web content. Every dataset is rights-reviewed for ownership and consents, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked; no method is perfect. Buyers can describe the corpus they need without naming suppliers.
For the general licensing process across data types, see how to license proprietary data for AI training, and for the trade-offs between source routes, licensed vs synthetic vs scraped data. The pretraining glossary entry defines the stage itself.
Source licensed pre-training text through SourceX
SourceX finds US companies that hold the operational text you describe and manages the commercial process, from assessing the data and licensing permissions to agreeing pricing and allowed uses in a license. Nothing is contracted until a supplier agrees, and every release is approved by the supplying company. Tell SourceX what text your pre-training mix is missing.
Frequently asked questions
Is there such a thing as a "copyright-cleared" text corpus?
No contract can remove all copyright risk, because the status of trained weights as derivative works is unsettled [4]. What you can get is documented ownership, lawful acquisition and an explicit training grant. Treat "cleared" in vendor copy as a claim to verify, not a guarantee.
Should we license non-English text separately?
Usually yes, because openly licensed volume per language is limited [5] and rights holders differ by market. See licensing non-English text corpora for language-specific sourcing and tokenizer effects.
Does a pre-training license cover fine-tuning and evaluation?
Only if it says so. List pre-training, continued pre-training, fine-tuning and evaluation as separate permitted uses, and confirm whether outputs from each may be used commercially.
Sources
- Kilpatrick Townsend, "Parties Reach a Landmark Settlement in the Bartz v. Anthropic Litigation" (2025). https://ktslaw.com/insights/alert/2025/9/parties-reach-a-landmark-settlement-in-the-bartz-v-anthropic-litigation
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
- arXiv, "GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training" (2026). https://arxiv.org/pdf/2604.00920
- arXiv, "The German Commons: 154 Billion Tokens of Openly Licensed Text for German Language Models" (2025). https://arxiv.org/html/2510.13996v1
- arXiv, "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
- Source Library, "Licensing". https://sourcelibrary.org/licensing
- Syndigate, "Fully Licensed Premium Datasets for Training LLMs and AI Applications". https://www.syndigate.info/fully-licensed-premium-datasets-for-training-llms-and-ai-applications/
- TAUS, "Data for AI". https://www.taus.net/data-solutions/data-for-ai
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- arXiv (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.