Text and language data
Text Datasets for LLM Training: Licensed Corpora, Language Data and How Buyers Source Them
Quick answer
Text datasets for LLM training fall into five rights tiers: web-crawled corpora built on Common Crawl, openly licensed and public-domain collections, research corpora under distributor licenses, publisher-licensed books, news, journals and forums, and proprietary business text such as support tickets and internal documentation. A corpus's tier, not its download link, decides whether a commercial model can use it. Verify each dataset's license, then license what the open tiers lack, such as gated business text, lower-resource languages and text with a documented acquisition history.
By SourceX Editorial · Updated
Five rights tiers of LLM text data
Sort text sources by the legal basis that lets you train on them, because that basis, not size or host, decides commercial usability. Catalogs of open text are large: one 2024 survey reviews 444 datasets across 32 domains and 8 language categories [1]. Their license labels are unreliable: the Data Provenance Initiative's audit of more than 1,800 text datasets reported licenses omitted for over 70% of datasets on popular hosting sites and license error rates above 50% [2].
| Tier | Typical sources | Rights basis | Check before commercial training |
|---|---|---|---|
| 1. Web-crawled | Common Crawl snapshots and derivatives such as C4 and mC4; web-archive extractions such as HPLT | No author license; reliance on fair use in the US, or on text-and-data-mining (TDM) exceptions subject to opt-outs in the EU | Robots.txt and opt-out handling, Terms of Service restrictions, personal data, duplication |
| 2. Openly licensed and public domain | Common Pile v0.1, German Commons, US federal works and USPTO patent grant text | CC0, CC BY, CC BY-SA, public-domain status | Attribution, ShareAlike, NonCommercial and NoDerivatives terms, source by source |
| 3. Research corpora | Linguistic Data Consortium (LDC) and ELRA catalogs, academic parallel corpora such as JParaCrawl | Distributor license or paid membership | Whether rights reach commercial products and models trained on the data |
| 4. Publisher and platform licensed | Book backlists, news archives, scholarly journals, forum and Q&A content | Negotiated license, often term-limited | Who holds training rights (author, publisher, platform or user), term, output limits |
| 5. Proprietary business text | Support tickets, chat logs, email, wikis, standard operating procedures, translation memories | The company's ownership or control of its records, plus consents | Customer and employee personal data, confidentiality duties, de-identification method |
Web-crawled text carries crawl-time obligations, not a license. Common Crawl says it follows robots.txt and avoids paywalled and login-protected pages, and it argues that opt-outs should be applied by whoever uses its archive for AI training [3]. In the EU, Article 53(1)(c) of the AI Act requires providers of general-purpose AI (GPAI) models to keep a copyright policy that identifies and honors rights reservations made under Article 4(3) of Directive (EU) 2019/790 [4]. Signatories of the GPAI Code of Practice also commit, when crawling the web, to follow robots.txt as specified in RFC 9309 and not to circumvent paywalls [5].
Openly licensed text is now a credible baseline. The Common Pile v0.1 collects 8 TB of openly licensed and public-domain text from 30 sources, and its authors report that 7B-parameter models trained on it perform competitively with Llama 1 and Llama 2 7B models trained on similar compute [6]. Single-language collections exist too, such as the German Commons at 154 billion tokens [7]. The USPTO's weekly patent grant full text, available from 1976, is listed in the federal data catalog as public domain [8].
Research corpora come with research-first licenses. LDC's for-profit membership agreement licenses data for research and technology development at the sites listed in an exhibit, and lets members use portions in commercial work products only as far as copyright law and user agreements allow [9]. NTT's English-Japanese JParaCrawl corpus is research-only, and its terms extend that restriction to derived data and to translators trained on it [10].
Publisher-licensed text is negotiated deal by deal. Ithaka S+R tracks generative AI licensing agreements by publisher, purchaser, deal type and size [11]. Trade press reported that HarperCollins's 2024 deal with Microsoft offered $5,000 per title, split evenly between author and publisher, for a three-year license covering selected nonfiction whose authors opted in [12], so author consent, term and title selection all sat inside one grant.
Proprietary business text rarely appears in crawls because it sits behind logins [3]. Its rights question is mostly whether the company may share records that contain customer and employee data. For individual record types, see workplace email and chat datasets and licensed chat logs.
What open and crawled text no longer covers
Licensed text earns its price where the open tiers are thin: domains that now refuse crawling, lower-resource languages, gated business language, and text whose acquisition history you can document.
- Domains that withdrew consent. The Consent in Crisis audit of 14,000 web domains found that between 2023 and 2024 about 5% of all C4 tokens, and over 28% of its most actively maintained critical sources, became fully restricted from use. Counting Terms of Service crawling restrictions, about 45% of C4 is restricted [13].
- Languages and task types. The Data Provenance Initiative found that closed or non-commercial datasets dominate lower-resource languages and creative tasks [2]. Crawl-derived multilingual sets such as HPLT v2, extracted from 4.5 PB of Internet Archive and Common Crawl data [14], depend on automatic language identification, whose core failure modes include incorrect metadata, leakage from high-resource languages and confusion between closely related languages, according to the GlotLID authors [15].
- Gated business language. Ticket notes, escalation threads and engineering postmortems use vocabulary and structure that public prose rarely shows; see proprietary text data beyond web crawls and operational free-text notes.
- Documented acquisition. EU and California disclosure rules ask where training data came from [16][17]. The Bartz v. Anthropic class settlement was given final approval in July 2026; it is a settlement, not a ruling on fair use [18].
Matching the text source to the training stage
Specify the training stage first, because each stage buys a different unit: pre-training buys tokens, supervised fine-tuning buys conversations or examples, evaluation buys held-out items, and translation work buys aligned segments.
| Stage | Unit to specify | What decides value | Read next |
|---|---|---|---|
| Pre-training and continued pre-training | Tokens under your tokenizer, by domain, language and date | Novelty against web crawls, deduplication, document length | Licensed corpora for pre-training; domain corpora for continued pre-training |
| Supervised fine-tuning and preference tuning | Conversations, instruction-response pairs, draft-to-final edits | Verified human authorship, turn structure, consent | Human-to-human dialogue corpora; fine-tuning datasets hub |
| Evaluation | Held-out items with references or rubrics | Absence from the public web and training mixes | LLM evaluation datasets |
| Multilingual and translation | Native documents, parallel segments, translation memories (TMX), termbases (TBX) | Native vs translated origin, rights in derived models | Multilingual data types compared; licensing translation memories |
| Retrieval and grounding | Document collections queried at run time | Display and citation rights rather than training rights | Grounding vs training licenses |
Eight checks before a text corpus enters a commercial training run
Apply these to open and licensed text alike: most problems surface in the license text, the acquisition record or a sample scan, not the dataset card.
- Read the license text behind every subset. Hosting-site license labels are often missing or wrong [2]. Decide NonCommercial, ShareAlike and NoDerivatives terms per source, because one bundle can mix them.
- Confirm the grant reaches models and outputs. Research licenses can exclude derived models, as JParaCrawl's terms do [10]; see rights grants for pre-training data.
- Document opt-out handling for EU-facing models. GPAI providers have had to honor Article 4(3) reservations and publish a training-content summary since 2 August 2025 [4], using the Commission template dated 24 July 2025 [16]. Under Articles 111(3) and 113 of the AI Act as adopted in 2024, the Commission's enforcement and fining powers over GPAI providers apply from 2 August 2026, and models placed on the market before 2 August 2025 must comply by 2 August 2027 [27]; check the consolidated text, amended in July 2026, for any change. See EU TDM opt-outs under DSM Article 4.
- Collect the fields California's disclosure law asks for. California AB 2013 requires developers of generative AI systems offered in California to post documentation of dataset sources or owners, whether datasets include copyrighted or licensed material or personal information, collection periods and use of synthetic data. Posting was due by 1 January 2026 and again before each new release or substantial modification [17]. Ask suppliers to deliver those fields.
- Record how the text was acquired. The US Copyright Office's Part 3 report, a May 2025 pre-publication version, concludes that copying works into training datasets may be prima facie infringing unless fair use or another exception applies [19]. Rulings remain case-specific: in Kadrey v. Meta, a June 2025 partial ruling found fair use on that record, and as of October 2026 the case is ongoing [20].
- Measure duplication and crawl overlap before pricing. Lee et al. found a single sentence repeated more than 60,000 times in C4, and deduplicated training cut memorized output about tenfold [21]. Ask for MinHash near-duplicate clusters and an n-gram overlap report against recent Common Crawl snapshots; see near-duplicate detection with MinHash and LSH and novelty testing against web corpora.
- Scan inputs as well as targets for personal data. Carlini et al. extracted hundreds of verbatim training sequences from GPT-2, including contact details [22]. Ask for the de-identification method and a measured miss rate on a labeled sample; see PII redaction for LLM training data.
- Ask which quality and toxicity filters were applied. In a controlled study of 28 pretrained 1.5B-parameter models, quality filtering raised both downstream performance and toxic generation, while toxicity filtering reduced toxic output at some cost to generalization [23].
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
A per-document record that keeps rights attached to the text
Ask for one metadata record per document so rights, provenance and quality evidence travel with the text. JSON Lines suits this: each line is one valid JSON value, encoded as UTF-8 without a byte-order mark [24]. It is shown pretty-printed below; in delivery, each record sits on one line.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "tkt-000184233",
"record_type": "support_ticket_thread",
"source_system": "helpdesk ticket export",
"language": "en-US",
"lang_id_confidence": 0.98,
"authorship": "human: customer and agent",
"created_range": ["2021-03-02", "2021-03-09"],
"token_count": 742,
"tokenizer": "buyer-specified",
"deidentification": {
"method": "NER replacement, manual check of sample",
"entity_types": ["PERSON", "EMAIL", "PHONE", "ACCOUNT_ID"]
},
"license_ref": "license schedule A, item 3",
"permitted_uses": ["pretraining", "sft"],
"near_dup_cluster": null,
"crawl_overlap_13gram": 0.0,
"text": "Customer reports a duplicate charge after a plan downgrade..."
}
For the dataset-level description, the Data Statements schema (Version 2, 2021) includes language-data fields such as speech context, speaker demographic and annotator demographic [25]. The dataset delivery hub covers file formats and transfer.
Mistakes that shrink the value of purchased text
Most wasted spend on text comes from paying for tokens the model has already seen, rights the license does not grant, or volume counted in the wrong unit.
- Paying for crawl duplicates. Test a sample before agreeing a price, as in testing licensed data for overlap with public web crawls.
- Quoting volume in gigabytes or words. Token counts depend on the tokenizer and the language, so ask for counts under yours; see token count estimation and tokenizer fertility and coverage.
- Treating translated text as native. Translated documents can look native in metadata; see native vs translated text.
- Assuming a content license equals platform permission. Stack Exchange content is licensed CC BY-SA, yet in 2024 the company put its data dump behind an agreement not to use it for AI training [26]. Review platform and API terms separately; see user-generated text via official APIs.
Start here: guides by text source
| If you need | Start with |
|---|---|
| A commercially usable open baseline | Openly licensed text corpora and where they stop |
| Books, news, journals or forums | Book corpus licensing, news archive text, scholarly journal licensing, forum content licensing |
| Corpora from research distributors | Getting commercial rights for research-only datasets |
| Public filings, patents and law | SEC filings corpora, patent full text, deposition and court transcripts |
| Non-English and low-resource languages | Licensing non-English corpora, low-resource language corpora, parallel corpus licensing |
| Human-written text for post-training | Verifying human authorship, long-form professional writing |
Compare licensed, synthetic and scraped training data, or see the AI training data licensing hub for rights questions beyond text. Other data types start from the AI data buyer's hub.
Where SourceX fits: business text from US companies
SourceX works in the fifth tier: it sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. The kinds of data it sources include support and sales histories, engineering records, documents, and finance and legal workflows. It does not source scraped public web content or train AI models. Datasets are sourced on request, not held in stock, and a request does not guarantee a matching dataset.
Rights review checks that the business owns or may share the records and that required consents are in place, and each dataset is delivered under a license that defines the records, permitted uses, term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded, and a processed sample is checked, though no de-identification method is perfect. You can describe the business text you need on the SourceX buyer page.
Tell SourceX which text your model is missing
If the gap in your text mix is business language that crawls and open corpora do not contain, describe the record types, languages, date range, volume in tokens and permitted uses you need. SourceX looks for US companies that hold that text, checks their licensing permissions, and manages the license and delivery. Start a text data request with SourceX.
Guides in this section
- Book Corpora for LLM Training: Licensing and AcquisitionHow AI teams source book text for LLM training: publisher and author licenses, per-title terms, opt-ins, output caps, acquisition evidence and formats.
- Human-to-Human Dialogue Data With Commercial Training RightsHow to license real multi-turn conversations between people for commercial AI training: what open sets lack, fields to require, consent, de-identification.
- Licensed Text Corpora for LLM Pre-Training: Buyer GuideWhere licensed LLM pre-training text comes from, how sources compare on volume and rights, and what to verify on tokens and grants before signing.
- Licensing Forum and Community Content for AI TrainingHow AI buyers license forum and community content for training: platform terms, contributor licenses, private-message exclusions and community backlash.
- Licensing Non-English Text Corpora for LLM Pre-TrainingHow to buy natively written non-English text for LLM pre-training: open baselines, language specs, native vs translated checks, token math, licenses.
- Licensing Scientific Papers for AI Training: Buyer GuideHow to license scholarly journal articles for LLM training: publisher vs aggregator deals, author opt-in and opt-out, version and retraction metadata.
- Long-Form Professional Writing Datasets for LLM TrainingHow to source reports, analyses, memos and proposals written by working professionals, with commercial training rights, metadata and de-identification.
- News Archive Data for AI Training: Depth, Metadata, RightsHow to license news archives for model training: archive depth, term, article metadata, wire and freelance carve-outs, and training vs display rights.
- Parallel Corpus Commercial Licensing: Bitext PitfallsHow to license parallel corpora for commercial MT and LLM translation: research-only traps, derived-model clauses, card conflicts and bitext pricing.
- Proprietary Text Data for LLMs: What Web Crawls MissWhich genres of business text never reach Common Crawl-derived corpora, how to measure their marginal value, and how to license them for LLM training.
- Verifying Human-Written Text Data Before You BuyHow AI data buyers verify that licensed text was written by humans: authorship evidence, sample tests, detector limits and license terms.
- Can API Data Train AI? UGC Text Rights, Tiers and DeletionOfficial API access is not a training license. How buyers check API terms, commercial tiers and deletion duties before training on user-generated text.
- Chatbot Conversation Logs as Training Data: Rights GuideHow to source chatbot conversation logs for SFT, preference and eval data: user-turn privacy, model-output terms, role-tagged schemas and buyer checks.
- Dialect and Regional Variety Text Data for LLMsHow to specify, source and verify regional variety text data for LLMs: variety labels, Spanish dialect coverage, generated-text checks and bias evaluation.
- Domain Vocabulary Data for LLMs: Jargon and AbbreviationsHow to test whether domain text covers a vertical's jargon, acronyms and controlled vocabularies, and which glossaries to request with the corpus.
- Estimating Text Dataset Token Counts Before You LicenseConvert a supplier's bytes, words or documents into tokens for your tokenizer and language mix, with sampling, deductions and an auditable counting basis.
- Licensing Customer Review Text Data for AI TrainingHow to license customer review text for sentiment and aspect models: platform, reviewer and repackager rights, filtering and a license checklist.
- Low-Resource Language Text Data: Licensing and SourcingHow to license text in under-resourced languages: Creative Commons conflicts, community licenses, publisher deals, consent records and volumes.
- Machine Translation Post-Editing Data: Triplets and MQMHow to source MT post-edit triplets and MQM error annotations for APE, quality estimation and translation reward models: fields, checks and rights.
- Multilingual Training Data Types for LLMs ComparedNative text, parallel corpora, translation memories, termbases or post-edits? Match each multilingual data type to the LLM capability it improves.
- Native vs Translated Text in Multilingual Training DataHow translated and machine-translated text affects multilingual LLM training, and how to require, label and verify natively authored text from suppliers.
- Openly Licensed Text Corpora for Commercial LLM TrainingCompare Common Pile, Common Corpus, German Commons and GPT-NL: license filters, NC/SA/ND traps, coverage gaps, and when to license text commercially.
- Operational Free-Text Notes Data for NLP and LLM TrainingHow to license terse free-text notes from case, work-order and CRM systems for extraction, summarization and eval: traits, pairing, PII review and specs.
- Patent Drafting AI Training Data: Disclosures to ResponsesHow to source patent drafting AI training data: invention disclosures, draft claims and office action responses, with privilege and confidentiality checks.
- Patent Full-Text Corpora for LLM Training: Grants and ClaimsHow to build a patent text dataset for LLM training from USPTO grant and application full text: coverage, XML parsing, claims, families and leakage.
- PMC Open Access Subset and arXiv: Commercial Training RightsWhich open-access papers can train commercial models: PMC OA license tiers, arXiv per-paper licenses, bulk access channels and a filtering checklist.
- Q&A Threads With Accepted Answers and Votes for LLM TrainingHow accepted-answer flags, vote scores and edit history in Q&A threads become SFT data, preference pairs and retrieval evals, and which fields to require.
- Recurring Text Data Feeds for Continual LLM TrainingHow to specify an ongoing text data feed for LLM training: cadence, date stamps, deltas, corrections, deletions, cutoffs and refresh license terms.
- SEC Filings Text Corpora for Financial LLM TrainingHow EDGAR-derived 10-K, 10-Q and proxy corpora are built for financial LLMs: section splitting, inline XBRL tables, dedup, mixtures and eval leakage.
- Speech Transcripts as LLM Training Text: Normalization GuideHow to clean, normalize and label ASR or human transcripts before LLM training: verbatim vs clean read, disfluencies, WER metadata and de-identification.
- Technical Manual Corpora for LLM Training: Rights and SpecsHow to source technical manuals, service guides and engineering documentation for LLM training: rights beyond retrieval, versions, dedup and metadata.
- Termbase and TBX Terminology Data for LLM TranslationHow to evaluate and license multilingual termbases (TBX, glossaries) for LLM translation: entry fields, term status, ownership checks and uses.
- Tokenizer Fertility: Testing a Corpus for Vocab ExtensionHow to measure tokenizer fertility, script coverage and byte fallback on a candidate corpus, and decide whether it can support vocabulary extension.
Sources
- arXiv (arXiv:2402.18041), "Datasets for Large Language Models: A Comprehensive Survey" (2024). https://arxiv.org/pdf/2402.18041
- Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Common Crawl Foundation, "Common Crawl's Submission to the UK's Copyright and AI Consultation". https://commoncrawl.org/uk-copyright-and-ai-consultation
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- Kandpal et al., NeurIPS 2025 Datasets and Benchmarks Track, "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://proceedings.neurips.cc/paper_files/paper/2025/hash/52acc050138d6f40dad6f12f91a4ce22-Abstract-Datasets_and_Benchmarks_Track.html
- arXiv (arXiv:2510.13996), "The German Commons – 154 Billion Tokens of Openly Licensed Text for German Language Models" (2025). https://arxiv.org/html/2510.13996v1
- Data.gov catalog (U.S. Patent and Trademark Office), "Patent Grant Full Text (1976 - Present)". https://catalog.data.gov/dataset/patent-grant-full-text-1976-present
- Linguistic Data Consortium, University of Pennsylvania, "LDC For-Profit Membership Agreement". https://Catalog.Ldc.Upenn.Edu/license/ldc-for-profit-membership.pdf
- NTT Communication Science Laboratories, "JParaCrawl". https://www.rd.ntt/cs/team_project/icl/lirg/jparacrawl/
- Ithaka S+R, "Generative AI Licensing Agreement Tracker". https://sr.ithaka.org/our-work/generative-ai-licensing-agreement-tracker/
- EMARKETER, "Microsoft and HarperCollins sign AI licensing deal, but author opt-in still required" (2024). https://www.emarketer.com/content/microsoft-harpercollins-sign-ai-licensing-deal--author-opt-in-still-required
- Longpre et al. (arXiv:2407.14933; NeurIPS 2024 Datasets and Benchmarks Track), "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
- arXiv (arXiv:2503.10267), "An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)" (2025). https://arxiv.org/pdf/2503.10267
- Kargaran et al. (Findings of EMNLP 2023; arXiv:2310.16248), "GlotLID: Language Identification for Low-Resource Languages" (2023). https://arxiv.org/pdf/2310.16248
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
- Lee et al. (ACL 2022; arXiv:2107.06499), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- Carlini et al. (USENIX Security 2021), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- Longpre et al. (NAACL 2024; arXiv:2305.13169), "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023). https://arxiv.org/pdf/2305.13169
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/
- DevClass, "Stack Exchange restricts access to dump of user-contributed data as critics complain license permits reuse for any purpose" (2024). https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose
- EUR-Lex, Official Journal of the European Union, "Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act)" (2024). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.