Skip to content

Text and language data

Text Datasets for LLM Training: Licensed Corpora, Language Data and How Buyers Source Them

Quick answer

Text datasets for LLM training fall into five rights tiers: web-crawled corpora built on Common Crawl, openly licensed and public-domain collections, research corpora under distributor licenses, publisher-licensed books, news, journals and forums, and proprietary business text such as support tickets and internal documentation. A corpus's tier, not its download link, decides whether a commercial model can use it. Verify each dataset's license, then license what the open tiers lack, such as gated business text, lower-resource languages and text with a documented acquisition history.

By SourceX Editorial · Updated

Five rights tiers of LLM text data

Sort text sources by the legal basis that lets you train on them, because that basis, not size or host, decides commercial usability. Catalogs of open text are large: one 2024 survey reviews 444 datasets across 32 domains and 8 language categories [1]. Their license labels are unreliable: the Data Provenance Initiative's audit of more than 1,800 text datasets reported licenses omitted for over 70% of datasets on popular hosting sites and license error rates above 50% [2].

TierTypical sourcesRights basisCheck before commercial training
1. Web-crawledCommon Crawl snapshots and derivatives such as C4 and mC4; web-archive extractions such as HPLTNo author license; reliance on fair use in the US, or on text-and-data-mining (TDM) exceptions subject to opt-outs in the EURobots.txt and opt-out handling, Terms of Service restrictions, personal data, duplication
2. Openly licensed and public domainCommon Pile v0.1, German Commons, US federal works and USPTO patent grant textCC0, CC BY, CC BY-SA, public-domain statusAttribution, ShareAlike, NonCommercial and NoDerivatives terms, source by source
3. Research corporaLinguistic Data Consortium (LDC) and ELRA catalogs, academic parallel corpora such as JParaCrawlDistributor license or paid membershipWhether rights reach commercial products and models trained on the data
4. Publisher and platform licensedBook backlists, news archives, scholarly journals, forum and Q&A contentNegotiated license, often term-limitedWho holds training rights (author, publisher, platform or user), term, output limits
5. Proprietary business textSupport tickets, chat logs, email, wikis, standard operating procedures, translation memoriesThe company's ownership or control of its records, plus consentsCustomer and employee personal data, confidentiality duties, de-identification method

Web-crawled text carries crawl-time obligations, not a license. Common Crawl says it follows robots.txt and avoids paywalled and login-protected pages, and it argues that opt-outs should be applied by whoever uses its archive for AI training [3]. In the EU, Article 53(1)(c) of the AI Act requires providers of general-purpose AI (GPAI) models to keep a copyright policy that identifies and honors rights reservations made under Article 4(3) of Directive (EU) 2019/790 [4]. Signatories of the GPAI Code of Practice also commit, when crawling the web, to follow robots.txt as specified in RFC 9309 and not to circumvent paywalls [5].

Openly licensed text is now a credible baseline. The Common Pile v0.1 collects 8 TB of openly licensed and public-domain text from 30 sources, and its authors report that 7B-parameter models trained on it perform competitively with Llama 1 and Llama 2 7B models trained on similar compute [6]. Single-language collections exist too, such as the German Commons at 154 billion tokens [7]. The USPTO's weekly patent grant full text, available from 1976, is listed in the federal data catalog as public domain [8].

Research corpora come with research-first licenses. LDC's for-profit membership agreement licenses data for research and technology development at the sites listed in an exhibit, and lets members use portions in commercial work products only as far as copyright law and user agreements allow [9]. NTT's English-Japanese JParaCrawl corpus is research-only, and its terms extend that restriction to derived data and to translators trained on it [10].

Publisher-licensed text is negotiated deal by deal. Ithaka S+R tracks generative AI licensing agreements by publisher, purchaser, deal type and size [11]. Trade press reported that HarperCollins's 2024 deal with Microsoft offered $5,000 per title, split evenly between author and publisher, for a three-year license covering selected nonfiction whose authors opted in [12], so author consent, term and title selection all sat inside one grant.

Proprietary business text rarely appears in crawls because it sits behind logins [3]. Its rights question is mostly whether the company may share records that contain customer and employee data. For individual record types, see workplace email and chat datasets and licensed chat logs.

What open and crawled text no longer covers

Licensed text earns its price where the open tiers are thin: domains that now refuse crawling, lower-resource languages, gated business language, and text whose acquisition history you can document.

  • Domains that withdrew consent. The Consent in Crisis audit of 14,000 web domains found that between 2023 and 2024 about 5% of all C4 tokens, and over 28% of its most actively maintained critical sources, became fully restricted from use. Counting Terms of Service crawling restrictions, about 45% of C4 is restricted [13].
  • Languages and task types. The Data Provenance Initiative found that closed or non-commercial datasets dominate lower-resource languages and creative tasks [2]. Crawl-derived multilingual sets such as HPLT v2, extracted from 4.5 PB of Internet Archive and Common Crawl data [14], depend on automatic language identification, whose core failure modes include incorrect metadata, leakage from high-resource languages and confusion between closely related languages, according to the GlotLID authors [15].
  • Gated business language. Ticket notes, escalation threads and engineering postmortems use vocabulary and structure that public prose rarely shows; see proprietary text data beyond web crawls and operational free-text notes.
  • Documented acquisition. EU and California disclosure rules ask where training data came from [16][17]. The Bartz v. Anthropic class settlement was given final approval in July 2026; it is a settlement, not a ruling on fair use [18].

Matching the text source to the training stage

Specify the training stage first, because each stage buys a different unit: pre-training buys tokens, supervised fine-tuning buys conversations or examples, evaluation buys held-out items, and translation work buys aligned segments.

StageUnit to specifyWhat decides valueRead next
Pre-training and continued pre-trainingTokens under your tokenizer, by domain, language and dateNovelty against web crawls, deduplication, document lengthLicensed corpora for pre-training; domain corpora for continued pre-training
Supervised fine-tuning and preference tuningConversations, instruction-response pairs, draft-to-final editsVerified human authorship, turn structure, consentHuman-to-human dialogue corpora; fine-tuning datasets hub
EvaluationHeld-out items with references or rubricsAbsence from the public web and training mixesLLM evaluation datasets
Multilingual and translationNative documents, parallel segments, translation memories (TMX), termbases (TBX)Native vs translated origin, rights in derived modelsMultilingual data types compared; licensing translation memories
Retrieval and groundingDocument collections queried at run timeDisplay and citation rights rather than training rightsGrounding vs training licenses

Eight checks before a text corpus enters a commercial training run

Apply these to open and licensed text alike: most problems surface in the license text, the acquisition record or a sample scan, not the dataset card.

  1. Read the license text behind every subset. Hosting-site license labels are often missing or wrong [2]. Decide NonCommercial, ShareAlike and NoDerivatives terms per source, because one bundle can mix them.
  2. Confirm the grant reaches models and outputs. Research licenses can exclude derived models, as JParaCrawl's terms do [10]; see rights grants for pre-training data.
  3. Document opt-out handling for EU-facing models. GPAI providers have had to honor Article 4(3) reservations and publish a training-content summary since 2 August 2025 [4], using the Commission template dated 24 July 2025 [16]. Under Articles 111(3) and 113 of the AI Act as adopted in 2024, the Commission's enforcement and fining powers over GPAI providers apply from 2 August 2026, and models placed on the market before 2 August 2025 must comply by 2 August 2027 [27]; check the consolidated text, amended in July 2026, for any change. See EU TDM opt-outs under DSM Article 4.
  4. Collect the fields California's disclosure law asks for. California AB 2013 requires developers of generative AI systems offered in California to post documentation of dataset sources or owners, whether datasets include copyrighted or licensed material or personal information, collection periods and use of synthetic data. Posting was due by 1 January 2026 and again before each new release or substantial modification [17]. Ask suppliers to deliver those fields.
  5. Record how the text was acquired. The US Copyright Office's Part 3 report, a May 2025 pre-publication version, concludes that copying works into training datasets may be prima facie infringing unless fair use or another exception applies [19]. Rulings remain case-specific: in Kadrey v. Meta, a June 2025 partial ruling found fair use on that record, and as of October 2026 the case is ongoing [20].
  6. Measure duplication and crawl overlap before pricing. Lee et al. found a single sentence repeated more than 60,000 times in C4, and deduplicated training cut memorized output about tenfold [21]. Ask for MinHash near-duplicate clusters and an n-gram overlap report against recent Common Crawl snapshots; see near-duplicate detection with MinHash and LSH and novelty testing against web corpora.
  7. Scan inputs as well as targets for personal data. Carlini et al. extracted hundreds of verbatim training sequences from GPT-2, including contact details [22]. Ask for the de-identification method and a measured miss rate on a labeled sample; see PII redaction for LLM training data.
  8. Ask which quality and toxicity filters were applied. In a controlled study of 28 pretrained 1.5B-parameter models, quality filtering raised both downstream performance and toxic generation, while toxicity filtering reduced toxic output at some cost to generalization [23].

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

A per-document record that keeps rights attached to the text

Ask for one metadata record per document so rights, provenance and quality evidence travel with the text. JSON Lines suits this: each line is one valid JSON value, encoded as UTF-8 without a byte-order mark [24]. It is shown pretty-printed below; in delivery, each record sits on one line.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_id": "tkt-000184233",
  "record_type": "support_ticket_thread",
  "source_system": "helpdesk ticket export",
  "language": "en-US",
  "lang_id_confidence": 0.98,
  "authorship": "human: customer and agent",
  "created_range": ["2021-03-02", "2021-03-09"],
  "token_count": 742,
  "tokenizer": "buyer-specified",
  "deidentification": {
    "method": "NER replacement, manual check of sample",
    "entity_types": ["PERSON", "EMAIL", "PHONE", "ACCOUNT_ID"]
  },
  "license_ref": "license schedule A, item 3",
  "permitted_uses": ["pretraining", "sft"],
  "near_dup_cluster": null,
  "crawl_overlap_13gram": 0.0,
  "text": "Customer reports a duplicate charge after a plan downgrade..."
}

For the dataset-level description, the Data Statements schema (Version 2, 2021) includes language-data fields such as speech context, speaker demographic and annotator demographic [25]. The dataset delivery hub covers file formats and transfer.

Mistakes that shrink the value of purchased text

Most wasted spend on text comes from paying for tokens the model has already seen, rights the license does not grant, or volume counted in the wrong unit.

Start here: guides by text source

If you needStart with
A commercially usable open baselineOpenly licensed text corpora and where they stop
Books, news, journals or forumsBook corpus licensing, news archive text, scholarly journal licensing, forum content licensing
Corpora from research distributorsGetting commercial rights for research-only datasets
Public filings, patents and lawSEC filings corpora, patent full text, deposition and court transcripts
Non-English and low-resource languagesLicensing non-English corpora, low-resource language corpora, parallel corpus licensing
Human-written text for post-trainingVerifying human authorship, long-form professional writing

Compare licensed, synthetic and scraped training data, or see the AI training data licensing hub for rights questions beyond text. Other data types start from the AI data buyer's hub.

Where SourceX fits: business text from US companies

SourceX works in the fifth tier: it sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. The kinds of data it sources include support and sales histories, engineering records, documents, and finance and legal workflows. It does not source scraped public web content or train AI models. Datasets are sourced on request, not held in stock, and a request does not guarantee a matching dataset.

Rights review checks that the business owns or may share the records and that required consents are in place, and each dataset is delivered under a license that defines the records, permitted uses, term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded, and a processed sample is checked, though no de-identification method is perfect. You can describe the business text you need on the SourceX buyer page.

Tell SourceX which text your model is missing

If the gap in your text mix is business language that crawls and open corpora do not contain, describe the record types, languages, date range, volume in tokens and permitted uses you need. SourceX looks for US companies that hold that text, checks their licensing permissions, and manages the license and delivery. Start a text data request with SourceX.

Guides in this section

Sources

  1. arXiv (arXiv:2402.18041), "Datasets for Large Language Models: A Comprehensive Survey" (2024). https://arxiv.org/pdf/2402.18041
  2. Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  3. Common Crawl Foundation, "Common Crawl's Submission to the UK's Copyright and AI Consultation". https://commoncrawl.org/uk-copyright-and-ai-consultation
  4. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  5. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  6. Kandpal et al., NeurIPS 2025 Datasets and Benchmarks Track, "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://proceedings.neurips.cc/paper_files/paper/2025/hash/52acc050138d6f40dad6f12f91a4ce22-Abstract-Datasets_and_Benchmarks_Track.html
  7. arXiv (arXiv:2510.13996), "The German Commons – 154 Billion Tokens of Openly Licensed Text for German Language Models" (2025). https://arxiv.org/html/2510.13996v1
  8. Data.gov catalog (U.S. Patent and Trademark Office), "Patent Grant Full Text (1976 - Present)". https://catalog.data.gov/dataset/patent-grant-full-text-1976-present
  9. Linguistic Data Consortium, University of Pennsylvania, "LDC For-Profit Membership Agreement". https://Catalog.Ldc.Upenn.Edu/license/ldc-for-profit-membership.pdf
  10. NTT Communication Science Laboratories, "JParaCrawl". https://www.rd.ntt/cs/team_project/icl/lirg/jparacrawl/
  11. Ithaka S+R, "Generative AI Licensing Agreement Tracker". https://sr.ithaka.org/our-work/generative-ai-licensing-agreement-tracker/
  12. EMARKETER, "Microsoft and HarperCollins sign AI licensing deal, but author opt-in still required" (2024). https://www.emarketer.com/content/microsoft-harpercollins-sign-ai-licensing-deal--author-opt-in-still-required
  13. Longpre et al. (arXiv:2407.14933; NeurIPS 2024 Datasets and Benchmarks Track), "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
  14. arXiv (arXiv:2503.10267), "An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)" (2025). https://arxiv.org/pdf/2503.10267
  15. Kargaran et al. (Findings of EMNLP 2023; arXiv:2310.16248), "GlotLID: Language Identification for Low-Resource Languages" (2023). https://arxiv.org/pdf/2310.16248
  16. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  17. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  18. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  19. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  20. Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
  21. Lee et al. (ACL 2022; arXiv:2107.06499), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  22. Carlini et al. (USENIX Security 2021), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  23. Longpre et al. (NAACL 2024; arXiv:2305.13169), "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023). https://arxiv.org/pdf/2305.13169
  24. jsonlines.org, "JSON Lines". https://jsonlines.org/
  25. University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/
  26. DevClass, "Stack Exchange restricts access to dump of user-contributed data as critics complain license permits reuse for any purpose" (2024). https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose
  27. EUR-Lex, Official Journal of the European Union, "Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act)" (2024). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data