Text and language data
Openly Licensed Text Corpora for Commercial LLM Training: What They Cover and Where They Stop
Quick answer
Openly licensed corpora such as the Common Pile v0.1 (8TB from 30 sources), Common Corpus (launched in 2024 as public-domain text in six European languages), German Commons (about 154B tokens) and the GPT-NL Public Corpus give a pre-training team a defensible, commercially usable baseline. They are not uniform: filters differ on share-alike, NoDerivatives and dataset-level licenses, and all of them are thin on business registers, recent content and domain-specific non-English text. Treat them as the floor, then license the gaps.
By SourceX Editorial · Updated
This guide sits under text datasets for LLM training and catalogs the corpora themselves. License-compatibility rules and volume planning are covered in licensed text corpora for LLM pre-training.
The four corpora that define the openly licensed baseline
As of October 2026, four corpora are the reference points for a license-clean text baseline, and each was built with a different filter. Knowing the filter matters more than the headline size, because the filter determines what your legal reviewers still have to check.
| Corpus | Composition | License filter (as described by authors) | Main limitation for commercial buyers |
|---|---|---|---|
| Common Pile v0.1 | 8TB from 30 sources: research papers, code, books, encyclopedias, educational material, audio transcripts [1] | Public domain plus "openly licensed" text; code, data mixture and checkpoints released [1] | Share-alike components carry attribution and license-propagation duties; English-heavy |
| Common Corpus | At its 2024 launch, public-domain text in English, French, Dutch, Spanish, German and Italian [4] | Public domain at launch [4]; check the current release's dataset card for later additions | Skews toward older digitized material; little contemporary professional text |
| German Commons | About 154B tokens of openly licensed German [2] | Openly licensed German sources [2] | Smaller than web-derived German alternatives [2] |
| GPT-NL Public Corpus | Dutch-first, permissively licensed [3] | Excludes Creative Commons NC and SA material for commercial usability [3] | Narrower pool by design; Dutch focus |
The Common Pile result is the strongest evidence that the approach is viable: the authors report that their 7B-parameter Comma v0.1 models are competitive with models trained on unlicensed text at similar compute budgets [1]. That is a parity claim at one scale, not a guarantee for frontier-scale runs or for domains the corpus barely covers.
How license filters differ: SA, NC and ND are the deciding clauses
The practical difference between "open" corpora is which Creative Commons modules they let through, and the strictest commercial filter excludes NC, SA and ND. GPT-NL is explicit that it drops NC and SA material to keep the corpus commercially usable [3], while broader "openly licensed" definitions can retain share-alike text. If your model weights or outputs could be argued to be adaptations, share-alike terms raise questions your counsel should answer before mixing.
NoDerivatives is the quiet trap. Recent work on low-resource corpus licensing notes that an ND clause can be read to forbid tokenization, annotation and other transformations, which means an ND component can be unusable for training even when commercial use is permitted [8]. Check ND-licensed subsets separately rather than trusting a corpus-level label.
NC is the obvious exclusion, but it hides in mixed collections. The PubMed Central Open Access Subset carries machine-readable Creative Commons or similar licenses that allow broader reuse than the rest of PMC [7], yet that subset mixes commercial and non-commercial terms. Filter on the per-article license field, not on membership in the subset; the open-access scientific full-text guide covers those license tiers.
Dataset-level licenses do not settle rights in the underlying text
A permissive license on a dataset card covers the compilation, not necessarily every document inside it. Web-derived corpora such as C4 and RefinedWeb publish dataset-level licenses, but those terms come from the distributor and do not convert the copyrighted pages they contain into openly licensed text. They do not belong in a license-clean baseline unless your counsel has accepted a separate legal theory for them.
The reverse also happens. Pile of Law collects about 256GB of court opinions, filings, agency publications, statutes and casebooks from 35 sources [5]; much of the underlying government text is public domain, yet the compiled dataset card carries its own terms. Read both layers.
Label errors are common enough to plan for. The Data Provenance Initiative audited more than 1,800 text datasets and found license omission above 70% and license error rates above 50% on popular hosting sites [6]. Treat the card as a lead, then verify against the original source's terms; the copyright-status classification guide gives a four-way scheme for recording the result.
Build the baseline with a per-document license manifest
The safest way to use open corpora commercially is to carry license metadata down to the document, so a later filter change or takedown can be applied without re-crawling. Most openly licensed corpora preserve a source and license field; keep them through deduplication and tokenization rather than dropping them at the shard stage.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "cp-v0.1:pubmed-oa:000184223",
"corpus": "common-pile-v0.1",
"upstream_source": "pmc-oa",
"upstream_url": "https://example.org/article/000184223",
"license_spdx": "CC-BY-4.0",
"license_evidence": "article XML <license> element",
"nc": false, "sa": false, "nd": false,
"attribution_required": true,
"attribution_string": "Author A, Author B (2021), Journal X",
"language": "en",
"first_published": "2021-05-14",
"copyright_status": "licensed-open",
"filter_version": "baseline-2026-10-a",
"dedup_cluster": "minhash-7f3c91"
}
Use SPDX identifiers (CC-BY-4.0, CC-BY-SA-4.0, CC0-1.0, MIT) so policy filters are machine-checkable. Keep filter_version so you can show which rule set produced a given training mix. The metadata fields guide lists the fields to require when you later add licensed corpora on top.
Commercial-use filter checklist
- Drop any document whose license includes NC, unless counsel has approved a specific exception.
- Route SA and ND documents to a separate review queue instead of silently including them.
- Reject documents with no license evidence beyond the dataset card.
- Preserve attribution strings for CC BY material and confirm how you will satisfy attribution.
- Record
first_publishedso you can show the temporal coverage and find post-cutoff gaps. - Rerun near-duplicate detection across corpora; Common Pile and Common Corpus can overlap on digitized public-domain books.
Where open corpora stop: the coverage gaps to license
Open corpora are structurally thin wherever the text is commercially valuable, recent or private, and those gaps decide what you license next. Public-domain collections lean on older books and government documents, and openly licensed collections lean on research papers, code and encyclopedic text [1][4]. Neither source type produces much of how businesses actually write.
Gaps that pre-training teams typically find after building the baseline:
- Professional and business registers. Support transcripts, sales emails, engineering tickets, contracts, invoices and internal procedures are almost never released under open licenses. See proprietary text beyond web crawls.
- Post-cutoff content. Public-domain text outside government works is mostly old, and openly licensed sources update unevenly, so recent terminology and events are underrepresented.
- Domain-specific non-English text. German Commons and GPT-NL show that open non-English corpora are possible, but per-language volumes trail web alternatives [2][3], and domain text in those languages is thinner still. The non-English text licensing guide covers the commercial route.
- Dialogue and multi-turn interaction. Transcripts in open corpora are mostly lectures and proceedings, not two-party working conversations.
Before buying, test whether a candidate licensed corpus actually adds novelty over what you already hold; the method is in testing licensed data for overlap with public web corpora.
Documentation duties apply regardless of license
Using only open text does not remove disclosure duties; it makes them easier to meet. Under EU AI Act Article 53(1)(c), general-purpose model providers must maintain a copyright compliance policy that respects machine-readable reservations under Article 4(3) of the DSM Directive, and Article 53(1)(d) requires a public summary of training content [9]. The Commission's template for that summary is dated 24 July 2025 [10]. As of October 2026 these duties apply, with AI Office enforcement for new models from 2 August 2026.
In California, AB 2013 requires developers of generative AI systems made available to Californians to post documentation about training data, with a deadline of January 1, 2026 [11]. A per-document manifest like the one above lets you answer both regimes from the same records. For the broader trade-off between open, licensed and scraped inputs, see licensed vs synthetic vs scraped data and public vs proprietary data.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Filling gaps with licensed operational text
Once the open baseline is fixed, the remaining gaps are usually operational text held by companies, and that requires a negotiated license rather than a download. SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. You can describe the text your baseline lacks without naming specific companies. Foundation-model pre-training teams can scope these requests by register, language and date range.
Sourcing the text open corpora do not cover
SourceX helps AI teams wherever they are based find US companies that hold the operational text an open baseline lacks, then manages the process from assessment through licensing and ongoing purchases. Every release is approved by the supplying company, and personal details are removed or replaced before delivery. Start a buyer request at sourcex.si/buyers.
Frequently asked questions
Is the Common Pile safe for commercial training without further review?
It is designed for that use, but its "openly licensed" definition is broader than a strict no-NC/no-SA filter like GPT-NL's [1][3]. Run your own license policy over its per-source metadata, especially for share-alike and attribution duties.
Does a CC BY corpus allow commercial LLM training?
CC BY permits commercial use subject to attribution, which is why most commercial baselines accept it. The open question is how you satisfy attribution at training scale, so keep attribution strings in your manifest.
Can I count C4 or RefinedWeb as openly licensed?
Not on the strength of their dataset cards. Those licenses cover the compilation, not rights in the crawled pages, so classify them separately from public-domain and openly licensed text.
Sources
- arXiv (Kandpal et al., EleutherAI and collaborators), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
- arXiv, "The German Commons: 154 Billion Tokens of Openly Licensed Text for German Language Models" (2025). https://arxiv.org/html/2510.13996v1
- arXiv, "GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training" (2026). https://arxiv.org/pdf/2604.00920
- NISO Information Standards Quarterly / NISO I/O, "Common Corpus: A Multilingual Data Set for Training LLMs" (2024). https://www.niso.org/niso-io/2024/03/common-corpus-multilingual-data-set-training-llms
- arXiv (Henderson et al.), "Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset" (2022). https://arxiv.org/pdf/2207.00220
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- National Library of Medicine via Data.gov, "PubMed Central Open Access Subset (PMC OA)". https://catalog.data.gov/dataset/pubmed-central-open-access-subset-pmc-oa
- arXiv, "Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages" (2026). https://arxiv.org/abs/2606.28867
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.