Skip to content

Text and language data

Licensing Non-English Text Corpora for LLM Pre-Training

Quick answer

To license non-English text for pre-training, first measure what open web corpora such as HPLT v2 and ROOTS already give you in each language, then buy only what they lack: natively written, rights-documented text in specific registers and locales. Specify ISO 639-3 language codes, ISO 15924 scripts, locales, register mix and a ceiling on translated content. Count tokens with your own tokenizer, and get a license that names pre-training, derivative models and every territory you deploy in.

By SourceX Editorial · Updated

What open multilingual corpora already cover

Open corpora set the baseline, so a licensed purchase only makes sense where they are thin, noisy or legally uncertain for your use. HPLT v2 was extracted from 4.5 PB of Internet Archive and Common Crawl data [1], and BigScience's ROOTS assembled 1.6TB across 59 languages (including 46 natural languages) for a 176B-parameter model [2]. Both are broad, but both inherit the web's skew toward marketing copy, forums, boilerplate and machine-translated pages, especially outside the largest languages.

Openly licensed alternatives reduce rights risk but shrink volume. The Common Pile v0.1 gathers 8TB of public domain and openly licensed text from 30 sources [5], and the German Commons reaches 154 billion German tokens of openly licensed text [4], which is substantial yet smaller than web-derived German. The GPT-NL team reports that finding enough curated, copyright-compliant data is especially hard for languages other than English [3]. That scarcity is where licensed native text earns its price.

Before you write a request, run a gap analysis per language against these baselines. Our guide to openly licensed text corpora for commercial training covers where those collections stop.

Native text versus translated text: why it decides value

Natively written text is the asset; translated text mostly teaches a model translationese. Translated output carries source-language word order, calqued idioms and flattened register, and machine-translated pages on the web are hard to filter once mixed in. If you need translation pairs, that is a different product with different pitfalls, covered in parallel corpus licensing.

Ask every supplier to label each document's origin: native, human translation, machine translation or post-edited machine translation. Then verify on a sample. Practical checks include a translationese classifier, looking for paired documents in another language under the same identifier, and checking whether CMS or localization-system fields (source locale, translation job ID) exist in the export.

Operational business text tends to be native by default. Support tickets written by local agents, sales emails, internal wikis, contracts and engineering notes in a German or Brazilian subsidiary were written for local readers, not localized from English. The comparison of multilingual training data types explains when native text, translation memories or post-edits each fit.

Writing a language-specific corpus request

A useful request specifies language, script, locale, register and provenance per slice, not just "Spanish, 10B tokens." Language labels alone hide large differences: es-MX and es-ES differ in vocabulary, forms of address and legal terms, and evaluation work such as La Leaderboard treats Spanish varieties and the languages of Spain and Latin America separately [9]. Serbian appears in Cyrillic and Latin scripts; Chinese needs Hans versus Hant; Norwegian needs nb versus nn.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
language (ISO 639-3)spaUnambiguous code; avoids "Spanish (other)" buckets
script (ISO 15924)LatnSeparates script variants such as srp-Cyrl and srp-Latn
locale (BCP 47)es-MX, es-ARRegional vocabulary and forms of address
register mix40% support dialogue, 30% business documents, 20% technical, 10% legalPre-training mix you will actually use
native shareat least 95% native; MT and post-edited MT flaggedLimits translationese
date range2015-2026Currency of terminology
target volumestated in your tokenizer's tokens, per localeAvoids word-count versus token-count disputes
formatJSON Lines, UTF-8, one document per linePredictable ingestion
exclusionsno scraped web pages, no contact listsKeeps the slice distinct from what you already have

Keep the request about the data, not about named companies. A supplier search works better when it describes record types, languages and volumes and lets the intermediary find holders.

Token math: counting volume with your tokenizer

Volume quoted in words or bytes does not convert evenly to tokens across languages, so price and plan in tokens from your own tokenizer. Fertility, the average number of tokens per word, is typically higher for morphologically rich languages and non-Latin scripts, which is why work on adapting pretrained models to new languages pairs vocabulary extension with continued pre-training [8]. A million Finnish or Thai words can cost far more context and compute than a million English words.

Ask for a 1-5 MB sample per locale and run three numbers yourself: tokens per word, tokens per byte, and the share of bytes in your target script after Unicode NFC normalization. Watch for mojibake from Windows-1252 or Shift_JIS exports, mixed-script contamination, and boilerplate signatures that inflate counts. The guide on tokenizer fertility and corpus coverage covers whether to extend your vocabulary first.

Deduplicate before you sign a volume number. Run MinHash near-duplicate detection within the sample and against your existing web mix; templated emails and ticket macros can cut effective volume sharply. See near-duplicate detection with MinHash and LSH.

Commercial market signals and pricing references

Public pricing for licensed multilingual text is rare, and the few published figures are vendor self-reports for translation data rather than monolingual corpora. TAUS describes a 7.4-billion-word collection across 483 language pairs sold by language pair, with historical list rates of about EUR 1,500-2,500 per million words that a 2024 sale cut temporarily (self-reported; the sale ended April 30, 2024) [6]. Treat that as a data point for parallel data, not a benchmark for native monolingual text.

Monolingual operational text is usually priced per deal, driven by exclusivity, language rarity, register, de-identification effort and refresh cadence. Budget for your own costs too: language-identification QA, PII review by native speakers, and deduplication.

Rights review across jurisdictions

Non-English text often comes from non-US authors or entities, so rights review has to work across copyright regimes rather than assume US fair use. Statutory text and data mining exceptions are narrow for commercial training: UK guidance on the research exception states that contract research for an outside company is unlikely to count as non-commercial [11], and as of October 2026 CDPA s29A remains non-commercial only. In the EU, check DSM Article 4 opt-outs for any web-sourced component, covered in EU TDM opt-outs.

License metadata on public datasets is unreliable. The Data Provenance Initiative audited more than 1,800 text datasets and found license omission above 70% and license error rates above 50% on popular hosting sites [7]. Do not inherit a license label; get a chain of title per source.

If you place a general-purpose model on the EU market, the AI Act's public summary of training content (template published 24 July 2025 under Article 53(1)(d)) asks you to describe data sources [10]. Licensed corpora should come with enough provenance metadata to complete it. Our article on pre-training data license rights lists the grants to request: pre-training, derivative models, weights distribution, territory and term.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Privacy and de-identification for native-language business text

Native business text carries personal data in local formats, so redaction must be language- and locale-aware. English-trained NER misses Spanish compound surnames, German Steuer-ID numbers, Brazilian CPF numbers, Japanese addresses and honorific-marked names. Ask suppliers which tools and patterns were used per language, and audit a sample with native speakers. PII redaction for LLM training data covers measuring miss rates.

Delivery format and acceptance checks

Agree on format and acceptance tests before delivery so that disputes are about data, not parsing. JSON Lines works well for text corpora: files must be UTF-8 without a byte order mark, and each line must be a valid JSON value [12]. Require per-document fields for language, script, locale, origin (native or translated), source type, date and a stable document ID; the list in metadata fields for licensed text corpora is a good template.

Illustrative example: invented to show structure; it does not describe an available dataset.

Acceptance checklist per locale:

  • Language ID agreement of at least 98% between supplier labels and your own classifier (fastText lid.176 or GlotLID).
  • Native-origin share at or above the contracted floor on a random sample reviewed by a native speaker.
  • Token count within an agreed tolerance of the contracted volume, measured with your tokenizer after deduplication.
  • No residual direct identifiers in a sampled audit, using locale-specific patterns.
  • Valid UTF-8, NFC-normalized, no BOM, no empty lines, schema-valid on every record.

How SourceX handles non-English text requests

SourceX sources operational datasets from US companies on request; it does not hold non-English corpora in stock, and a request does not guarantee a match. US businesses with operations abroad may hold native-language support histories, sales correspondence, documents and engineering records, and SourceX looks for the businesses that hold the data a buyer describes. It does not source scraped web content or standalone contact lists.

Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. You can describe a non-English text need to SourceX; background on record types is on the multilingual enterprise models page, and supplier-side context is in do AI labs buy non-English business data and what if my data includes non-English content. For the wider cluster, start at text datasets for LLM training or the AI data hub.

Request licensed non-English text for pre-training

Tell SourceX the languages, locales, registers and volume you need, and it looks for US businesses that hold matching data. Each dataset is assessed for licensing permissions and delivered under a license that defines records, uses, term and delivery, and nothing is contracted until a supplier agrees. Start a buyer request.

Sources

  1. arXiv (HPLT project), "An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT v2)" (2025). https://arxiv.org/pdf/2503.10267
  2. arXiv (BigScience), "The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset" (2023). https://arxiv.org/pdf/2303.03915
  3. arXiv, "GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training" (2026). https://arxiv.org/html/2604.00920v1
  4. arXiv, "The German Commons: 154 Billion Tokens of Openly Licensed Text for German Language Models" (2025). https://arxiv.org/html/2510.13996v1
  5. arXiv (EleutherAI and collaborators), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
  6. TAUS, "TAUS data sale to boost multilingual LLMs" (2024). https://www.taus.net/resources/blog/taus-data-sale-to-boost-multilingual-llms
  7. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  8. arXiv, "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
  9. arXiv, "La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America" (2025). https://arxiv.org/pdf/2507.00999
  10. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  11. UK Intellectual Property Office, "Exceptions to copyright: Research" (2014). https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/375954/Research.pdf
  12. jsonlines.org, "JSON Lines". https://jsonlines.org/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data