Text and language data
Multilingual Training Data Types Compared: Native Text, Parallel Corpora, TMs, Termbases and Post-Edits
Quick answer
Buy multilingual training data by capability, not by language count. Native monolingual text builds fluency in a language; parallel corpora and translation memories (TMs) teach translation; termbases and post-edited machine translation fix terminology and domain register; native instruction and dialogue data shape multilingual chat. Most teams need a mix, and the binding constraint is often rights rather than volume: much public bitext forbids commercial models trained on it. Decide the capability first, then the data type, then check the license.
By SourceX Editorial · Updated
This page sits in the text and language data hub and compares the five main types side by side. Each type has its own deep-dive, linked below, and translation memories have a dedicated owner page at license translation memories for AI training.
Which data type fixes which capability
Each multilingual capability fails for a different reason, so each needs a different data type. A model that writes stilted Polish has a fluency gap that bitext will not close. A model that translates fluently but renders "Kündigungsfrist" inconsistently has a terminology gap that more monolingual German will not close either.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Capability gap | Primary data type | Typical source | Main rights pitfall | Deep-dive |
|---|---|---|---|---|
| Fluency and world knowledge in language X | Native monolingual text | Publishers, enterprise document stores, open web-derived corpora | Web-derived sets carry unclear upstream licenses | Licensing non-English text corpora |
| General translation quality, pair A-B | Parallel corpus (sentence-aligned bitext) | Public crawls, LDC-style catalogs, commercial collections | Research-only terms that extend to derived translators | Parallel corpora licensing pitfalls |
| Domain translation (legal, medical, product) | Translation memories (TMX) | Enterprise localization teams and language-service providers | Client confidentiality clauses in the original translation contract | Translation memories owner page |
| Terminology adherence | Termbases (TBX) plus term-annotated segments | Enterprise terminology management systems | Termbase owned by the client, not the translation vendor | This page |
| Fixing systematic MT errors, preference data | Post-edits (MT output, human edit, edit metadata) | Localization workflows using MT plus human review | Engine provider terms on the raw MT output | This page |
| Multilingual chat and instruction following | Native-written instructions and dialogues | Commissioned writers, support and sales histories | Translated SFT data inherits translationese | Non-English instruction data |
Native monolingual text: the only route to real fluency
Fluency in a new or weak language comes from large volumes of native text, not from tokenizer tricks or translated data. Research on vocabulary expansion with continual pre-training reports that adding target-language tokens can degrade translation quality and still needs large target-language corpora to pay off [1]. If your model underperforms in Vietnamese or Swahili, the first purchase is native Vietnamese or Swahili text.
Open baselines exist. HPLT releases a large multilingual web-derived dataset [2], and BigScience's ROOTS is a 1.6TB composite multilingual corpus built from many sources [3]. These are useful for scale, but web-derived text skews toward a few registers, carries boilerplate and near-duplicates, and has upstream license uncertainty; deduplication alone measurably changes model behavior [9].
Licensed native text earns its cost where the web is thin: professional registers, enterprise documents, regulated domains and low-resource languages. See how much text you need to adapt an LLM to a new language for sizing and native vs translated text for why machine-translated "native" corpora underdeliver.
Parallel corpora: translation signal with the heaviest rights risk
Parallel corpora are the direct training signal for translation, but their licenses are the most likely of all multilingual types to block commercial use. As of October 2026, NTT's JParaCrawl, one of the largest public English-Japanese corpora, is licensed for research only, and its terms say translators trained on it are excluded from commercial use, with commercial licensing handled separately [4]. That clause reaches your model weights, not just the raw file.
The problem is systemic. The Data Provenance Initiative audited more than 1,800 text datasets and found license omission above 70% and license error rates above 50% on popular hosting sites [8]. A permissive tag on an aggregator mirror does not override the original publisher's terms.
Catalog and commercial routes exist. The Linguistic Data Consortium distributes domain bitext such as Chinese-English sentences extracted from patents under its own license agreements [5], and TAUS sells human-translation collections by language pair [6]. Before buying, ask for:
- Alignment method (manual, Bleualign-style, LASER or LaBSE embedding scores) and the score threshold used.
- Direction of original authorship per segment, because source-original and target-original text behave differently.
- Whether segments were machine-translated before human review.
- Deduplication against public crawls you already hold.
Translation memories: parallel data with provenance and domain
A translation memory is parallel data that a professional workflow produced, so it usually carries domain, client, date and reviewer context that crawled bitext lacks. TMs are typically exported as TMX, an XML format where each <tu> translation unit holds <tuv xml:lang="..."> variants plus attributes such as creationdate, changedate, creationid and custom <prop> fields for project or domain. That metadata is what lets you filter by recency, domain and translator.
The rights question differs from bitext. A language-service provider may hold the TM file while the end client owns the source content under the original services contract, so a TM sale needs authority from whoever owns the text. SourceX's insight on what makes translation memories valuable for AI covers quality signals; the translation memory owner page covers licensing.
Common TM failure modes to screen for: fuzzy-match leftovers stored as confirmed units, segments with tags (<bpt>, <ept>, <ph>) stripped inconsistently, untranslated segments copied source-to-target, and decades-old units with obsolete terminology.
Termbases: the fix for terminology accuracy
Termbases address a gap that more bitext does not close: strong models still get domain terms wrong. On the WMT23 terminology task, evaluated models reached only about 52-53% term accuracy for Chinese-English and 37-39% for German-English [7]. For regulated, legal or product content, that error rate is the main quality complaint.
A termbase is a concept-oriented glossary, usually exported as TBX (ISO 30042) or CSV from a terminology management system. Useful fields include concept ID, subject field, preferred and admitted terms per language, forbidden terms, part of speech, definition, context sentence and usage status. Forbidden-term entries are especially valuable because they give negative examples for constrained decoding, reward models or terminology-aware RL.
Termbases are small, so their value comes from pairing. Buy them alongside TMs or post-edits from the same organization, where the term decisions can be verified in real segments.
Post-edits: error-correction and preference data
Post-edited machine translation records a source segment, the raw MT output and the human-corrected version, which makes it a natural source of edit, preference and reward-model data. The pair (MT output, post-edit) is effectively a rejected-and-chosen example, and the diff shows exactly which errors humans fix. Small curated supervised sets can shape model behavior strongly [10], which suggests that, for a narrow domain, a modest volume of high-quality post-edits may be worth more than a much larger pile of crawled sentence pairs.
Ask for edit metadata: engine and version, edit distance or HTER, time spent, quality tier (light vs full post-editing) and reviewer ID. Without engine identity you cannot tell whether the errors represent your model's failure modes. Check the MT provider's terms on output reuse, which some enterprise contracts restrict.
Multilingual instruction and dialogue data
Multilingual chat quality depends on instructions and conversations originally written in each language, not English SFT data run through a translator. Translated instructions carry English discourse patterns, culturally misplaced examples and translationese, which models then imitate. Our guide to non-English instruction data compares native and translated SFT in detail.
Operational records are a strong native source: multilingual support tickets, chat transcripts and sales emails capture real register, code-switching and regional variety. Those records contain personal data, so require documented redaction and a language identification QA pass before training.
Building a purchase mix
A defensible mix starts from your evaluation gaps and allocates budget to the type that addresses each one. Run per-language evaluations first (fluency, translation by pair and domain, term accuracy, chat preference), then map each failing metric to a row in the table above.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: multilingual-mix-q4
capability_gaps:
- gap: "Weak fluency in pt-BR and id-ID legal register"
data_type: native_monolingual
format: JSONL, one document per line, fields [doc_id, lang, domain, source_type, created_at, license_ref]
- gap: "de-DE <-> en-US product translation, term errors"
data_type: [translation_memory, termbase]
format: [TMX 1.4b, TBX]
must_have: [creationdate, domain prop, forbidden terms]
- gap: "MT error correction for ja-JP support content"
data_type: post_edits
fields: [source, mt_output, mt_engine, post_edit, hter, pe_level]
rights_required: commercial training and model distribution, including derived translators
exclusions: research-only bitext, scraped web content, machine-translated "native" text
Keep each line tied to a measurable gap so you can verify value after training. For volume planning across pre-training sets see licensed text corpora for LLM pre-training, and for rights language see the pre-training rights grant you need.
How SourceX sources multilingual data
SourceX sources operational datasets from US companies, such as support and sales histories, documents, and finance and legal workflows, on request rather than from stock; a request does not guarantee a match. Buyers describe the data they need, every release is approved by the supplying company, and each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. SourceX does not source scraped web content. For multilingual enterprise use cases, see training data for multilingual enterprise models or describe your requirement to SourceX.
Source multilingual training data for your model
If your evaluation shows a gap that native text, TMs, termbases, post-edits or native dialogue data from US businesses could close, describe the data rather than the supplier. SourceX looks for companies that hold it, assesses data and licensing permissions, and nothing is contracted until a supplier agrees. Start a buyer request at SourceX.
Sources
- ACL Anthology, "An Information-Theoretic Approach to Reducing Fertility in LLMs for Manipuri Machine Translation" (2025). https://aclanthology.org/people/priyankoo-sarmah/unverified/
- arXiv, "An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)" (2025). https://arxiv.org/pdf/2503.10267
- arXiv, "The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset" (2023). https://arxiv.org/pdf/2303.03915
- NTT Communication Science Laboratories, "JParaCrawl". https://www.rd.ntt/cs/team_project/icl/lirg/jparacrawl/
- Linguistic Data Consortium, "Chinese-English Parallel Sentences Extracted from Patents (LDC2016T22)" (2016). https://catalog.ldc.upenn.edu/LDC2016T22
- TAUS, "Data for AI". https://www.taus.net/data-solutions/data-for-ai
- arXiv, "How Well Do Large Reasoning Models Translate? A Comprehensive Evaluation for Multi-Domain Machine Translation" (2025). https://arxiv.org/pdf/2505.19987
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- arXiv (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- arXiv (Zhou et al.), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.