Text and language data
Native vs Translated Text in Multilingual Training Data
Quick answer
Natively authored text and translated text are different training inputs and should be specified, labeled and priced separately. Native text carries the target language's own syntax, idiom, register and regional knowledge; translated text, especially machine translation, carries source-language interference and English-centric content. For pre-training and evaluation, require native text with per-document origin labels. Accept human translation or post-edited MT only where you deliberately want parallel signal, and keep LLM-generated expansions out of anything you call "native."
By SourceX Editorial · Updated
Why translated text behaves differently in training
Translated text is a shifted distribution, not a noisy copy of native text. Translation studies describe "translationese": calqued constructions, source word order, simplified vocabulary, explicitation and normalized register. A model trained heavily on it learns those patterns as if they were the target language, which shows up as stilted output, wrong formality levels and missing local references.
Web crawls make this worse in practice. Content that appears in many languages at once is often machine translated, and lower-resource languages have little native web text to dilute it. Treat it as a working hypothesis that a "native" crawl in a lower-resource language is partly English content run through MT, and test that on your own sample.
Translation also imports the source culture. A Spanish rendering of a US support ticket still describes ZIP codes, 401(k) plans and US holidays. That is fine for parallel training and harmful when you are trying to teach a model how a Madrid or Monterrey business actually writes.
Where translated or generated text is acceptable
Translated text is acceptable when the translation itself is the signal you are buying; it is a poor substitute when you need native distribution. Use the matrix below when writing your specification. For a broader map of the data types (parallel corpora, translation memories, termbases, post-edits), see multilingual training data types compared.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Use | Native text | Human translation | Post-edited MT | Raw MT | LLM-generated or expanded |
|---|---|---|---|---|---|
| Monolingual pre-training | Preferred | Small share, labeled | Small share, labeled | Avoid | Avoid unless separately budgeted |
| MT or cross-lingual alignment training | Useful as target side | Preferred (parallel) | Acceptable with QE scores | Avoid | Avoid |
| SFT / instruction tuning | Preferred | Acceptable if localized, not literal | Acceptable with review | Avoid | Only with provenance records |
| Evaluation and benchmarks | Required | Only for translation-quality tests | Avoid | Never | Never |
| Tokenizer or vocabulary extension | Required | Avoid | Avoid | Avoid | Avoid |
Evaluation is the strictest row. Benchmarks built by translating English resources inherit English-centric content and miss regional and cultural knowledge, so prefer test items written natively in each region, such as local exam or workplace material. Work on Spanish regional varieties shows a related risk: expert-reviewed test questions probe whether models identify regional varieties and favor particular dialects [3]. A translated or homogenized evaluation set cannot surface that kind of bias, because it contains no real regional variation to test against. If your evaluation set is translated, your scores mostly measure performance on translationese.
For instruction data specifically, the Aya project separated human-written annotations from templated and translated collections, which is the labeling discipline you want from any supplier [1]. Our page on native vs translated SFT data covers instruction tuning in depth.
The LLM-generated share is the new provenance gap
New multilingual resources increasingly mix human seed text with LLM-generated expansions, so ask for the generated share explicitly. For example, dialect-parallel Spanish work has expanded seed text from public-domain sources using LLM-based generation [2]. That is a legitimate research technique, but the output is not native authorship and should never be delivered under a "human-written" label.
Ask suppliers three questions: what fraction of documents or tokens was generated or rewritten by a model, which model and prompts produced it, and how generated records are flagged in the data. Treat an answer of "none" as a claim to verify, not a fact. Our guides to synthetic data provenance records and detecting model-generated content in purchased "human" data describe what a complete record looks like.
Per-document fields to require from suppliers
Origin should be a per-document field, not a dataset-level adjective. Provenance metadata is often missing or wrong even on popular dataset hosts [8], and standard dataset-card metadata records language but not whether text was authored in that language [9]. Add the fields below to your data specification and acceptance criteria; the general field list lives in metadata fields to require with licensed text corpora.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "kb-es-MX-000184",
"language": "es-MX",
"script": "Latn",
"origin_type": "native",
"original_language": "es-MX",
"translation_chain": [],
"generated_share": 0.0,
"generation_model": null,
"authoring_context": "internal support macro written by Mexico-based agents",
"created_at": "2023-04-17",
"lid_model": "glotlid",
"lid_label": "spa_Latn",
"lid_confidence": 0.98,
"review": { "method": "native-speaker spot check", "sampled": true }
}
Use a closed vocabulary for origin_type: native, human_translation, post_edited_mt, raw_mt, llm_generated, llm_expanded, mixed. For anything other than native, translation_chain should list each step in order, with method (human, engine name and version, LLM), source language and date. A document translated English to French to Wolof needs both hops recorded, because pivot translation compounds errors.
Language-code granularity matters. Use BCP 47 tags with a region where variety matters (pt-BR vs pt-PT, es-419 vs es-ES), ISO 15924 script codes for languages written in more than one script, and a separate flag for code-switched documents. See dialect and regional variety coverage for how to set regional quotas.
How to verify native authorship in a sample
Verification combines provenance evidence, automated screening and native-speaker review; no single signal is conclusive. The heuristics below are editorial guidance from practice, not validated thresholds, and should be calibrated on a sample where you know the answer.
Illustrative example: invented to show structure; it does not describe an available dataset.
Native-authorship verification checklist
- Source-system evidence. Ask where documents were created: a ticketing system, CRM notes field or document store used by staff working in that language is strong evidence. Exports from a CMS translation plugin, XLIFF or TMX files, or pages with
hreflangalternates point to translation. - Language ID agreement. Run an open LID model such as GlotLID, which targets low-resource languages, and compare its label with the declared tag [5]. Mismatches and low confidence on short documents are where mislabeled or mixed text concentrates; our language identification QA guide covers thresholds.
- Cross-language near-duplicates. Embed a sample with a multilingual sentence encoder and look for documents that align closely with English (or other-language) documents in the same delivery. Content that is parallel across many languages is a strong MT signal.
- Locale artifacts. Check dates, currencies, units, phone formats, legal entities and holidays. A German invoice note with MM/DD dates, dollar amounts and US state names was probably written in English first.
- Translationese markers. Native reviewers should flag calques, unnatural word order, over-explicit connectives, uniform formality and English named entities left untranslated. Compare lexical diversity (for example type-token ratio on fixed-length windows) against a known-native reference set.
- Generated-text screening, with caution. AI-text detectors are unreliable on writing that departs from typical native style; one study found them biased against non-native English writers [6], and translated text may trigger similar false positives. Use detector scores to prioritize review, never as an acceptance gate.
- Native-speaker review on a stratified sample. Sample by language, variety, source system and document length. Where translation quality itself is in scope, use professional translators with an MQM-style error typology rather than crowd ratings [7].
Record the sample size, method and failure rate in your acceptance report, and tie remediation (relabeling, removal or price adjustment) to defined failure rates in the license schedule.
Should you translate English datasets yourself?
Translating your own English data is a reasonable bootstrap for cross-lingual transfer, but it does not replace native data and should be labeled as translated throughout your pipeline. It is cheapest for instruction formats and alignment pairs, and least useful for evaluation, tokenizer work and domain register.
If you do it, keep the English source ID on every translated record, store the engine and version, run quality estimation, and cap the translated share per language in each training mixture so ablations can show its effect. Buying professional human translation is a separate product, typically sold by language pair as translation memories or parallel collections [4]; see licensing translation memories if parallel data is what you need.
For enterprise models that must read and write real business correspondence, native operational text is usually the gap. The multilingual enterprise models page describes that requirement, and licensing non-English text corpora covers the rights side. Back to the text and language data hub.
Sourcing natively authored business text through SourceX
SourceX sources operational datasets from US companies on request, including support and sales histories, documents and finance and legal workflows; nothing is held in stock and a request does not guarantee a match. Describe the language, variety, origin type and fields you need, and SourceX looks for US businesses that hold that data, with every release approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery. You can describe your multilingual text requirement as a specification like the one above.
Request natively authored non-English text
SourceX sources operational text from US companies on request and manages the process from assessment of data and licensing permissions through agreement, transaction and ongoing management. Nothing is contracted until a supplier agrees. Share your origin, language and metadata requirements at SourceX for buyers.
Frequently asked questions
Is post-edited machine translation "native" text?
No. Post-editing fixes errors but usually keeps the MT output's structure and source-language word order, so label it posteditedmt and keep it separate from text authored in the target language.
How much translated text can a pre-training mix tolerate?
There is no published universal threshold. Cap the translated share per language, keep it labeled, and run ablations on native evaluation sets such as region-sourced exams to measure the effect on your own model.
Can a supplier prove text was never translated?
Not with certainty. Strong evidence is a source system used by staff working in that language, creation timestamps, absence of cross-language near-duplicates and a clean native-speaker sample review; ask for all four.
Sources
- Singh et al., arXiv:2402.06619, "Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning" (2024). https://arxiv.org/pdf/2402.06619
- ACL Anthology, "Jessica Claribel Ramirez Vidal (ACL Anthology author page)". https://aclanthology.org/people/jessica-claribel-ramirez-vidal/
- Universidad Politecnica de Madrid open archive, "Regional Spanish variety identification and dialect bias in LLMs (data article)" (2025). https://oa.upm.es/91162/3/S2352340925008108.pdf
- TAUS, "Data for AI". https://www.taus.net/data-solutions/data-for-ai
- Kargaran et al., arXiv:2310.16248 (Findings of EMNLP 2023), "GlotLID: Language Identification for Low-Resource Languages" (2023). https://arxiv.org/pdf/2310.16248
- Liang et al., arXiv:2304.02819, "GPT detectors are biased against non-native English writers" (2023). https://arxiv.org/pdf/2304.02819
- Freitag et al., arXiv:2104.14478 (TACL 2021), "Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation" (2021). https://arxiv.org/abs/2104.14478v1
- Longpre et al., arXiv:2310.16787, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.