Fine-tuning and post-training data
Non-English instruction data: native vs translated SFT data
Quick answer
Translate English SFT data when you need broad coverage fast and the target behavior is format-following; source native-language data when your assistant must sound natural, handle local conventions, and survive real user phrasing such as code-switching. Many teams end up with a hybrid: native prompts collected from real interactions, responses written or post-edited by fluent experts, and a thin translated layer for coverage. The deciding factors are your evaluation set, your languages' resource level, and whether native sources can be licensed.
By SourceX Editorial · Updated
Translated SFT data is cheap coverage with predictable failure modes
Machine-translated instruction data reliably teaches task format, but it carries English-centric content and "translationese" into every target language. The research record reflects this: the authors of a 2025 TACL paper note that large-scale multilingual instruction data remains limited and that many efforts simply extend English data by automatic translation [2]. Translation preserves the prompt distribution of the source set, so your Spanish or Thai model learns to answer questions English-speaking (often US) users ask, about their institutions, in their registers.
The failure modes are specific and testable:
- Cultural transfer errors. Prompts about 401(k) plans, ZIP codes or Thanksgiving survive translation intact and teach the model the wrong defaults. Human evaluation work on LLM translation argues that grammar- and token-level benchmarks miss exactly this kind of cultural nuance [7].
- Register drift. Formal and informal address (tú/usted, du/Sie, tu/vous, Japanese keigo) are often chosen inconsistently by MT, so the fine-tuned assistant switches register mid-conversation.
- Answer-prompt mismatch. Code, units, dates and named entities get translated or localized in the prompt but not in the response, producing pairs where the "correct" answer is wrong.
- Lost code-switching. Real users in Manila, Mumbai or Miami mix languages within a sentence; translated sets are almost always monolingual per turn.
Translation is still the right tool for some layers: tool-call schemas, JSON-output tasks and safety refusals, where the structure matters more than idiom. See structured-output fine-tuning data for that layer.
Native instruction data captures how people actually ask
Native data is written or spoken originally in the target language, so prompt distribution, phrasing and cultural assumptions come from real speakers rather than from an English template. The clearest open example is the Aya Dataset, which collected human-written instruction-completion pairs from native speakers in 65 languages, released alongside the much larger Aya Collection of 513 million templated and translated instances across 114 languages [1]. That pairing is itself a useful lesson: a small native core and a large translated or templated shell are complements, not substitutes.
Native data matters most when:
- Your evaluation set comes from real users in that market, not from translated benchmarks.
- The language has strong regional variation. Spanish projects such as #Somos600M build instruction resources across varieties from Latin America, the Caribbean and Spain [5], and La Leaderboard evaluates Spanish varieties and languages of Spain and Latin America separately [6].
- Domain vocabulary is local: tax forms, legal procedure, payment rails and product names differ by country even within one language.
The native vs translated text guide covers the same trade-off for pretraining and general text; this page is about instruction and preference pairs.
Volume matters less than curation in non-English SFT
A few thousand well-reviewed native pairs per language can do more than a much larger translated set, because, as the LIMA authors argue, SFT mostly teaches format and style on top of what pretraining already knows. LIMA fine-tuned a 65B LLaMA model on just 1,000 carefully curated prompt-response pairs [10]; the authors also note that such curation is labor-intensive. For multilingual work that labor multiplies by language, which is why buyers should budget review time per language rather than per record.
Large open multilingual sets are best treated as starting material. MURI-IT, for instance, lists 2.2 million instruction-output pairs across 200 languages [3], and M3IT extends instruction tuning to multimodal, multilingual tasks [4]. Filter them with the methods in instruction-tuning data quality filtering before mixing them with native data.
Decision table: translate, source native, or go hybrid
The right mix depends on the job each slice of data does in your post-training stack. Use the table below per language, not once for the whole project.
| Situation | Translate English SFT | Source native data | Hybrid (native prompts, expert responses) |
|---|---|---|---|
| Format-following, JSON, tool calls | Good fit | Unnecessary | Optional |
| Customer-facing chat in a high-resource language | Weak alone | Strong | Strongest |
| Regional variety matters (es-MX vs es-ES, pt-BR vs pt-PT) | Poor | Strong | Strong |
| Low-resource language, little native text exists | Only option at scale | Hard to find; check licensing early [12] | Native prompts plus expert-written responses |
| Domain tasks (support, finance, legal) | Wrong defaults | Strong if records are licensable | Strong |
| Preference or DPO pairs | Weak; translated "chosen" responses inherit errors | Strong | Strong with native raters |
The hybrid column reflects a working hypothesis many teams share rather than a settled result: real native prompts fix the input distribution, and fluent experts writing or post-editing responses fix quality and policy alignment. If you already hold aligned bilingual assets, translation memories can seed parallel tasks, and multilingual training data types compared explains how they differ from native text.
Where native-language instruction data comes from
Native instruction data usually comes from three places: open community datasets, commissioned writing, and operational records from businesses that already serve customers in that language. Each has a different rights profile.
- Open datasets (Aya, #Somos600M, MURI-IT) are fast to start with, but check the license on every component. The Data Provenance Initiative audited more than 1,800 text datasets and found license omission rates above 70% and error rates above 50% on popular hosting sites [9]. The open fine-tuning datasets for commercial use page lists what to verify.
- Commissioned writing by native-speaker annotators gives you control over register, domain and dialect, but output can drift toward a "writer's voice" that does not match real users.
- Operational records such as Spanish-language support tickets, chat transcripts, sales emails, or internal knowledge-base articles carry genuine phrasing, code-switching and domain vocabulary. They need conversion into instruction-response pairs, as covered in turning business records into instruction-response pairs, plus de-identification and a license that permits fine-tuning.
Many US companies run support, sales and back-office work in Spanish and other languages, which is why buyers ask whether AI labs buy non-English business data. If that is the source you need, you can describe it on the SourceX buyer request page.
Multilingual preference data needs native raters, not translated rankings
Preference pairs encode judgments of naturalness, politeness and helpfulness, and those judgments are language- and culture-specific. DPO fits a policy directly to chosen-versus-rejected pairs without a separate reward model [11], so any bias in the pairs passes straight into the policy. Translating English preference data keeps the English raters' choices, including their tolerance for directness or verbosity, and attaches them to text they never read.
For each language, ask for rater metadata: native language, country, years of fluency, and the rubric they applied. Require inter-rater agreement per language, because a single overall agreement number can hide one language where raters disagree constantly. See rubric-based rewards data for rubric design.
Acceptance checks for a multilingual SFT delivery
Run automated language and script checks first, then native-speaker review on a stratified sample per language. Language ID is not trivial for low-resource languages: GlotLID was built precisely because common identifiers misclassify many of them [8]. Run it on both prompt and response, and flag pairs where they disagree.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_id": "es-mx-000412",
"lang_tag": "es-MX",
"script": "Latn",
"origin": "native_operational",
"source_type": "support_chat_transcript",
"prompt": "Oye, me llegó doble cargo en mi tarjeta, ¿me lo pueden reembolsar?",
"response": "Claro, con gusto lo reviso. ¿Me confirma la fecha del cargo y los últimos cuatro dígitos que aparecen en su estado de cuenta?",
"register": "usted",
"code_switching": false,
"response_author": "native_expert_post_edit",
"deid_method": "named entities replaced with typed placeholders",
"reviewer_lang": "es-MX",
"review_score": 4
}
Illustrative example: invented to show structure; it does not describe an available dataset.
Delivery acceptance checklist (per language)
- BCP 47 language tag and ISO 15924 script code on every record; reject records where detected language disagrees with the tag.
originfield distinguishes native, translated, post-edited and synthetic records, with counts per language.- Translationese spot check: a native reviewer flags calques and unnatural word order on a random sample.
- Register consistency: formal and informal address does not switch within a conversation.
- Locale correctness: currency, date formats, phone formats and institutions match the tag.
- Code-switched turns are labeled, not stripped.
- De-identification method is documented and checked on a sample.
- License covers fine-tuning, derivative models and the languages delivered; confirm with the fine-tuning-only data license guide.
- Hold out a native-written evaluation set per language that never enters training.
The general checks in evaluating a fine-tuning dataset before you buy still apply on top of these.
Sourcing native-language instruction data with SourceX
SourceX sources operational datasets, such as support and sales histories, from US companies on request and manages licensing; it holds no stock, so a request does not guarantee a match. Every dataset is rights-reviewed, personal details are removed or replaced before delivery, and each release is approved by the supplying company. Describe the languages, varieties and interaction types you need on the SourceX buyer page.
For the wider cluster, start at the fine-tuning and post-training data hub, the SFT sourcing guide, or training data for multilingual enterprise models. Background on the term: instruction tuning.
Frequently asked questions
Is post-edited machine translation "native" data?
No. Post-editing fixes errors, but the prompt distribution and topic choices remain those of the English source. Label post-edited records separately and weight them below truly native ones.
How much native data do I need per language?
There is no fixed number. Start with a native evaluation set, add native SFT in increments, and stop when held-out native-user metrics stop improving. Curated sets in the low thousands can move behavior noticeably [10].
Can I use one Spanish dataset for every Spanish-speaking market?
Only if your users tolerate a neutral variety. Resources such as La Leaderboard evaluate varieties separately [6], which is a reason to test per market before assuming one set covers all.
Sources
- arXiv (Cohere For AI community), "Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning" (2024). https://arxiv.org/pdf/2402.06619
- Köksal et al., "MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions" (TACL, 2025). https://aclanthology.org/2025.tacl-1.48/
- Köksal et al., "MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions" (TACL, 2025). https://github.com/akoksal/muri
- arXiv, "M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning" (2023). https://arxiv.org/pdf/2306.04387
- arXiv (SomosNLP), "The #Somos600M Project: Generating NLP resources that represent the diversity of the languages from LATAM, the Caribbean, and Spain" (2024). https://arxiv.org/pdf/2407.17479
- arXiv, "La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America" (2025). https://arxiv.org/pdf/2507.00999
- Madison Van Doren et al. (ACL Anthology, GEM 2026), "“Be My Cheese?”: Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs". https://aclanthology.org/2026.gem-main.6/
- arXiv (EMNLP Findings 2023), "GlotLID: Language Identification for Low-Resource Languages" (2023). https://arxiv.org/pdf/2310.16248
- Longpre et al. (Nature Machine Intelligence 6, 2024), "A large-scale audit of dataset licensing and attribution in AI". https://www.nature.com/articles/s42256-024-00878-8
- arXiv (Zhou et al.), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- arXiv (Rafailov et al.), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- arXiv, "Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo" (2025). https://arxiv.org/pdf/2501.11003
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.