Document AI data
Multilingual and Mixed-Script Business Document Data for OCR and Extraction
Quick answer
A multilingual document dataset for Document AI should contain real document images in your target languages and scripts, with OCR text, field-level key-value labels and a language tag on every field, not just on the page. Public sets such as XFUND cover seven languages of forms, and many "multilingual document" datasets are actually text-only corpora. For production extraction, buyers usually need licensed business documents, including mixed-language invoices and forms, labeled by native-reading annotators under one schema.
By SourceX Editorial · Updated
Why public multilingual document sets rarely cover production needs
Public sets prove a model can generalize across languages, but they rarely match the document types, scripts and noise of your production traffic. XFUND, released by Microsoft Research alongside the LayoutXLM model, contains human-annotated forms with key-value pairs in seven languages: Chinese, Japanese, Spanish, French, Italian, German and Portuguese. It follows the labeling style of FUNSD, which is 199 noisy scanned English forms annotated for text detection, OCR, layout and entity linking [1]. That makes XFUND a useful benchmark, but it is small per language and is not made of invoices, purchase orders, customs declarations or bank statements.
Search results for "multilingual document dataset" also mislead. Many hits are NLP corpora: DocHPLT, for example, aligns web document pairs across 50 languages but contains no page images, bounding boxes or layout [3]. The closest public image set in the same results, MMDoc, is described on its dataset card as pairing document images with OCR and translations across 10 language pairs under Apache 2.0 [2], which suits translation-aware OCR rather than key-value extraction.
Business-document benchmarks such as DocILE show what production-grade labels look like: field localization plus line-item recognition across thousands of annotated documents [4]. When you need that depth in Spanish, Vietnamese or Arabic, you generally have to license real documents or annotate your own. Our guide on annotating your own documents versus licensing pre-labeled ones covers that trade-off.
Label language per field, not per page
Mixed-language documents are the normal case in cross-border business, so a page-level language tag hides most of the failure modes. A US-printed shipping form filled in by hand in Spanish, a bilingual English/French invoice, or a Japanese purchase order with English part numbers each mix languages inside one page. If your schema records only page_language: es, you cannot measure whether the model fails on Spanish handwriting or on the English template text.
Ask for these attributes on every labeled field or text region:
langas a BCP 47 tag (for examplees-MX,zh-Hant,ar), so regional variants are explicit.scriptas an ISO 15924 code (Latn,Cyrl,Arab,Hebr,Hans,Hant,Jpan,Deva), since script drives OCR difficulty.direction(ltr,rtl, or mixed) for Arabic and Hebrew lines that embed Latin numbers or SKUs.modality(printed, handwritten, stamp) because handwriting in a second language behaves differently from print.normalized_valuealongside the verbatim transcription, so dates such as09/10/2026resolve to an ISO date under the document's locale.
The last point matters more than teams expect. Decimal commas, day-first dates, and currency placed after the amount (1.234,56 €) break extraction evaluation when gold values are compared as raw strings. For more on field schemas, see key-value extraction labels.
Script and font coverage drives OCR error more than language count
For OCR, the unit of coverage is the script and its rendering conditions, not the language name. Two Latin-script languages share most glyphs, while diacritic-heavy languages such as Vietnamese, Polish or Czech fail when accents are dropped by low-resolution scans or fax compression. CJK documents add vertical text, full-width punctuation and very large character inventories, and Arabic adds contextual letter shapes, right-to-left reading order and ligatures.
Common failure modes to test for explicitly:
- Diacritic stripping:
Pérezread asPerezsilently breaks entity matching and deduplication. - Reading-order inversion: RTL text lines read left to right, or embedded LTR numbers reversed inside RTL lines.
- Mixed-width confusion: full-width digits in Japanese invoices read as different characters from half-width digits.
- Unicode normalization drift: NFC and NFD forms of the same accented character fail string equality in gold comparisons.
- Font and stamp occlusion: red hanko seals or company chops overlapping printed totals in Chinese and Japanese documents.
Ask any supplier for a coverage table broken down by script, modality and document type, and for character error rate measured separately per script. Our page on OCR ground truth data explains transcription alignment conventions that apply here too.
Annotators must read the language
Labeling quality in non-English documents depends on annotators who can read the language and know local document conventions. An annotator who cannot read Thai will draw plausible boxes but mislabel which value is the tax ID and which is the branch code. In Spanish-speaking markets, fields such as RFC (Mexico), NIF (Spain) or CUIT (Argentina) look similar but carry different validation rules.
Budget native-speaker QA as a separate pass, not a sample check by the original annotator. A practical pattern is double annotation on a stratified slice of each language and script, with inter-annotator agreement reported per field. Store the annotator's language proficiency and the QA reviewer ID in the label metadata so audit and documentation formats such as Croissant-RAI can carry it [5].
Native documents versus translated or synthetic ones
Translated or synthetic multilingual documents help with coverage, but they lose the layouts, abbreviations and handwriting that real local businesses produce. A machine-translated English invoice keeps an English template, English field order and US date format, so it trains the model on a document that does not exist in the target market. Generated documents also tend to use clean fonts and miss stamps, carbon copies and local tax layouts.
Use synthetic data to stress specific scripts or fonts, then evaluate on real documents only. The trade-offs are covered in synthetic vs real documents, and the parallel question for text appears in native vs translated text in multilingual training.
Request template for a multilingual document dataset
A precise request names languages, scripts, document types and the mix of mixed-language pages you expect. The template below shows a structure buyers can adapt.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example entry |
|---|---|
| Target task | Key-value extraction and line items for AP automation; held-out evaluation set |
| Languages and variants | es-MX, es-US, pt-BR, vi, zh-Hant, ar |
| Scripts | Latn, Hant, Arab; minimum share of handwritten fields per script |
| Document types | Supplier invoices, purchase orders, delivery notes, remittance advice |
| Mixed-language share | At least 20% of pages with two or more languages; English template plus local-language entries |
| Capture conditions | Scans at 200-300 dpi, phone photos, fax-degraded pages; native PDFs flagged separately |
| Label schema | Field boxes with lang, script, direction, modality, verbatim and normalized_value |
| Formats | Page images (PNG/TIFF), OCR in hOCR or PAGE XML, labels in JSON Lines |
| QA | Native-speaker second pass; agreement reported per language and field |
| Privacy | Names, account numbers, tax IDs replaced; method documented per language |
| Origin | Country of the issuer and of the recipient recorded per document |
An illustrative label record for one mixed-language field might look like this:
Illustrative example: invented to show structure; it does not describe an available dataset.
{"doc_id": "inv-000123", "page": 1, "field": "total_amount",
"bbox": [812, 1440, 1010, 1478], "lang": "es-MX", "script": "Latn",
"direction": "ltr", "modality": "handwritten",
"text": "$12,450.00 M.N.", "normalized_value": {"amount": 12450.00, "currency": "MXN"},
"template_lang": "en-US", "qa": {"native_review": true}}
Privacy and cross-border questions for non-US documents
Documents issued outside the US can bring foreign data protection rules into the deal even when the holder is a US company. A US distributor's archive may contain invoices from Mexican or EU suppliers with personal names, tax IDs and bank details, and some of those jurisdictions restrict transfers or reuse of personal data. Record the issuer country per document and ask counsel which rules apply before you set scope.
Redaction is harder in multilingual documents, because named-entity detectors trained on English miss names in other scripts and local ID formats. Test redaction recall per language, and check the OCR text layer and PDF metadata as well as the pixels. Our guide on redacting PII in scanned documents details those layers.
How SourceX approaches multilingual document requests
SourceX sources operational datasets from US companies, including documents and finance workflows, and manages the commercial process, including licensing agreements. Data is sourced on request rather than held in stock, so a described language and script mix does not guarantee a match. Buyers describe the documents they need, SourceX looks for US businesses that hold them, and every release is approved by the supplying company.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe your multilingual document requirements to SourceX as a starting point. For broader context, see training data for multilingual enterprise models and whether AI labs buy non-English business data.
Source multilingual business documents for OCR and extraction
SourceX sources documents and finance workflow data from US companies on request, with rights review, recorded de-identification and a license defining records, uses, term and delivery. AI teams wherever they are based can describe the languages, scripts and document types they need, and nothing is contracted until a supplier agrees. Start at SourceX for AI data buyers.
More guides in the Document AI data hub and the AI data hub.
Frequently asked questions
Is XFUND enough to train a production multilingual extractor?
Usually not on its own. XFUND covers forms in seven languages with key-value labels, which suits benchmarking cross-lingual transfer, but it does not cover invoices, line items or right-to-left scripts. Check its license terms in the official release before commercial use.
Should I evaluate per language or per script?
Report both. Language-level scores show business coverage, while script-level character error rate isolates OCR problems such as diacritics, RTL ordering or CJK confusions. Keep mixed-language pages as their own slice, as described in document extraction evaluation sets.
Can US companies hold useful non-English documents?
Often yes. Importers, logistics firms, and companies serving Spanish-speaking customers can hold supplier invoices, customs paperwork and customer forms in other languages, frequently mixed with English templates. See trade document extraction data for related document types.
Sources
- arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
- Hugging Face (rileykim/MMDoc), "MMDoc dataset card (README)". https://huggingface.co/datasets/rileykim/MMDoc/blob/main/README.md
- arXiv, "DocHPLT: A Massively Multilingual Document-Level Translation Dataset" (2025). https://arxiv.org/pdf/2508.13079
- arXiv (Šimsa et al.), "DocILE Benchmark for Document Information Localization and Extraction" (2023). https://arxiv.org/pdf/2302.05658
- arXiv (Jain et al., MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.