Retrieval, RAG and grounding data
Cross-language retrieval datasets: sourcing queries, documents and bilingual judgments
Quick answer
A cross-language retrieval dataset pairs queries in one language with documents in another, plus relevance judgments made by assessors who can read the document language. For evaluation, the reference design is a TREC-style test collection: a fixed corpus snapshot, topics translated or written natively, pooled runs and graded qrels [1]. For training, you need query-passage pairs with language-matched hard negatives. Public sets cover Wikipedia and news well, while enterprise documents in mixed languages usually have to be sourced and judged under license.
By SourceX Editorial · Updated
What makes cross-language retrieval data different from monolingual sets
Cross-language data adds a second language to every judgment, so the cost and the failure modes move to the assessor and the topic translation. In a monolingual set, queries and passages share a language and native-speaker assessors judge within it. In cross-language information retrieval (CLIR), an English topic may retrieve Persian or Chinese documents, and the person judging must read the document, not only the query.
Three things change in practice:
- Topic development. Topics are written in one language and translated, or written natively per language. Translation drift (a narrower or broader term in the target language) silently changes what counts as relevant.
- Judgment. Assessors need reading proficiency in the document language and a stable understanding of the topic narrative in the query language.
- Pool bias. If most submitted systems translate the query with the same machine translation (MT) engine, documents that only a different translation would surface never enter the pool.
The TREC NeuCLIR track addressed this by pooling diverse runs and collecting graded relevance judgments on news documents in Chinese, Persian and Russian against English topics [1]. If your users are multilingual, treat that collection design as the minimum bar for an internal benchmark, then adapt it to your own corpus as covered in building a private domain retrieval test collection.
The four data types a multilingual retrieval program needs
Most programs need four distinct assets, and each one is sourced and licensed differently. Buying one undifferentiated "multilingual corpus" usually leaves the judgment layer missing.
| Asset | What it contains | Typical source | Main sourcing risk |
|---|---|---|---|
| Document corpus | Native-language documents with stable doc IDs, language tags, script, timestamps | Licensed business documents, news, manuals | Language ID errors, machine-translated pages posing as native text [4] |
| Queries / topics | Real user queries or written topics with title, description, narrative per language | Search logs, support tickets, assessor-written topics | Translated topics that do not match how users actually ask |
| Relevance judgments (qrels) | topic_id, doc_id, graded label, assessor_id, judging language | Bilingual assessors; pooled runs [1][6] | Shallow pools, untrained assessors, LLM labels without calibration |
| Training pairs and hard negatives | query, positive passage, mined negatives across languages | Public sets [3], licensed logs, synthetic generation | License limits, false negatives that are actually relevant |
The corpus and the judgments are the parts most often underbudgeted. For a broader comparison of native text, parallel corpora and translation memories, see multilingual training data types compared; this page focuses on retrieval-specific layers.
Building a CLIR test collection buyers can trust
A trustworthy CLIR test collection freezes the corpus, documents how topics were created per language, and pools diverse systems before bilingual assessors judge. The steps below follow the TREC pattern [1][6].
- Snapshot the corpus. Assign immutable doc_ids, record language and script per document, and store a hash so later re-runs compare like with like.
- Write topics natively where possible. If topics are translated, keep both versions and record the translator and method (human, MT plus post-edit).
- Pool diverse runs. Include BM25 with query translation, BM25 with document translation, a multilingual dense retriever and a reranker. Pool the top-k from each.
- Judge on a graded scale. NeuCLIR used graded relevance rather than binary labels [1]; record the scale definitions in the judging guide.
- Measure agreement. Double-judge a sample per language and report agreement by language, since quality often differs between high- and low-resource languages.
- Publish qrels separately. Keep qrels as their own versioned file, as TREC does [6], so you can add judgments without touching topics.
Budget for bilingual assessors as the dominant cost; this is a working assumption to price into any vendor quote rather than a published benchmark figure. LLM assessors can extend pools, but calibrate them against human labels per language before trusting them, as discussed in LLM vs human relevance labels. For sourcing the judgment layer itself, see buying relevance judgments (qrels).
Multilingual RAG evaluation: from ranked lists to grounded reports
Multilingual RAG evaluation tests whether a system can retrieve across languages and then write a correct, cited answer in the user's language. The TREC 2025 RAGTIME track is a current reference: systems drew on Arabic, Chinese, English and Russian documents to produce English reports, and evaluation relied on nuggets (atomic facts) linked to supporting documents [2].
For an internal multilingual RAG evaluation dataset, add three fields beyond standard qrels:
- nugget_id and nugget text in the report language, with links to supporting doc_ids in their original languages.
- citation_language, so you can see whether the generator over-cites English sources when the best evidence is in another language.
- answer_language_required, so a correct answer in the wrong language is scored as a failure.
A common failure mode is translation-induced hallucination: the system retrieves the right Russian passage, mistranslates a figure or negation, and cites it correctly. Nugget-level judgments by assessors who read the source language catch this; English-only answer review does not.
Multilingual hard negatives and training pairs
Multilingual hard negatives are passages that look relevant to the query but are not, mined in the same and other languages so the retriever learns meaning rather than surface overlap. Public resources exist: recent multilingual QA datasets ship mined hard negatives for dense retrieval [3], and Wikipedia-based multilingual retrieval sets provide human-judged positives within each language.
Three checks before you train on them:
- False negatives. Mined negatives from one language are often relevant documents nobody judged. Sample and review a few hundred per language with bilingual reviewers.
- Cross-lingual negatives. Many public sets only mine negatives within a language. If your users query in English against German contracts, you need English-query, German-document negatives, including translations of the positive document's near-duplicates. See superseded versions and near-duplicates for building those distractors.
- License and domain. Many public retrieval datasets carry non-commercial terms or unclear provenance; one large audit found license omissions above 70% on popular hosting sites [5]. Check each component against which public retrieval datasets allow commercial use.
Wikipedia-derived sets also underrepresent the vocabulary of invoices, service tickets and engineering notes. Domain shift, not language, is often the larger gap for enterprise search.
Sourcing non-English queries and documents from operating businesses
Enterprise-domain CLIR data usually comes from companies that already run multilingual support, sales or document workflows. Support ticket histories with an English knowledge base and customer messages in Spanish or Portuguese are natural cross-language query-document pairs, provided ticket-to-article links were recorded. Product manuals published in several languages give aligned document versions; product manuals and service documentation covers that corpus type.
Watch for these quality problems in sourced data:
- Language tags that lie. Web-crawled multilingual corpora frequently carry wrong or missing language labels [4]. Run your own language identification and compare it with the supplier's tags; language identification QA explains the checks.
- Machine translation presented as native text. Translated knowledge base articles inflate cross-language scores because their phrasing mirrors the English source.
- Code-switching and mixed script. Real support messages mix English product terms into other languages; keep them, but tag them.
- Personal data in queries. Customer messages contain names, emails and account numbers. Plan de-identification that keeps retrievable terms intact, as described in retrieval-preserving de-identification.
If you need parallel text to train translation components rather than retrieval data, translation memories are a separate purchase. For model training beyond retrieval, see training data for multilingual enterprise models.
Request specification for a cross-language retrieval dataset
A precise request names the language pairs, the direction of retrieval, the judgment design and the permitted uses, so suppliers can confirm fit before any sample. Document the delivered dataset with a datasheet covering composition, collection and recommended uses [7].
Illustrative example: invented to show structure; it does not describe an available dataset.
request: cross-language retrieval evaluation + training pairs
retrieval_direction:
- query_lang: en doc_langs: [es, pt-BR, de]
- query_lang: es doc_langs: [en]
corpus:
source_type: support knowledge base + resolved tickets
doc_fields: [doc_id, lang, script, lang_id_confidence, created_at, version, is_machine_translated]
snapshot: frozen, hashed
topics:
count_per_pair: to be agreed
creation: native-written, with human translation to query_lang
fields: [topic_id, title, description, narrative, source_lang]
judgments:
scale: 0 not relevant / 1 related / 2 partially answers / 3 fully answers
pooling: BM25+QT, BM25+DT, multilingual dense, reranker; top-k per run
assessors: reads doc_lang at professional level; double-judged sample per language
record: [topic_id, doc_id, label, assessor_id, judged_at, doc_lang]
training_pairs:
hard_negatives: same-language and cross-language, reviewed sample per language
rag_eval:
nuggets: report language en, linked to source doc_ids
privacy: names, emails, phone and account numbers replaced; method documented
license_scope: evaluation and retriever training; embeddings and index rights stated
A sample qrels row in TREC layout plus a language column might read T014 0 KB-de-00431 2 de. Keeping the document language in the row lets you report nDCG@10 per language pair rather than one blended number that hides weak pairs.
License terms specific to multilingual retrieval data
The license for retrieval data should name each use separately: evaluation, retriever training, embedding and indexing, and display in generated answers. Grounding rights and training rights are not interchangeable, as explained in grounding license vs training license. For multilingual sets, also confirm whether you may machine-translate documents for document-translation pipelines, since a translation can be a derivative work under the supplier's terms. Ask how vector indexes are handled when the license ends; see embedding and vector index rights.
How SourceX sources cross-language retrieval data
SourceX sources operational datasets from US companies on request, including support and sales histories and business documents, and manages the licensing process; categories are not inventory, and a request does not guarantee a match. You describe the data, such as language pairs, document types and judgment needs, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and delivery runs through private, access-controlled workflows under a license defining records, uses, term and delivery. Start with the buyer request form, or read the retrieval and RAG data hub first.
Describe the cross-language retrieval data you need
If public collections do not cover your language pairs or document types, describe the queries, documents and judgments you need. SourceX serves AI teams wherever they are based and moves each request through Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Describe your cross-language retrieval dataset request.
Sources
- arXiv (TREC NeuCLIR organizers), "Overview of the TREC 2022 NeuCLIR Track" (2023). https://arxiv.org/pdf/2304.12367
- arXiv (TREC RAGTIME organizers), "Overview of the TREC 2025 RAGTIME Track" (2026). https://arxiv.org/pdf/2602.10024
- arXiv, "WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval" (2026). https://arxiv.org/pdf/2602.17327
- arXiv (Kreutzer et al.), "Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets" (2021). https://arxiv.org/pdf/2103.12028
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- NIST Text REtrieval Conference (TREC), "Data - English Relevance Judgements". https://trec.nist.gov/data/reljudge_eng.html
- arXiv (Gebru et al.), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.