Retrieval, RAG and grounding data
Building a private domain retrieval test collection: corpus snapshot, topics, qrels and nugget links
Quick answer
A domain-specific retrieval benchmark is a test collection you control: a frozen snapshot of documents from your vertical, topics written from real user needs, graded relevance judgments (qrels) that point at document IDs in that exact snapshot, and, if you score RAG answers, nuggets linked to the documents that support them. Public suites such as BEIR show that retriever rankings shift across domains [1], so a private collection built from production-like records is the most reliable way to choose a retriever for your workload.
By SourceX Editorial · Updated
Why public retrieval benchmarks mispredict your domain
Public benchmarks mispredict domain performance because retriever quality depends on the vocabulary, document structure and query style of the collection being searched. The BEIR study evaluated retrievers zero-shot across heterogeneous corpora and found that in-domain results were a poor predictor of how models ranked on other domains [1]. A dense model tuned on MS MARCO web passages can rank well on open-web questions and still lose to BM25 on contract clauses, part numbers, ticket shorthand or ICD-coded clinical notes.
Two further problems hit enterprise buyers. First, popular public sets leak into pretraining and fine-tuning corpora, so scores drift upward without real gains; the LiveBench authors document how test-set contamination can make a benchmark obsolete quickly [8]. Second, open-web corpora do not contain what your users search: internal wikis, support threads, engineering change orders, policy versions. That gap is why researchers have started building benchmarks around company-internal knowledge [9]. A held-out private collection solves both, provided it is never fed back into training.
If you are still deciding between a private set and a public one, the trade-offs are laid out in private evaluation sets vs public benchmarks.
The four components and how they bind together
A usable test collection has four artifacts that must share one set of identifiers: the corpus, the topics, the qrels and (optionally) the nuggets. TREC distributes judgments per document collection, and the qrels are only valid against the collection they were made on [3]. The TREC 2025 RAG track follows the same pattern, publishing corpus, topics, relevance judgments and nuggets as separate downloads over a single corpus [5].
| Component | What it is | Typical format | Binds to |
|---|---|---|---|
| Corpus snapshot | Frozen, versioned set of documents or passages | corpus.jsonl with _id, title, text, metadata | Snapshot ID and content hashes |
| Topics | Information needs with a short query and a narrative | queries.jsonl or TREC topic XML with num, title, desc, narr | Topic IDs |
| Qrels | Graded judgments of topic-document pairs | TREC topic_id 0 doc_id grade or BEIR query-id corpus-id score TSV | Topic IDs and corpus _ids |
| Nuggets (optional) | Atomic facts a good answer must contain | nuggets.jsonl with nugget ID, text, importance, support_doc_ids | Topic IDs and corpus _ids |
The most common failure is silent ID drift: a re-chunking job renames passage IDs, the qrels still reference the old ones, and recall drops to near zero for reasons that have nothing to do with the retriever. Store a manifest with the snapshot ID, the chunking parameters and a SHA-256 per document, and reject any evaluation run whose corpus hash does not match. The mechanics of pinning are covered in pinned corpus snapshots for RAG evaluation.
Freezing the corpus snapshot
Freeze the corpus at the document level before a single judgment is made, and treat every later change as a new collection version. Judgments refer to a specific document collection [3]; if you add, delete or edit documents afterward, unjudged documents enter the ranking and judged ones may vanish, which biases metrics such as recall and nDCG.
Practical rules for the snapshot:
- Keep realistic noise. Include superseded policy versions, drafts and near-duplicate tickets, because production retrievers must rank the current version above stale ones. See distractor and near-duplicate documents.
- Keep the metadata your filters use. Effective dates, document type, product line, access group and source system often decide relevance; the RAG corpus metadata requirements page lists fields worth requiring.
- Store both raw documents and passages. Keep the source file (PDF, HTML, DOCX, ticket JSON) plus the passage split, so you can re-chunk later without re-judging from scratch.
- De-identify without breaking matches. Replacing account numbers with random tokens can destroy the exact-match queries users actually issue; use consistent pseudonyms. See retrieval-preserving de-identification.
Size the corpus so it is large enough to contain real distractors but small enough to judge sensibly. There is no standard size; what matters more is that the corpus mirrors the mix of sources your production index will hold, such as the multi-source enterprise search corpora that combine wikis, chat and tickets.
Sourcing topics from real user needs
Topics drawn from real user behavior, such as search logs, support tickets, escalation notes or analyst requests, predict production performance better than questions invented by annotators reading the corpus. Annotator-written questions tend to reuse document wording, which inflates lexical-match scores and hides the vocabulary gap that dense retrievers are supposed to close (this is our working hypothesis, consistent with BEIR's finding that dataset construction affects which models win [1]).
Write each topic in three layers, following TREC convention: a short title query as a user would type it, a desc sentence stating the need, and a narr paragraph defining what counts as relevant and what does not. The narrative is what keeps assessors consistent. Aim for 50 to 150 topics stratified by query type: known-item lookup, procedural how-to, troubleshooting, policy or eligibility, and multi-document comparison.
Keep topic sources out of anything you train on. If topics come from tickets, record the ticket IDs and exclude those tickets and their resolutions from fine-tuning or embedding training data.
Building graded qrels by pooling
Graded qrels are built by pooling: run several diverse retrievers over the frozen corpus, take the union of each run's top-k documents per topic, and have assessors judge that pool on a graded scale [6]. Diversity in the pool matters more than depth. Combine BM25, at least two dense or learned-sparse models, a hybrid run and, if possible, a reranked run, so the pool is not biased toward any single model you are about to evaluate.
Use a graded scale with written anchors, for example 0 not relevant, 1 related but does not answer, 2 partially answers, 3 fully answers. Graded labels let you report nDCG@10 alongside Recall@100 and MRR. Assessor guidance is covered in relevance assessment guidelines and graded scales, and the cost and quality trade-off of model-generated labels in LLM relevance labels vs human assessors.
Open tooling such as trectools can build pools from run files and score runs against qrels [7]. If you only need the judgment layer on a corpus you already hold, the narrower procurement question is covered in buying relevance judgments (qrels).
Two checks before you trust the qrels: double-judge 10 to 20 percent of pairs and report agreement (Cohen's kappa or Krippendorff's alpha), and after adding a new retriever, measure how many of its top-10 documents are unjudged. A high unjudged rate means the pool needs another round.
Linking nuggets to supporting documents
If you score RAG answers with nuggets, each nugget must carry the IDs of the corpus documents that support it; otherwise the nuggets measure answer content but cannot tell you whether retrieval found the evidence. The TREC 2025 RAGTIME overview notes that nuggets extracted from pooled documents without supporting-document annotations are limited for retrieval evaluation [4]. Public tracks also ship qrels and nuggets as separate files over one corpus [5], so join both to the same frozen snapshot ID.
With support links in place you can compute nugget-level evidence recall (the share of vital nuggets for which at least one supporting document appears in the top-k) alongside answer-level nugget coverage. That separates retrieval failures from generation failures. Record the passage offsets, not just the document ID, when documents are long.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"snapshot_id": "claims-kb-2026-09-15-v1", "topic_id": "T042",
"title": "water damage exclusion slow leak",
"narr": "Relevant: policy text or adjuster guidance stating whether gradual leaks over 14 days are covered. Not relevant: sudden pipe burst guidance.",
"qrels": [{"doc_id": "pol-HO3-2024-s4.2", "grade": 3},
{"doc_id": "kb-adj-0817#p3", "grade": 2},
{"doc_id": "pol-HO3-2019-s4.2", "grade": 1}],
"nuggets": [{"nugget_id": "T042-N1", "importance": "vital",
"text": "Continuous or repeated seepage over 14+ days is excluded",
"support_doc_ids": ["pol-HO3-2024-s4.2"],
"support_spans": [[1180, 1342]]}]}
Note how the superseded 2019 policy is judged 1, not 3: the collection rewards a retriever that surfaces the current version.
Sourcing checklist for a private collection
A buyer sourcing a domain test collection should settle rights, provenance and format before any annotation starts. Use this checklist when scoping a request with any supplier or internal data owner.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Item | What to require |
|---|---|
| Corpus provenance | Source systems, date range, record types, and who approved release |
| Rights | Written permission for evaluation use; check whether any public component you mix in requires its own license, as some BEIR datasets do [2]; see public retrieval datasets and commercial-use licenses |
| Snapshot manifest | Snapshot ID, document count, chunking parameters, per-document hashes |
| Topics | Origin (logs, tickets, requests), stratification by query type, held-out guarantee |
| Qrels | Pool construction, scale definitions, assessor qualifications, agreement scores |
| Nuggets | Importance labels, support_doc_ids, span offsets |
| De-identification | Method used, fields affected, whether pseudonyms are consistent |
| Contamination controls | Exclusion of topics and judged documents from any training set |
| Allowed use | Whether the license is limited to evaluation or also permits training, as discussed in grounding license vs training license |
An evaluation-only scope can narrow the rights a supplier must grant, which may simplify the deal; that is a negotiating hypothesis, not a rule.
Where licensed operational records fit
Licensed operational records fit when the documents your retriever must search live inside companies, not on the open web: support knowledge bases and ticket histories, engineering records, contract and finance workflows. SourceX sources operational datasets from US companies on request and manages the licensing process; it holds no inventory, and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails and account numbers are removed or replaced before delivery with the method recorded, and every release is approved by the supplying company. You can describe the collection you need on the SourceX buyers page.
For packaged RAG question-answer sets rather than IR test collections, see RAG evaluation datasets from real company documents and evaluation datasets built from real business work. The broader map of retrieval data is on the retrieval and RAG data hub, and before licensing at scale you can pilot-test a content source for retrieval lift.
Request a domain retrieval test collection
Describe the documents, user needs and judgment depth your vertical requires, and SourceX looks for US businesses that hold matching operational data. Each dataset goes through rights review and is delivered under a license that defines the records, uses, term and delivery, and nothing is contracted until a supplier agrees. Start at sourcex.si/buyers.
Sources
- Thakur et al. (arXiv:2104.08663), "BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
- BEIR (beir-cellar, GitHub), "BEIR Wiki: Datasets available". https://github.com/beir-cellar/beir/wiki/Datasets-available
- NIST, Text REtrieval Conference (TREC), "Data - English Relevance Judgements". https://trec.nist.gov/data/reljudge_eng.html
- arXiv:2602.10024, "Overview of the TREC 2025 RAGTIME Track" (2026). https://arxiv.org/pdf/2602.10024
- NIST, Text REtrieval Conference (TREC), "TREC 2025 RAG Track data" (2025). https://trec.nist.gov/data/rag2025.html
- arXiv:2304.12367, "Overview of the TREC 2022 NeuCLIR Track" (2023). https://arxiv.org/pdf/2304.12367
- GitHub (lgienapp/trectools), "trectools: an open-source Python library for TREC-style evaluation". https://github.com/lgienapp/trectools
- White et al. (arXiv:2406.19314; ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- arXiv:2605.05253, "EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge" (2026). https://arxiv.org/pdf/2605.05253
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.