Retrieval, RAG and grounding data
Buying relevance judgments (qrels) for retrieval evaluation
Quick answer
A relevance judgments dataset (qrels) is a set of human grades linking each test query to documents in one fixed corpus snapshot. To buy or commission one, specify the corpus version, the topics, the pooling method (the union of top-k results from several systems), a graded scale with written guidelines, the TREC four-column delivery format, and assessor quality controls. Budget for mostly non-relevant judgments, and license the grades for evaluation and for reranker training if you plan both.
By SourceX Editorial · Updated
Qrels are the expensive, slow-to-rebuild half of a test collection. Public sets such as MS MARCO are useful for method development, but Microsoft states its datasets are intended for non-commercial research only [5], and none of them grade your corpus against your users' queries. This guide covers what to put in the statement of work. The wider decisions about corpus snapshots and topic design live in building a private domain retrieval test collection, and the cluster overview is the retrieval and RAG data buyer's guide. For evaluation data beyond retrieval, see AI evaluation data and the knowledge retrieval evaluation use case; a finished qrels set is one form of golden dataset.
What a qrels deliverable actually contains
A usable qrels deliverable is four linked artifacts, not one file: a pinned corpus, a topic set, the judgments and the assessment guidelines. NIST describes TREC judgments as "complete" only for the specific set of documents judged that year, meaning enough results were assessed that most relevant documents are assumed found, and the judgments must be used with the matching document collection [1]. If document IDs drift because the corpus was re-chunked, re-crawled or de-duplicated, the grades silently point at the wrong text.
Ask for each of these as a named, versioned file:
- Corpus manifest: document IDs, content hashes, snapshot date and the chunking or passage-splitting rule. See pinned corpus snapshots for reproducible RAG evaluation.
- Topics: query text plus a description and narrative field stating what counts as relevant, in the TREC topic style.
- Qrels: the grades themselves, in TREC format.
- Guidelines: the grade definitions, worked borderline examples and the escalation rule for disagreements.
- Pool provenance: which systems or runs contributed candidates, at what depth, and the run files themselves.
The run files matter more than buyers expect. Without them you cannot later tell whether a new system is penalized for retrieving relevant documents that no pooled system surfaced.
The TREC qrels format and how evaluators read it
Deliver judgments as whitespace-separated TREC qrels: topic ID, iteration, document ID and relevance grade, one judgment per line [3]. The iteration column is conventionally 0 and ignored by most tools. Evaluators such as trec_eval and trectools [6] treat any document missing from the qrels as non-relevant [3], which is why pool depth and judged coverage control how trustworthy your metrics are.
Illustrative example: invented to show structure; it does not describe an available dataset.
Q0412 0 kb-2026-03:art-18872#p3 3
Q0412 0 kb-2026-03:art-18872#p4 1
Q0412 0 tkt-2026-03:case-559102 0
Q0412 0 kb-2026-03:art-20410#p1 2
Q0413 0 kb-2026-03:art-00931#p2 0
Two details prevent rework. First, require document IDs that embed or reference the snapshot (here, kb-2026-03), so a qrels file cannot be paired with the wrong corpus by accident. Second, agree whether the unit of judgment is the document or the passage; mixing them breaks nDCG@10 and recall@k comparisons.
Ask for a sidecar file alongside the qrels with per-judgment metadata that the four-column format cannot carry: assessor ID (pseudonymous), timestamp, time spent, confidence, whether the item was double-judged and the adjudicated final grade. This is what lets you audit agreement later.
Pooling: deciding which documents get judged
Pooling means judging the union of the top-k results from several independent retrieval systems for each topic, rather than the whole corpus. In the TREC 2022 NeuCLIR track, pools were built from the top-ranked documents of submitted runs, with pool depth varying by run type [2]. The design choice for a buyer is which systems feed the pool and how deep each one goes.
Practical pooling rules for a commissioned collection:
- Diversify the contributing systems. Include at least a lexical system (BM25), a dense retriever and, if you have one, your current production stack plus a cross-encoder reranker. A pool built only from your current system will make that system look complete.
- Vary depth by system strength. Deeper pools for weaker or more different systems surface relevant documents others miss.
- Hold back a system. Leave one strong run out of the pool, then measure what fraction of its top 10 is judged. A low judged@10 for an unpooled system means the collection will penalize genuinely new retrievers.
- Add targeted distractors. Superseded policy versions and near-duplicates test whether the grades separate current from stale content; see superseded versions, drafts and near-duplicates for retrieval testing.
If your team plans to replace human assessors with model-generated labels for the deeper pool layers, decide that before commissioning. The trade-offs are covered in LLM relevance labels vs human assessors.
Graded scales and assessor guidelines
Use a graded scale with written definitions, not binary relevant or not relevant. NeuCLIR assessors sorted documents into four categories that were then mapped to a 3/1/0 point scale for scoring [2], which shows a useful pattern: assessors work with meaningful labels, and the gain values used by nDCG are a separate, documented mapping.
A defensible guideline set for a domain corpus specifies:
- Grade definitions tied to the task. For a support-deflection use case, "fully answers the question and is current" is a different grade from "on topic but requires a second document."
- Recency and supersession rules. Whether a correct answer from a superseded article is graded down, and by how much.
- Partial and fragment rules. How to grade a passage that contains the answer only when read with its neighbor.
- Assessor expertise. Which topics require a domain specialist (claims adjusters, clinical coders, network engineers) and which a trained generalist can judge.
Detailed scale design, including how to write borderline examples for domain experts, is on relevance assessment guidelines and graded scales.
How many judgments you need and what they cost to produce
Most of your judging budget will go to non-relevant documents. An analysis of the TREC-3 to TREC-8 ad hoc collections found that on average about 94% of judged documents were non-relevant [4]. That ratio is expected, because pools are deliberately wide, but it means "number of relevant documents found" is a poor unit for pricing or progress.
Estimate volume from three numbers: topics, pool depth and pool overlap. The arithmetic below is a planning sketch, not a benchmark.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Parameter | Planning value | Effect on volume |
|---|---|---|
| Topics | 150 | Linear |
| Contributing systems | 4 | Sublinear, because results overlap |
| Depth per system | 20 (BM25, dense, production), 30 (reranker) | Roughly linear per system |
| Raw candidates per topic | 90 | Before de-duplication |
| Unique documents per topic after overlap | about 55 | Depends on how similar the systems are |
| Double-judged share | 15% | Adds about 8 judgments per topic |
| Estimated judgments | about 9,500 | 150 x (55 + 8) |
Two levers matter most. Topic count drives statistical power for comparing systems, so prefer more topics with moderate depth over few topics judged very deeply. Depth drives how reusable the collection is for future, unpooled systems, so check judged@k on a held-back run before deciding depth is sufficient.
Quality controls to write into the statement of work
Quality in qrels is measured by agreement, coverage and consistency, and each needs a contractual number or procedure. Specify a double-judging rate (a stratified sample across topics and grades), an agreement statistic such as Cohen's kappa or Krippendorff's alpha reported per topic, and an adjudication rule naming who resolves disagreements and whether the original grades are retained.
Additional checks a buyer can run on delivery:
- Gold insertion. Seed known-relevant and known-irrelevant documents into the pool; track each assessor's accuracy on them.
- Time-per-judgment outliers. Grades made in a few seconds on long documents usually signal skimming.
- Grade distribution per assessor. One assessor grading far more documents as relevant than peers on the same topics needs review.
- Corpus integrity. Every document ID in the qrels resolves to a hash in the corpus manifest [1].
- Format validation. The file loads in trec_eval or trectools [6] without dropped lines.
Licensing qrels for evaluation versus reranker training
Grades bought for evaluation are often reused as training labels, so the license should say which uses are allowed. Reranker fine-tuning consumes query and candidate-document pairs labeled positive or negative [7], and qrels map directly onto that shape, including the hard negatives from the pool. If the same topics feed training, your evaluation numbers are contaminated, so split topics before the work starts and record which split each topic belongs to.
Points to settle in the license:
- Permitted uses: evaluation only, evaluation plus reranker or embedding training, or grounding.
- Corpus rights: whether you may retain and query the underlying documents, not only the grades. See grounding license vs training license.
- Topic provenance: whether queries derive from real user logs and how personal data in them was removed. See de-identifying a RAG corpus without breaking retrieval.
- Refresh terms: whether new topics or re-judging after a corpus update can be added later.
Where the corpus is operational business data, such as support tickets linked to knowledge articles, some relevance signal already exists in the source system; support tickets linked to knowledge articles explains how to use it as a seed before human grading. Teams that need the corpus itself from a business that holds it can describe the data to SourceX's buyer desk.
Buyer request template for commissioned qrels
A request that a supplier can scope and price names the corpus, the use, the scale and the controls. Copy and adapt the template below.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: relevance_judgments
corpus:
description: "Product support knowledge base articles and resolved tickets"
snapshot_required: true
unit_of_judgment: passage # document | passage
passage_rule: "split at H2, max 300 tokens"
topics:
count: 150
source: "de-identified user queries + analyst-written"
fields: [title, description, narrative]
split: {eval: 100, train: 50}
pooling:
systems: [bm25, dense_retriever, production_stack, cross_encoder_rerank]
depth: {bm25: 20, dense_retriever: 20, production_stack: 20, cross_encoder_rerank: 30}
holdout_system_for_judged_at_10: true
scale:
labels: [highly_relevant, relevant, related_not_answering, not_relevant]
gain_mapping: [3, 1, 0, 0]
qa:
double_judged_share: 0.15
agreement_metric: cohen_kappa_per_topic
adjudication: "senior assessor; keep original grades"
gold_items_per_assessor: 20
delivery:
qrels_format: trec_4col
sidecar: [assessor_id, timestamp, seconds_spent, confidence, adjudicated]
include_run_files: true
license_uses: [evaluation, reranker_training]
Sourcing the corpus behind your relevance judgments
SourceX sources operational datasets, such as support and sales histories, engineering records and documents, from US companies on request and manages the licensing, with every release approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and nothing is contracted until a supplier agrees. Tell SourceX what corpus and judgments you need.
Frequently asked questions
Can I reuse public TREC qrels for a commercial product evaluation?
Check each collection's terms individually. MS MARCO, a common starting point, states its datasets are for non-commercial research only [5]. Public qrels also grade a different corpus, so they cannot measure retrieval quality on your own documents; see which public retrieval datasets allow commercial use.
Are unjudged documents really scored as non-relevant?
Yes, by default in standard evaluators [3]. That is why a new system retrieving relevant but unpooled documents scores lower than it should. Report judged@k alongside nDCG, and add new runs to the pool before drawing conclusions.
Should qrels be graded at document or passage level for RAG?
Judge at the unit your retriever returns. If your pipeline retrieves chunks, document-level grades overstate precision because a relevant document can contain many irrelevant chunks. Fix the chunking rule before judging, since re-chunking invalidates the passage IDs.
Sources
- NIST TREC, "Data - English Relevance Judgements". https://trec.nist.gov/data/reljudge_eng.html
- arXiv (Lawrie et al.), "Overview of the TREC 2022 NeuCLIR Track" (2023). https://arxiv.org/pdf/2304.12367
- Tessl registry, "judgment-list-author". https://tessl.io/registry/testland/judgment-list-author/files
- NTCIR / National Institute of Informatics (Sukomal Pal et al.), "Relevance assessment in TREC ad hoc collections (EVIA 2010)" (2010). https://research.nii.ac.jp/ntcir/workshop/OnlineProceedings8/EVIA/02-EVIA2010-SukomalP.pdf
- Microsoft, "MS MARCO Datasets". https://microsoft.github.io/msmarco/Datasets.html
- GitHub (joaopalotti/trectools), "trectools". https://github.com/joaopalotti/trectools
- NVIDIA, "Reranking Model Fine-Tuning Recipe". https://docs.nvidia.com/nemotron/nightly/nemotron/rerank/README.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.