Skip to content

Retrieval, RAG and grounding data

Reranker training data: query-document pairs, listwise groups and distillation labels

Quick answer

Reranker training data is a set of real domain queries, each paired with candidate passages and a relevance signal. The signal can be a binary label (pointwise), a preferred-versus-rejected pair (pairwise), one positive grouped with several hard negatives (listwise), or teacher scores for distillation. Take negatives from your own first-stage retriever and remove known positives. Split by query, deduplicate passages across splits, and evaluate both in-domain and out-of-domain before you ship a fine-tuned cross-encoder.

By SourceX Editorial · Updated

What a cross-encoder actually learns from

A cross-encoder reranker learns from the joint text of a query and one candidate passage, so every training example has to carry the exact passage text your retriever will hand it at inference time. That is the main difference from bi-encoder data. An embedding model (see embeddings) learns to place queries and documents near each other in vector space, and it usually benefits from in-batch negatives. A reranker scores pairs that a first stage has already shortlisted, so its useful negatives are the near-misses that the first stage ranks highly.

In practice this means three things for data. First, keep the same chunking, title prefixing and truncation (for example, a 512-token limit on query plus passage) in training that you use in production. Second, store the retrieval context: which retriever produced each candidate and at what rank. Third, keep the query text as users actually wrote it, typos included. If you also need data for the first stage, the companion guide on training data for domain-specific embedding models covers contrastive pairs and in-batch negative formats.

The four reranker dataset formats

Pick the format by the label source you actually have, not by the loss you want to try. Most training stacks can turn a richer format into a simpler one, but not the reverse.

FormatRecord shapeTypical label sourceCommon loss familyMain failure mode
Pointwise(query, passage, label 0/1 or graded 0-3)Gold answer sets, qrels, ticket-to-article linksBinary cross-entropyClass imbalance; easy negatives teach little
Pairwise(query, positive, negative)Clicks versus skips, A/B preferencesMargin ranking / RankNet-stylePosition bias in click data
Listwise group(query, 1 positive, k hard negatives)Gold positive plus mined negativesSoftmax cross-entropy over the groupFalse negatives inside the group
Distillation(query, passages, teacher scores)Larger cross-encoder or LLM judgeMSE on scores or score margins, KL over listsStudent inherits teacher errors and bias

Pointwise 0/1 labels from a gold set are the simplest starting point. One biomedical QA system labels question-document pairs against a gold set of relevant documents, treating everything else retrieved for the question as negative [1]. Listwise groups are common in open reranker fine-tunes: one public model card trains on groups of one positive and seven hard negatives built from Natural Questions [2]. Pairwise data mostly comes from behavior logs, which the guide to search query and click logs for retriever training treats in depth, including position-bias corrections.

Where hard negatives come from, and how they go wrong

Good negatives come from your first-stage retriever's results for the same query, minus every document known to be relevant. The biomedical QA work does exactly this: it takes index results and subtracts the gold set [1]. A Polish passage-retrieval system samples 100 negatives per positive from the top-2000 BM25 results [3]. Sampling from a wide band rather than only the top few gives the model both hard and moderate contrasts.

The dominant failure is the false negative. A passage that ranks highly for a query is, by construction, more likely to be relevant but unlabeled, and recent dense-retrieval research documents how often mined "negatives" turn out to answer the question [5]. In enterprise corpora the risk is higher because the same answer often lives in a knowledge article, a pasted ticket reply and a wiki copy. Three controls help:

  • Deduplicate near-identical passages (MinHash or exact-hash on normalized text) before mining, so copies of the positive cannot become negatives.
  • Skip the top 1-3 retrieved candidates when sampling negatives, or relabel them with a strong judge before use.
  • Audit a sample of negatives by hand; label errors exist even in curated benchmarks, with one audit estimating at least 3.3% on average across widely used test sets [6].

Linked operational records reduce this problem because the positive is an observed resolution rather than a guess. Support tickets linked to knowledge articles are a good example: the agent's cited article is a ready-made positive, and other articles retrieved for the ticket text are candidate negatives.

Distillation labels: what to store and why

Distillation labels are continuous relevance scores from a stronger teacher, and you should store the raw scores, the teacher identity and the prompt, not just a thresholded 0/1. A typical record holds one query, the candidate list your first stage returned, and a teacher score per candidate. The student then learns either the scores (MSE), the score margin between a positive and a negative (Margin-MSE), or the teacher's distribution over the list (KL divergence).

Record the teacher model and version, the scoring prompt or checkpoint, the decoding settings, and the date. Without that provenance you cannot rerun labels after a teacher upgrade or explain a regression. If the teacher is a hosted LLM, check its terms for restrictions on using outputs to train other models before you build a training set on them. The trade-offs between machine and human judgments are covered in LLM relevance labels vs human assessors.

Distillation does not remove the need for a human-labeled test set. The student can match the teacher closely and still be wrong where the teacher is wrong, so keep at least one evaluation slice judged by domain experts.

How many examples to fine-tune a reranker

There is no universal count; the useful number is the one where your in-domain metric stops improving while your out-of-domain metric holds. For scale, a public bge-reranker-base fine-tune used 2,034 listwise groups from Natural Questions and reports an in-domain gain with near-zero change out of domain [2]. Treat that as one data point, not a target.

Run a learning curve instead of guessing. Train on 10%, 25%, 50% and 100% of your queries, holding the evaluation queries fixed, and plot nDCG@10 or MRR@10 for each. If the curve is still rising at 100%, more queries will help more than more negatives per query. If it is flat, spend effort on label quality and negative quality instead.

Query diversity usually matters more than pairs per query. Five hundred distinct real queries with one verified positive each tend to be more useful than fifty queries with ten positives, because the reranker must generalize across intents, not memorize a few. Where labeled pairs are scarce, synthetic pipelines exist: NVIDIA publishes a six-stage recipe that generates synthetic training data from a domain corpus, fine-tunes a cross-encoder and deploys it [4]. Synthetic queries still need real queries for evaluation.

Splits, leakage and evaluation design

Split by query, never by row, so that no query appears in both training and test. The biomedical QA system splits by question for this reason [1]. Row-level splits put the same query's positive in training and its sibling pairs in test, which inflates metrics.

Two further leaks are specific to rerankers. Near-duplicate queries ("reset SSO password" and "how to reset sso pw") should land in the same split; cluster them by normalized text or embedding similarity first. Near-duplicate passages across splits also leak, and duplicated text is common in training corpora [7]. For time-stamped data such as tickets or logs, a temporal split (train on older months, test on the latest) is the most honest proxy for production.

Evaluate on at least two slices. The in-domain slice measures the gain you paid for. An out-of-domain slice, such as a public benchmark or a different product line, catches regressions that the in-domain metric alone would hide; the public fine-tune cited above reports its out-of-domain result next to its in-domain gain [2]. For building the in-domain slice properly, see building a private domain retrieval test collection and buying relevance judgments (qrels).

A record schema and acceptance checklist

A well-formed delivery uses one JSON Lines record per query group, with stable IDs that join back to a passage table. The example below shows the fields worth requesting from any supplier or internal team.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"query_id": "q_000184", "query": "warranty claim rejected missing serial plate", "split": "train", "query_cluster": "c_0217", "created_at": "2025-11-03",
 "positives": [{"passage_id": "kb_4471#c3", "label": 1, "label_source": "agent_cited_article"}],
 "candidates": [
   {"passage_id": "kb_4471#c3", "retriever": "bm25", "rank": 2, "teacher_score": 8.91},
   {"passage_id": "kb_0912#c1", "retriever": "bm25", "rank": 1, "teacher_score": 6.40},
   {"passage_id": "kb_3305#c7", "retriever": "dense_v2", "rank": 5, "teacher_score": 1.12}],
 "teacher": {"model": "teacher-xl-2026-08", "prompt_version": "rel_v3"},
 "chunking": {"method": "heading_split", "max_tokens": 384}}

Acceptance checklist before training:

  1. Every passage_id resolves to text in the passage table, with the same chunking used in production.
  2. No query_id or query_cluster appears in more than one split.
  3. Exact and near-duplicate passages are collapsed before negatives are assigned.
  4. Positives record label_source (gold set, click, cited article, human judge, teacher threshold).
  5. Each candidate records retriever and rank, so negatives can be resampled by band.
  6. A hand audit of at least a few hundred negatives estimates the false-negative rate.
  7. Personal data in queries and passages has been removed or replaced, with the method documented; see de-identifying search query logs.
  8. The license allows model training, not just grounding at query time; the difference is set out in grounding license vs training license.

Sourcing real domain pairs instead of generating them

Real query-passage pairs come from operational systems where a person searched or asked, and then used a specific document. Support desks (ticket text linked to the cited article), internal search logs with clicks and dwell, sales engineering Q&A linked to product documentation, and engineering incident records linked to runbooks all produce this structure as a side effect of work. Their main advantage over synthetic data is query realism: abbreviations, misspellings and partial context that generated questions rarely reproduce.

These records need rights review and de-identification before they leave the company that holds them. When you request them, describe the data structure you need (query field, linked document field, resolution signal, date range, volume of distinct queries) rather than a named business. For a broader view of retrieval data types and licensing, start at the retrieval and RAG data hub; for fine-tuning beyond ranking, see training data for domain-specific fine-tuning. Teams that want help locating such records can describe the dataset to SourceX.

Find domain query-document pairs for your reranker

SourceX sources operational datasets, such as support and sales histories, engineering records and documents, from US companies on request; categories are not inventory and a request does not guarantee a match. Each dataset is rights-reviewed and de-identified before delivery under a license that defines records, uses, term and delivery. Describe the reranker training data you need.

Frequently asked questions

Can I reuse my embedding model's training pairs for the reranker?

Partly. Positives transfer directly, but in-batch or random negatives are too easy for a cross-encoder. Re-mine negatives from the first-stage retriever you will deploy, so the reranker sees the near-misses it must actually separate.

Should labels be binary or graded?

Use graded labels (for example, 0-3) when you have them for evaluation, because nDCG rewards ordering among relevant passages. For training, binary labels plus listwise groups work well, and graded labels can be collapsed at a chosen threshold.

Do I need human labels if I distill from an LLM?

Yes, for the test set. Teacher scores are a cost-effective training signal, but only human or outcome-based judgments tell you whether the student and the teacher are both wrong in the same place.

Sources

  1. arXiv, "Beyond Retrieval: Ensembling Cross-Encoders and GPT Rerankers with LLMs for Biomedical QA" (2025). https://arxiv.org/pdf/2507.05577
  2. Hugging Face, "bge-reranker-base fine-tuned on Natural Questions (train split)". https://huggingface.co/sfczaa/bge-reranker-base-nq-ft
  3. arXiv, "Passage Retrieval of Polish Texts Using OKAPI BM25 and an Ensemble of Cross Encoders" (2024). https://arxiv.org/pdf/2410.04620
  4. NVIDIA, "Reranking Model Fine-Tuning Recipe". https://docs.nvidia.com/nemotron/nightly/nemotron/rerank/README.html
  5. arXiv, "ARHN: Answer-Centric Relabeling of Hard Negatives with Open-Source LLMs for Dense Retrieval" (2026). https://arxiv.org/pdf/2604.11092
  6. arXiv, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. arXiv, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data