Retrieval, RAG and grounding data
Hard negatives for retriever training: sourcing, mining and false-negative control
Quick answer
Hard negatives are passages that look relevant to a query but do not answer it, and they are what teach a dense retriever fine distinctions. The catch is that the closer a negative sits to the query, the more likely it is an unlabeled positive [1]. Useful hard negatives therefore come from a mining step plus a false-negative filter: a cross-encoder or positive-aware threshold, LLM or expert relabeling, and a held-out audit. Buy or build negatives with that provenance recorded, not as an anonymous column.
By SourceX Editorial · Updated
Why hard negatives matter for contrastive retriever training
Hard negatives matter because random and in-batch negatives are too easy to move a well-initialized encoder. A bi-encoder trained with InfoNCE-style loss learns from the margin between the positive and the negatives in each group; if every negative is topically unrelated, the gradient carries little signal about near-miss passages, which are exactly what your production index is full of. Typical near misses are the superseded version of a policy, the knowledge article for the adjacent product SKU, or the ticket with the same error code but a different root cause.
Mined negatives are therefore one of the highest-leverage choices in a retriever training run, and also one of the least documented. Two teams can start from the same queries and positives and get different models purely because one mined from BM25 ranks 1 to 10 and the other filtered dense-retriever candidates with a teacher model. If you are training a domain model, read training data for domain-specific embedding models alongside this page for the positive side of the dataset.
The false-negative problem in hard negative mining
The central failure mode is that a mined "negative" is often a relevant passage that nobody labeled. Benchmarks such as MS MARCO and Natural Questions label one or a few positives per query, so anything else the retriever ranks highly is treated as negative by default. Spot checks of a strong retriever's top-ranked unlabeled passages routinely turn up relevant ones, which is why every mining run needs an audit.
Training on those passages as negatives actively pushes relevant content away from the query. The AAAI 2024 work on contrastive confidence regularization frames the trade-off directly: harder negatives add more signal and more label noise at the same time [1]. ARHN reports that false negatives distort the geometry of the embedding space and that the damage shows up most in zero-shot and out-of-domain retrieval [2], which is where enterprise buyers usually deploy.
Domain corpora make this worse, not better. Enterprise collections are dense with duplicates, templated answers and near-identical revisions; for that class of document see superseded versions, drafts and near-duplicates for retrieval testing. A support knowledge base with 40 variants of a password-reset article will yield "hard negatives" that are simply alternate correct answers.
Mining methods and how each controls false negatives
Each mining method answers the same question differently: how do you decide that a close passage is really wrong? The table compares the common options.
| Method | How negatives are picked | False-negative control | Main risk |
|---|---|---|---|
| BM25 top-k | Lexical overlap with the query, minus labeled positives | None by default | Lexical paraphrases of the answer become negatives |
| Self-mined (ANCE-style) | Current or previous checkpoint's top-k from the dense index | None unless a filter is added | Feedback loop: the model's own blind spots become labels |
| Cross-encoder denoised | Retriever top-k, then drop items a cross-encoder scores as relevant | Teacher score threshold [3] | Teacher errors and domain mismatch |
| Positive-aware threshold | Teacher scores; keep negatives scoring below a fraction of the positive's score | Relative to each query's positive | Threshold tuning; weak positives pull the cut down |
| LLM relabeling | An LLM judges whether each candidate answers the query | Answer-centric judgment [2] | Cost, prompt sensitivity, LLM bias |
| Learned reweighting | An adapter estimates false-negative likelihood and downweights | Soft weights rather than hard drops [3] | Extra model to train and validate |
| Expert verification | Domain reviewers label a sample or all candidates | Human judgment with written guidelines | Cost; inter-annotator disagreement |
Positive-aware mining is a common, practical choice. A teacher model scores the labeled positive and every candidate, and candidates scoring above a set fraction of the positive's score are dropped as likely unlabeled answers. Because the cut is relative to each query, it adapts to queries where all scores are low; the teacher choice still matters, so test at least two teachers on a labeled sample. Cross-encoder denoising, answer-based heuristics and LLM relabeling are the other common filters reviewed in recent work [3].
Soft approaches avoid binary decisions. RRRA trains a retriever adapter that estimates how likely each negative is to be a false negative and uses that to resample and rerank rather than drop items outright [3]. Confidence regularization in the AAAI paper similarly makes the contrastive loss more robust to likely false negatives [1].
Choosing ratios, depth and group structure
Start with one to seven hard negatives per query in each training group and tune against a held-out set rather than copying a published ratio. More hard negatives per positive increase both signal and exposure to false negatives, so the right number depends on how clean your filter is. In-batch and cross-batch negatives remain useful as cheap easy negatives alongside the mined set.
Mining depth matters as much as count. Sampling from ranks 1 to 10 of a strong retriever produces the hardest, noisiest candidates; sampling from deeper bands (for example, ranks 30 to 200) gives negatives that are still on-topic but less likely to be unlabeled answers. Many teams mix bands and record the band in the data so the ratio can be changed later without re-mining.
Refresh negatives as the model improves. Negatives mined with the base checkpoint become easy after a few epochs; re-mining with the current checkpoint keeps the signal, but each refresh must rerun the false-negative filter. Track the filter's drop rate per refresh: a sudden rise usually means the model has learned to retrieve answers your positives never labeled. For reranker groups and distillation scores built from the same candidates, see reranker training data.
Where hard negatives come from in real business data
Real operational data supplies negatives that synthetic pipelines struggle to fake, because the confusable passages already exist. The best sources pair a real query with an observed resolution and an observed non-resolution.
- Support tickets linked to knowledge articles. The article an agent attached and the case closed on is a positive; articles the customer opened from search but that did not resolve the case, or that were attached and then replaced, are candidate hard negatives. See support tickets linked to knowledge articles.
- Search and click logs. Results shown above the clicked result and skipped are a weak negative signal; position bias means they need filtering. See search query and click logs for retriever training.
- Document version history. A superseded procedure is a strong negative for a query that needs the current one, provided effective dates are in the metadata.
- Engineering records. Incidents with the same error signature but different root causes give negatives that require reading the fix, not matching the stack trace.
Public sets are a starting point, not a finished answer. Some, such as WebFAQ 2.0, ship with mined hard negatives for dense retrieval [4]; before reusing those, check which retriever mined them, whether any false-negative filter was applied, and whether the license allows commercial training (see which public retrieval datasets allow commercial use). Synthetic queries can add coverage, with trade-offs covered in synthetic vs real queries for retriever training.
What to request from a supplier of hard negatives
Ask for negatives as documented records, not as a list of IDs. Data Cards-style documentation of upstream sources, collection and annotation methods and intended use applies directly here [5]: you need to know how each negative was chosen to know how much to trust it.
Supplier request checklist
- Mining method per negative (BM25, dense model name and checkpoint, rank band).
- False-negative filter used: teacher model and threshold, LLM prompt and model, or reviewer guideline version.
- Verification status per record: unverified, model-filtered, or expert-verified on a sample, with the sample size and disagreement rate.
- Corpus snapshot date and document version IDs, so negatives can be traced to the exact passage text.
- Known positives per query (all of them, not one), plus any "partially relevant" grade.
- De-identification method for ticket text and customer details, and the license terms covering training use.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"query_id": "q-018442",
"query": "printer shows E-37 after firmware update, still jams",
"positives": [{"doc_id": "kb-2291", "version": "v4", "grade": 2}],
"negatives": [
{
"doc_id": "kb-2291",
"version": "v2",
"source": "version_history",
"reason": "superseded_procedure",
"mined_by": null,
"teacher_score": 0.41,
"positive_score": 0.88,
"filter": "positive_aware_lt_0.90",
"verification": "expert_sample",
"reviewer_guideline": "neg-guide-1.3"
},
{
"doc_id": "kb-1876",
"version": "v7",
"source": "dense_topk",
"mined_by": "base-encoder@ckpt-12",
"rank": 46,
"teacher_score": 0.33,
"filter": "cross_encoder_lt_0.5",
"verification": "model_filtered"
}
],
"corpus_snapshot": "2026-09-30",
"pii_treatment": "names_and_serials_replaced"
}
Run an acceptance audit before training. Pull a random sample of a few hundred negatives stratified by rank band and method, have two reviewers judge each against the query, and compute the false-negative rate per stratum. Reject or re-filter any stratum whose rate is materially above what your validation experiments tolerate.
Evaluating whether negatives actually helped
Judge negatives by out-of-domain and in-domain retrieval metrics on a test set that was never mined from. Compare nDCG@10 and Recall@100 for three runs: in-batch negatives only, mined negatives without filtering, and mined negatives with your filter. If the unfiltered run wins in-domain but loses on held-out domains, false negatives are the likely cause, consistent with what ARHN reports [2].
Keep the evaluation qrels independent of the mining pipeline. If the same teacher that filtered training negatives also produced the test labels, the evaluation will reward agreement with the teacher rather than true relevance; buying relevance judgments (qrels) covers how to source independent judgments. For background on the vectors themselves, see the embeddings glossary entry.
How SourceX supports retrieval data requests
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and documents, and manages licensing and ongoing purchases; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed, personal details are removed or replaced with the method recorded and a sample checked, and delivery happens only after an executed license and supplier approval. You describe the query-passage data and negative signals you need on the SourceX buyer page, and terms are agreed per deal. More retrieval guides are in the retrieval, RAG and grounding data hub and the AI data hub.
Sourcing real query-passage data with verified negatives
If public mined negatives are too noisy for your domain, real operational records with observed resolutions are the usual next step. SourceX looks for US businesses that hold the data you describe, reviews ownership and consents, and delivers under a license that defines records, uses, term and delivery. Describe the retrieval training data you need.
Sources
- AAAI 2024 (via ML Anthology), "Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence Regularization" (2024). https://mlanthology.org/aaai/2024/wang2024aaai-mitigating
- arXiv, "ARHN: Answer-Centric Relabeling of Hard Negatives with Open-Source LLMs for Dense Retrieval" (2026). https://arxiv.org/pdf/2604.11092
- arXiv, "RRRA: Resampling and Reranking through a Retriever Adapter" (2025). https://arxiv.org/html/2508.11670v1
- arXiv, "WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval" (2026). https://arxiv.org/pdf/2602.17327
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.