Evaluation and benchmarking datasets
Unanswerable questions in RAG evaluation: testing abstention
Quick answer
An abstention test set mixes questions your corpus cannot answer into the evaluation, each labeled with an answerability class, the reason it is unanswerable, and the corpus snapshot it was checked against. You then score two errors separately: answering when the corpus has no support (fabrication) and declining when it does (over-abstention). The hard part is not the metric. It is finding realistic unanswerable questions and proving, at each corpus version, that they really are unanswerable.
By SourceX Editorial · Updated
What counts as unanswerable in a retrieval system
A question is unanswerable only relative to a specific corpus snapshot, so every label must name the snapshot it was verified against. The same question about a 2025 refund policy is answerable in one index build and unanswerable after the policy document is retired. Pin the corpus as described in pinned corpus snapshots for reproducible RAG evaluation before you label anything.
Within that frame, unanswerable items fall into distinct types that fail for different reasons:
- Out of corpus. The topic is plausible for the domain but no document covers it, such as a warranty term for a product line the company never sold.
- Near miss. The corpus covers a neighboring fact: a different year, region, product version or customer tier. A question about a 2022 award when the corpus only covers 2021 looks answerable to a retriever and invites a confident wrong answer.
- False premise. The question assumes something the corpus contradicts or never states ("Why was the Denver office closed?" when no closure is recorded).
- Partially answerable. The corpus supports one part of a compound question but not another. The right behavior is to answer the supported part and flag the gap.
- Out of scope by policy. The corpus may hold the fact, but the user is not entitled to it. That is an access-control test, covered separately in permission-aware RAG evaluation.
Keep retrieval misses out of this class. If the answer exists in the corpus but the retriever failed to surface it, the item is answerable and the failure belongs to retrieval recall, not abstention.
The answerability label schema
Use a three-way answerability label plus a reason code, because a binary flag hides the partial cases that cause most production complaints. Annotation guides for RAG ground truth start with exactly this step: the annotator decides whether the corpus can answer the question before writing any reference [2]. Store the evidence for the decision, including documents that look relevant but do not answer, so a reviewer can audit it later.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"question_id": "abst-000412",
"question": "What is the cancellation fee for the Enterprise Plus annual plan signed in 2024?",
"corpus_snapshot_id": "kb-2026-09-15-r3",
"answerability": "unanswerable",
"reason_code": "near_miss_version",
"reason_note": "Corpus covers Enterprise (2023 and 2024) and Enterprise Plus (2025 only). No 2024 Enterprise Plus terms.",
"nearest_distractor_doc_ids": ["pricing-terms-ent-2024.pdf#p3", "pricing-terms-entplus-2025.pdf#p2"],
"gold_doc_ids": [],
"expected_behavior": "abstain_and_point_to_nearest",
"acceptable_partial": null,
"source_of_question": "escalated_support_ticket",
"verification": {"method": "keyword+dense search over snapshot, two annotators", "agreement": true, "adjudicated_by": null},
"pii_handling": "account numbers replaced with tokens"
}
Useful expected_behavior values are answer, answer_partial_and_flag_gap, abstain, abstain_and_point_to_nearest and ask_clarifying_question. The last one matters for ambiguous questions, where declining outright is less useful than asking which plan or which year the user means. Answerable items in the same file should carry gold_doc_ids and the citation spans described in question-answer-citation triples for RAG evaluation, so one schema serves the whole suite.
Where realistic unanswerable questions come from
The most realistic unanswerable questions come from your own users, because they reflect the gaps people actually hit rather than the gaps an annotator imagines. Treat the following as working sources, ranked roughly by realism; the ranking is practitioner judgment, not a published standard.
- Escalated or unresolved support tickets. Tickets that went to a human because the help center had no answer are natural out-of-corpus questions. Check that the ticket's resolution was never written back into the knowledge base.
- Zero-result and low-engagement search queries. Internal search logs with no clicks, or with a follow-up "contact us" event, point at coverage gaps.
- Corpus ablation. Take a verified answerable question and remove its gold documents from a copy of the index. The question keeps its natural phrasing, and you know exactly why it became unanswerable.
- Controlled perturbation. Swap the entity, year, region or product version in an answerable question to create near misses that sit right next to real evidence.
- LLM-generated questions. Fast and cheap, but prone to surface cues: generated items tend to be shorter, vaguer or more obviously off-topic than real ones, so a model can spot them without checking evidence. Prompt for questions that resemble your answerable items in length, entities and phrasing.
A useful check on realism is whether a simple classifier using only the question text, with no retrieval, can separate answerable from unanswerable items. If it can, your set rewards spotting style rather than checking evidence.
Scoring abstention without rewarding refusal
Score abstention as a pair of error rates, never as a single accuracy figure, because a system that always says "I don't know" gets every unanswerable item right. The research literature on LLM abstention treats it as a trade-off between refusing when the model should and answering when it can, and reviews benchmarks and metrics for both sides [1]. Your suite needs the same two-sided view on your own corpus.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Gold label | System answered, supported | System answered, unsupported | System abstained | System answered part, flagged gap |
|---|---|---|---|---|
| Answerable | Correct | Wrong answer | Over-abstention | Incomplete |
| Partially answerable | Over-claim (no gap flagged) | Fabrication on missing part | Over-abstention | Correct |
| Unanswerable | Not possible by definition; re-check label | Fabrication | Correct abstention | Fabrication |
Report at least these numbers, each per reason code and per question source:
- Abstention recall: share of unanswerable items where the system declined or asked for clarification.
- Fabrication rate: share of unanswerable items where the system produced a substantive answer.
- Over-abstention rate: share of answerable items where the system declined.
- Unsupported-answer rate on answerable items: answers not grounded in retrieved context, scored with the same judge you use for faithfulness evaluation. Datasets that include both faithful and unfaithful answers, such as rag-qa-arena, are a practical way to validate that judge [3].
Detecting an abstention is itself a classification problem. Phrase lists miss hedged answers like "Typically, fees are around..." that sound cautious but still assert a fact. Use an LLM or classifier judge with an explicit rubric, then validate it against a few hundred human-labeled responses before trusting its numbers. If your system exposes a confidence score, plot a risk-coverage curve so product owners can choose a threshold rather than accept a default.
Mix ratio and sample size
No published standard sets the share of unanswerable items, so choose it from the decision you need to make and report slices separately. A blended score depends heavily on the mix: raise the unanswerable share and an over-cautious system looks better, lower it and a reckless one does. For that reason, many teams treat the mix as a design parameter and record it in the evaluation specification.
Two practical rules help. First, size each slice for the difference you need to detect, not the suite as a whole; a reason code with 30 items cannot distinguish a 10% from a 15% fabrication rate. Second, do not rely on normal-approximation error bars for small slices. A 2025 position paper argues that CLT-based intervals break down below a few hundred data points and come out too narrow [5]; use Wilson or Clopper-Pearson intervals for rates, or bootstrap at the question level.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Slice | Items | Purpose |
|---|---|---|
| Answerable, single document | 300 | Baseline accuracy and over-abstention |
| Answerable, multi-document | 150 | Over-abstention when evidence is spread out |
| Partially answerable | 120 | Gap flagging |
| Unanswerable, out of corpus (from tickets) | 150 | Fabrication on real gaps |
| Unanswerable, near miss (perturbation and ablation) | 150 | Fabrication under plausible distractors |
| Unanswerable, false premise | 80 | Premise correction |
The counts above are a starting template, not a recommendation for your traffic. Weight the reported headline by the slice frequencies you observe in production logs, and keep the unweighted per-slice table alongside it.
Failure modes that corrupt abstention sets
The most common failure is an item labeled unanswerable that some document in the corpus actually answers, which makes a correct system look like it fabricated. Test-set label errors are not rare even in well-known benchmarks; one audit of 10 widely used test sets estimated an average error rate of at least 3.3% [4]. Unanswerable labels are harder than most because proving absence requires searching everything.
Watch for these specific problems:
- Corpus drift. A re-ingest adds a document that answers a formerly unanswerable item. Re-verify every unanswerable label whenever
corpus_snapshot_idchanges, and treat new answers as label updates, not regressions. - Weak absence checks. One annotator running one keyword search will miss paraphrases, tables, scanned PDFs and slide text. Use both lexical and dense retrieval over the full snapshot, plus a second annotator.
- Conflicting and outdated sources. An answer that exists only in a superseded document is a different test; see evaluating RAG on versioned, outdated and conflicting documents.
- Style leakage. Generated unanswerable questions that are shorter, vaguer or more exotic than real ones inflate abstention recall.
- Contamination. If items are published or reused in prompts, models and judges may learn them. Keep the set private, as discussed in private evaluation sets vs public benchmarks.
- Personal data in tickets. Support logs carry names, emails and account numbers. Remove or replace them before items enter the shared evaluation store, and record the method.
Sourcing real unanswerable questions from business records
Realistic abstention tests need a corpus and the questions people actually asked of it, from the same organization, which public benchmarks cannot supply. Public reading-comprehension and RAG benchmarks test the skill in general, but they do not show whether your assistant fabricates answers about plan terms, internal procedures or engineering runbooks. The overview of LLM evaluation datasets maps where licensed data fits, and building a golden evaluation dataset from real business records covers the answerable side.
When you request this kind of data, describe it precisely: document types and date range for the corpus, question sources (escalations, unresolved tickets, search logs), the label schema above, and the evaluation-only use you intend. Ask any supplier how they verified absence, what snapshot the labels refer to, and how personal details were removed. A written evaluation dataset specification makes those answers comparable across offers.
SourceX sources operational datasets from US companies on request, including support histories, documents and engineering records, and nothing is held in stock, so a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and every release is approved by the supplying company. You can describe the corpus and question logs you need on the SourceX buyer page; see also RAG evaluation datasets from real company documents.
Get real-world data for RAG abstention testing
If your abstention suite needs real unanswerable questions paired with the corpus they were asked against, describe the data rather than the companies. SourceX looks for US businesses that hold it, assesses the data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Describe your RAG evaluation data needs.
Sources
- arXiv, "Know Your Limits: A Survey of Abstention in Large Language Models" (2024). https://arxiv.org/pdf/2407.18418
- OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
- Hugging Face (rajistics), "rag-qa-arena dataset card (README.md)". https://huggingface.co/datasets/rajistics/rag-qa-arena/blob/main/README.md
- Northcutt, Athalye, Mueller (NeurIPS 2021 Datasets and Benchmarks; arXiv), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.