Skip to content

Evaluation and benchmarking datasets

Question-answer-citation triples for RAG evaluation

Quick answer

A RAG evaluation dataset built from question-answer-citation triples pairs each question with a reference answer or a list of required claims, and with the evidence behind each claim: document IDs, versions and passage spans inside a frozen corpus snapshot. Keeping the three parts separate lets you score retrieval, answer correctness and citation support on their own. Build triples source-first, deciding answerability from the evidence before anyone writes an answer, so the ground truth is not a model's guess.

By SourceX Editorial · Updated

Why a QA pair alone cannot diagnose a RAG failure

A question with a reference answer tells you whether the final answer was right, but not whether retrieval, generation or citation broke. Gold evidence turns one pass/fail verdict on a retrieval-augmented generation system into a separate measurement for each stage.

In the Ragas v0.1.21 test format, each row holds a question, retrieved contexts, the generated answer and a ground_truth answer [1]; the system under test supplies contexts and answer, so the dataset contributes only a question and a reference. The annotated RAG QA Arena release adds the missing layer, with numbered citation markers in its answers and gold document IDs for retrieval evaluation [2]. TREC's 2025 RAG track distributes relevance judgments and nuggets (atomic facts an answer should contain) as separate downloads [3].

FailurePart of the triple that exposes itWhat to measure
Retriever misses the passage that holds the answerGold document IDs and spans vs retrieved resultsEvidence recall@k; rank of first gold passage
Generator ignores or misreads correct contextRequired claims vs the answer, where evidence was retrievedClaim recall conditional on retrieval
Answer is right but its citation does not support itCited passages vs gold evidence textCitation support rate
System answers when the corpus holds no answerAnswerability labelAbstention rate
Answer repeats a superseded or conflicting factForbidden claims tied to stale documentsForbidden-claim rate

FinanceBench pairs 10,231 questions with answers and evidence strings, and its authors report that GPT-4-Turbo with a retrieval system answered incorrectly or refused on 81% of a 150-case sample [4]. Splitting a failure rate like that between retrieval and generation requires knowing which passages hold each answer. For the same reason, the organizers of TREC's 2025 RAGTIME track note that its human-curated nuggets lack annotations of their supporting documents, which limits their use for evaluating retrieval [5].

Anatomy of a scoreable triple

A usable triple is one record with five groups of fields: the question and its origin, an answerability label, the answer broken into required and forbidden claims, evidence pinned to a document version and span, and annotation provenance. A practitioner guide to RAG ground truth proposes a similar record [6].

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "triple_id": "kbqa-000318",
  "set_version": "1.2.0",
  "corpus_snapshot": "helpcenter-2026-08-31",
  "question": {
    "text": "If we move our annual plan to a new workspace, do we keep the multi-year discount?",
    "origin": "support_ticket",
    "origin_record_id": "tkt-88412",
    "asked_at": "2026-05-12",
    "asker_scope": ["customer_admin"],
    "type": "conditional_policy",
    "documents_needed": 2
  },
  "answerability": "answerable",
  "required_claims": [
    {"id": "c1", "text": "An annual plan can move to another workspace on the same billing account.",
     "evidence_sets": [["e1"], ["e3"]]},
    {"id": "c2", "text": "The multi-year discount ends if the plan moves to a different billing account.",
     "evidence_sets": [["e2"]]}
  ],
  "forbidden_claims": [
    {"id": "f1", "text": "Moving a plan triggers a prorated refund.", "from": "kb-1042@v9 (superseded)"}
  ],
  "reference_answer": "Yes, as long as the new workspace is on the same billing account. Moving to a different billing account ends the multi-year discount.",
  "evidence": [
    {"id": "e1", "doc": "kb-1042@v12", "sha256": "3f9a1c0e", "start": 1180, "end": 1260,
     "quote": "Admins can transfer an annual plan to any workspace on the same billing account."},
    {"id": "e2", "doc": "kb-2210@v4", "sha256": "b71c55d2", "start": 402, "end": 507,
     "quote": "Multi-year pricing applies per billing account and ends if the plan moves to a different billing account."},
    {"id": "e3", "doc": "kb-0415@v3", "sha256": "9d02e4a7", "start": 2215, "end": 2278,
     "quote": "Plans can move between workspaces that share a billing account."}
  ],
  "annotation": {"author": "ann-07", "verifier": "ann-12", "adjudicated": false, "deid_method": "consistent_surrogates"}
}
  • Corpus snapshot and hashes. WixQA's authors argue that end-to-end RAG evaluation needs the specific knowledge-base snapshot the answers came from, not only question-answer pairs, and they release theirs [7]. Without one, an edited article silently breaks offsets and claims; see pinned corpus snapshots for RAG evaluation.
  • Required claims, not reference wording. Evidence and claims matter more than the reference answer's exact wording, because they let differently phrased answers score as correct [6]. Keep each claim atomic: one fact checkable against one passage.
  • Alternative evidence sets. A fact often appears in several valid places, such as a policy article and an FAQ. Listing each supporting set stops a system being penalized for citing the other one. TREC RAGTIME does the same for answers: every nugget is an OR nugget, so covering any one acceptable answer earns full credit for that nugget [5].
  • Answerability, scope and stale facts. Label items the corpus cannot answer (unanswerable questions in RAG evaluation), record the asker's access scope where documents carry permissions (permission-aware RAG evaluation), and tie each forbidden claim to its superseded source.

Choosing evidence granularity: document, passage, span or quote

Store evidence at the finest stable level, a character span inside a hashed document version plus the verbatim quote, and derive coarser labels from it when you score. Coarser labels are cheaper but tie the set to one chunking scheme or hide which sentence supports the claim.

GranularityExampleSurvives re-chunkingSupports claim-level citation checksMain weakness
Document IDkb-1042YesNoA long document counts as a hit even if the answering section was never retrieved
Document ID and versionkb-1042@v12YesNoSame, though stale versions stop counting
Chunk or passage IDkb-1042#c07NoPartlyStale whenever chunk size, overlap or parser changes
Character span with quotekb-1042@v12:1180-1260If text extraction is fixedYesOffsets shift if normalization changes; the quote detects it
Evidence string onlyThe quoted sentenceYesYesAmbiguous when text repeats across documents

Classic IR labels sit at the document level. A TREC qrels line holds a topic number, an iteration field, a document number and a relevance value, and NIST stresses that the document collection and qrels must match when runs are scored [8]. FinanceBench stores evidence strings instead [4]. Spans give you both: any chunk overlapping a gold span is relevant under that chunking, so two chunkers can be compared on one gold set.

Offsets refer to extracted text, so record the parser version for PDFs, slides and spreadsheets. Benchmarks already mix formats: RAG-Multi-Corpus, described in a 2026 preprint, spreads 786 curated query-answer pairs with ground-truth citations over 236 PDF, Markdown, HTML, DOCX and PPTX documents from five fictional organizations [9].

Building triples source-first from real questions

Start from a question and its authoritative evidence, decide whether the corpus answers it, and write claims and the reference answer last. The OneUptime guide sets out this order and warns that asking an LLM to invent both a question and its "truth" without verification produces a second model output, not ground truth [6].

  1. Capture the question in the asker's words, with its source record ID and date.
  2. Search the frozen snapshot for every passage that answers it, not only the first, and record each as an evidence set.
  3. Label answerability: answerable, partially answerable, unanswerable, or answerable only from a superseded version.
  4. Split the answer into atomic required claims linked to evidence, add forbidden claims from stale or conflicting documents, and write the reference answer from the claims.
  5. Have a second annotator re-derive spans and claims blind, then adjudicate differences.
Question sourceStrengthWeakness
Support tickets where the agent linked a knowledge articleReal wording plus a candidate evidence linkThe link may be stale, partial or wrong; tickets carry personal data
Internal Q&A threads, help-desk chats, search logsReal questions, including ones the documents never answerThe answer often sits in the thread, not the corpus
Expert-written questions from documentsCan target multi-document, numeric and conditional casesQuestions written while reading a passage tend to reuse its wording, which flatters keyword retrieval
LLM-generated questions from chunksCheap; Ragas offers synthetic test generation [1]Only the seed chunk is labeled, so other passages that answer the question score as misses

Tickets with linked articles make strong seeds because the link is already a relevance signal; see support tickets linked to knowledge articles. One of WixQA's three QA datasets is synthetic, generated from each knowledge-base article [7]; read when generated test sets mislead before leaning on that route, and building a golden dataset from business records for the wider process.

Worked example: scoring one answer against a triple

Scoring each part of the triple separately turns a vague "partly correct" verdict into a diagnosis that names the component to fix. Here one system response is scored against the record above.

Illustrative example: invented to show structure; it does not describe an available dataset.

The retriever returns, in order: kb-1042@v9 (superseded), kb-0877@v2, kb-1042@v12, kb-3301@v1, kb-0150@v6. The system answers: "Yes. Admins can transfer the annual plan to any workspace on the same billing account (cites kb-1042@v12), and the move triggers a prorated refund (cites kb-1042@v9)."

CheckResultReading
Claims with a full evidence set in the top 51 of 2: e1 at rank 3; e2 missingPricing article not retrieved
Same check, version-blind document IDsFirst "hit" at rank 1Superseded v9 counts, overstating retrieval
Required claims in the answerc1 present, c2 absentMissing claim matches missing evidence
Forbidden claimsf1 presentStale content reached the answer
Citation support1 of 2 citations points to current text that states its claimThe other cites superseded text

The fix belongs in retrieval (version filtering and recall on pricing content), not the prompt; a judge comparing the answer only with the reference would have reported a partial match and pointed nowhere. Building such items on purpose is covered in evaluating RAG on versioned, outdated and conflicting documents.

Exporting triples to Ragas, TREC qrels and nugget scorers

Keep one canonical triple file and generate each tool's input from it, so the gold data never forks. JSON Lines suits the canonical file: the specification requires UTF-8, one JSON value per line and a newline terminator [10], so records stream, diff and append cleanly.

  • Ragas rows. Map the question text to question and the reference answer to ground_truth; your pipeline fills contexts and answer [1]. Pin the library version: the cited page documents v0.1.21.
  • Qrels. Write one line per question and evidence document in the four-column TREC layout [8], using doc@version as the document number so stale versions cannot score. NIST's classic collections use binary judgments [8]; grading full versus partial support is your own design choice. For full test collections, see building a private domain retrieval test collection.
  • Nugget and claim scorers. Required claims map to nuggets. In RAGTIME, a nugget counts only if the report sentence is also grounded in its citations, and long topics were scored by a Llama3 70B Instruct judge working from human-curated nuggets [5]. Validate any such judge against answers labeled as grounded or unsupported.
  • Corpus manifest. List document ID, version, hash, MIME type, parser version and permission tags, so the snapshot can be rebuilt.

What public RAG sets carry, and where they stop

Public sets are useful for testing a harness and comparing scorers, but none uses your corpus, your users' wording or your access rules, and several carry license limits. The license and availability notes below reflect the cited pages as of October 2026; confirm current terms before reuse.

SetGold evidenceCorpusNote for buyers
RAG QA Arena (annotated)Citation markers; gold document IDs; faithful and unfaithful answers [2]Described on the dataset cardConfirm the license there before reuse
WixQAQA pairs plus the knowledge-base snapshot behind them [7]Customer-support knowledge baseClose public analog for support RAG
EnterpriseRAG-BenchAnswers scored with the IDs of documents relied on; distractors such as half-finished drafts and near-duplicate pages [11][12]A generated companyMIT license, per the vendor-hosted project page [12]
FinanceBenchAnswers and evidence strings [4]Public company filingsOpen subset only; full set licensed from the vendor, per its documentation [13]
TREC 2025 RAGRelevance judgments and nuggets as separate files [3]MS MARCOMicrosoft states MS MARCO datasets are for non-commercial research only [14]

BEIR's authors list dataset licenses from CC BY 4.0 and Apache 2.0 to share-alike CC BY-SA and non-commercial CC BY-NC, and the original authors of 4 of its 19 datasets report no license [15]. Check terms in which public retrieval datasets allow commercial use, and weigh exposure risk with private evaluation sets vs public benchmarks.

Acceptance checks for a delivered triple set

Accept a triple set only after mechanical checks pass on every record and a human audit passes on a sample, so broken links between claims, evidence and corpus surface before anyone reviews answer quality.

  • Every evidence quote equals the snapshot text at its offsets, and every doc@version exists in the corpus manifest with a matching hash.
  • Every required claim links to at least one evidence set; claims needing outside knowledge are flagged.
  • Unanswerable items record the search that failed to find evidence, and each forbidden claim names its stale or conflicting source.
  • A second annotator re-derives claims and spans on a random sample, with disagreements adjudicated. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets and showed such errors can change model rankings [16]; see gold-label audits for delivered eval sets.
  • Near-duplicate questions are grouped so one popular fact does not dominate a slice.
  • Each slice you plan to compare holds enough items; an ICML 2025 position paper argues that CLT-based confidence intervals tend to be too narrow on evals with fewer than a few hundred items [17].
  • The same surrogate replaces a name or account number in the question, the documents and the reference answer, or spans and claims stop matching. One guide warns that replacing every entity with REDACTED can destroy coreference, formatting or retrieval behavior [18]; see de-identifying evaluation data without breaking the test.
  • The license covers evaluation use of questions, answers and corpus, retention of the snapshot, and publication of scores; see evaluation-only data license terms.

Specifying a triple set built on licensed business records

When your own logs and documents cannot supply enough real questions, specify the corpus, the question source and the triple format together, because the evidence must come from the same snapshot the questions were asked against.

  • Corpus: document types (knowledge-base articles, policy pages, manuals, internal wikis), formats, date range, version history and permission metadata.
  • Questions: origin (linked tickets, internal Q&A, search logs), minimum share of real questions, and slices such as unanswerable, multi-document and numeric.
  • Evidence and claims: span-level evidence with quotes, versions and hashes; alternative evidence sets; atomic required and forbidden claims; double-annotation share and adjudication rule.
  • Privacy and rights: de-identification method and surrogate consistency; evaluation use, snapshot retention, publication of results, and whether items may be sent to hosted model APIs.
  • Delivery: triple file, corpus documents, manifest and derived qrels; see writing an evaluation dataset specification.

SourceX sources operational datasets from US companies, including documents and support histories, and manages the licensing process. It sources on request rather than from stock, so a request does not guarantee a match. Every dataset goes through rights review and is delivered under a license that defines the records included, permitted uses, term and delivery.

Personal details are removed or replaced before delivery and the method is recorded per dataset, so raise surrogate consistency during diligence. To start, describe the corpus and question source your triples need, or see the overview of RAG evaluation datasets from real company documents and knowledge retrieval evaluation. Training a model to cite from the same records is a different licensed use (retrieval-augmented fine-tuning data); the evaluation and benchmarking datasets hub maps the cluster.

Need real questions and documents for RAG evaluation?

Describe the corpus, the questions and the evidence your triples must carry, along with the uses you need licensed. SourceX looks for US companies that hold such records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Describe your RAG evaluation data needs.

Sources

  1. Ragas, "Prepare your test dataset" (v0.1.21 documentation). https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
  2. Hugging Face (rajistics), "rag-qa-arena" dataset card README. https://huggingface.co/datasets/rajistics/rag-qa-arena/blob/main/README.md
  3. NIST, Text REtrieval Conference, "2025 RAG Search Track" (2025). https://trec.nist.gov/data/rag2025.html
  4. Patronus AI, Contextual AI and Stanford researchers, "FinanceBench: A New Benchmark for Financial Question Answering," arXiv (2023). https://arxiv.org/abs/2311.11944v1
  5. arXiv, "Overview of the TREC 2025 RAGTIME Track" (2026). https://arxiv.org/pdf/2602.10024
  6. OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
  7. arXiv, "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/abs/2505.08643
  8. NIST, Text REtrieval Conference, "Data - English Relevance Judgements." https://trec.nist.gov/data/reljudge_eng.html
  9. arXiv, "Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems" (2026). https://arxiv.org/pdf/2604.04936
  10. jsonlines.org, "JSON Lines." https://jsonlines.org/
  11. arXiv, "EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge" (2026). https://arxiv.org/pdf/2605.05253
  12. Onyx, "EnterpriseRAG-Bench" (vendor project page). https://onyx.app/enterpriserag-bench
  13. Patronus AI, "FinanceBench" (vendor documentation). https://docs.patronus.ai/docs/financebench-1
  14. Microsoft, "MS MARCO Datasets." https://microsoft.github.io/msmarco/Datasets.html
  15. arXiv, "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
  16. Northcutt, Athalye and Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks," NeurIPS Datasets and Benchmarks (2021). https://arxiv.org/abs/2103.14749
  17. ICML, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
  18. OneUptime, "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data