Skip to content

Evaluation and benchmarking datasets

Faithfulness evaluation sets: answers labeled as grounded or unsupported

Quick answer

A RAG faithfulness evaluation set is a collection of (question, retrieved context, generated answer) records where humans have labeled whether each answer, or each claim inside it, is supported by the supplied context. Its job is to test the scorer, not the RAG system: you use it to measure whether a groundedness metric, NLI checker or LLM judge actually catches unsupported statements. A useful set carries claim-level labels, a deliberate mix of faithful and unfaithful answers, and error types that resemble your production failures.

By SourceX Editorial · Updated

Why labeled answers are a different asset from gold Q-A-citation triples

Faithfulness sets validate detectors, while gold triples score systems, so the label sits on a model output rather than on a reference answer. A question-answer-citation triple tells you what a correct answer and its supporting passage look like. A faithfulness record instead freezes one specific generated answer next to the exact context the generator saw and records a human verdict on whether that answer stays inside the evidence.

The distinction matters because metrics such as Ragas faithfulness take the context and answer fields as inputs and never consult a ground-truth answer for that score [1]. If you only own gold triples, you can compute faithfulness numbers but you cannot tell whether those numbers are right. The labeled set is the calibration layer underneath every automated groundedness dashboard, and it belongs alongside the broader LLM evaluation datasets map as its own procurement line.

Faithfulness is also narrower than correctness. An answer can be faithful to a retrieved passage that is outdated or wrong, and an answer can be factually true yet unsupported because the fact came from model memory. Label definitions must say which property you are measuring, and most groundedness work measures support by the context only.

Claim-level versus answer-level labels

Claim-level labels are the default for serious detector evaluation, because answer-level verdicts hide where and how often a model drifts from its evidence. An answer-level label (faithful, unfaithful) is cheap and works for a pass/fail release gate. It cannot distinguish a response with one minor unsupported date from one that invents a policy, and it gives a detector credit for flagging the right answer for the wrong span.

Granularity options run from word spans to whole answers. Span-level labels mark the exact tokens that are unsupported and suit products that highlight risky text; claim-level labels tie each atomic statement to a verdict and evidence chunk; answer-level labels serve release gates. Practitioner guidance offers a lighter route for teams without reference answers: list required and forbidden claims per question so faithfulness can be checked claim by claim [3].

A workable hierarchy for most teams is three levels. Decompose each answer into atomic claims, label each claim supported, contradicted, or not-in-context, then roll up to an answer label with an explicit rule (for example, any contradicted or not-in-context claim makes the answer unfaithful). Keep the rollup rule in the dataset card so a later reader can recompute it.

What a faithfulness record needs to contain

Every record should let a third party reproduce the label from the stored fields alone, which means the exact context the generator saw, not a pointer to a live index. Retrieval indexes change; a record that stores only document IDs will silently drift as chunks are re-embedded or re-split.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "faith-000417",
  "question": "What is the refund window for annual plans?",
  "contexts": [
    {"chunk_id": "kb-policy-12#c3", "doc_version": "2026-03-01",
     "text": "Annual plans may be refunded within 30 days of purchase. Monthly plans are non-refundable."}
  ],
  "answer": "Annual plans can be refunded within 30 days, and monthly plans within 7 days [1].",
  "generator": {"model": "internal-rag-v4", "prompt_template": "support_rag_v2", "temperature": 0.2},
  "claims": [
    {"claim_id": 1, "span": [0, 47], "text": "Annual plans can be refunded within 30 days",
     "label": "supported", "evidence": ["kb-policy-12#c3"]},
    {"claim_id": 2, "span": [49, 79], "text": "monthly plans within 7 days",
     "label": "contradicted", "evidence": ["kb-policy-12#c3"],
     "error_type": "fabricated_detail"}
  ],
  "citation_check": {"marker": "[1]", "supports_all_claims": false},
  "answer_label": "unfaithful",
  "rollup_rule": "any contradicted or not_in_context claim => unfaithful",
  "annotators": ["a07", "a12"], "adjudicated": true,
  "split": "test", "slice": ["policy", "numeric"]
}

The citation_check block covers a separate failure: a citation marker attached to a passage that does not support the sentence. Some public sets already pair faithful and unfaithful answers with citation markers [2], and buyers testing attribution features should insist that citation support is labeled independently of claim support.

Balancing faithful and unfaithful answers

A detector evaluation set needs enough unfaithful answers to measure recall with confidence, so the class mix should be designed rather than inherited from production rates. If only a small share of production answers are unsupported, a naturally sampled set of a few hundred records may contain too few positives to tell two detectors apart. Over-sample unfaithful cases, then report precision and recall per class and balanced accuracy, so a detector that flags nothing cannot look strong on a skewed set.

Balance the hard negatives too. The most informative unfaithful answers are near-misses: a correct entity with a wrong number, a right policy applied to the wrong plan tier, a merged statement drawn from two chunks that individually say something different. Equally, include faithful answers that look suspicious, such as paraphrases, unit conversions, or summaries that compress several sentences, so you can measure false positives. Use stratified sampling for rare and high-risk cases to set slice quotas.

Generated errors are useful but limited. Synthetic perturbations (swapped numbers, negations, entity substitutions) can help train a checker, but evaluating on the same kind of perturbations rewards detectors for spotting the generator's artifacts. Keep the test split built from real model outputs, and read the caveats on when synthetic evaluation data misleads.

Validating groundedness metrics and judges against the labels

Treat the labeled set as the reference against which every automated scorer is accepted or rejected, with a pre-agreed agreement threshold. Run Ragas faithfulness, an NLI model, a fine-tuned checker and your LLM judge over the same frozen records, then compare each scorer to the human labels at the granularity it claims to support [1]. Include small fine-tuned checkers in the comparison rather than assuming the largest judge wins; on a fixed labeled set, cost per correct flag is directly measurable.

For LLM judges, one practitioner loop is to score 150 to 300 stratified real inputs with two or three humans using the exact rubric text the judge receives, measure agreement with Cohen's kappa, revise the rubric, and repeat [4]. Measure inter-annotator kappa first; human-human agreement is a practical ceiling for interpreting judge-human agreement, so weak annotator agreement means the guidelines need work before the judge does. The LLM-as-a-judge calibration guide covers rubric versioning in more depth.

Label noise directly limits what the set can prove. Audits of widely used test sets estimated average label error rates of at least 3.3%, enough to change model rankings [5]. Double-annotate every unfaithful label, adjudicate disagreements, and re-audit a sample whenever two detectors differ by less than the estimated noise rate.

Illustrative example: invented to show structure; it does not describe an available dataset.

Acceptance checkWhat to measureTypical decision rule
Human agreementCohen's kappa between annotators on claim labelsFix guidelines before testing any scorer if agreement is weak
Detector recallShare of contradicted and not-in-context claims flaggedSet a floor per high-risk slice, not only overall
Detector precisionShare of flags that humans confirmTrack separately on paraphrase and numeric slices
Span accuracyOverlap of flagged span with labeled claim spanRequired only if the product highlights spans to users
Citation supportAgreement on whether each marker supports its sentenceRequired for any answer UI that shows citations
DriftScore change after a judge prompt or model updateRe-run the frozen set on every scorer change

Where real unsupported claims come from

The richest source of realistic unsupported claims is likely human correction of model drafts, though this is a working hypothesis rather than a measured result. When support agents, analysts or paralegals edit an AI-drafted answer before sending it, the diff between draft and sent version often marks exactly the sentences a human judged unsupported or wrong. Paired with the knowledge-base passages the draft was grounded on, that edit history can be converted into claim-level labels with far less annotation effort than labeling from scratch.

Treat these records with care. An edit can reflect tone or policy rather than faithfulness, so a reviewer still needs to classify each deleted or changed sentence. The original retrieval context must have been logged at draft time, and the supplying organization needs the rights to license both the drafts and the underlying documents. For context on the documents themselves, see RAG evaluation datasets from real company documents and evaluation datasets built from real business work.

The same records serve training as well as evaluation. Corrected drafts are a natural input to fine-tuning data that reduces hallucinations, but split them first and keep the evaluation partition out of any training run.

Specifying and documenting the set before you buy or build

Write the specification before collecting a single record, because label definitions and slice quotas are expensive to change after annotation starts. A complete evaluation dataset specification for faithfulness should state the following.

  • Label taxonomy: supported, contradicted, not-in-context, plus whether "partially supported" exists and how it rolls up.
  • Claim decomposition rules: what counts as an atomic claim, and how hedges, quotes and numbers are split.
  • Context policy: exact chunks stored, document version dates, maximum context length.
  • Generator diversity: which models and prompt templates produced answers, so the detector is not tuned to one model's style.
  • Class balance and slice quotas, including near-miss unfaithful and suspicious-but-faithful answers.
  • Annotation protocol: number of annotators, adjudication, agreement target, error audit.
  • Splits: a held-out test partition never used for prompt tuning of the judge.

Document the result in a dataset card. On the Hugging Face Hub, a dataset's README.md is its card and carries YAML metadata such as license, language and size [6]; internal sets benefit from the same structure even when they never leave your environment.

How SourceX fits a faithfulness evaluation project

SourceX sources operational datasets from US companies, such as support and sales histories, documents and engineering records, which are the kinds of material behind corrected drafts and grounded answers. Data is sourced on request rather than held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and every release is approved by the supplying company. Buyers describe the data they need on the SourceX buyer page, and SourceX does not train models.

Request faithfulness evaluation data

Describe the documents, answer types and label granularity your groundedness evaluation needs, and SourceX looks for US businesses that hold matching operational data. Each dataset goes through Find, Assess, Agree, Transact and Manage, is delivered under a license that defines records, uses, term and delivery, and nothing is contracted until a supplier agrees. Start your request at sourcex.si/buyers.

Sources

  1. Ragas, "Prepare data (Ragas v0.1.21 documentation)". https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
  2. Hugging Face (rajistics), "rag-qa-arena dataset card (README.md)". https://huggingface.co/datasets/rajistics/rag-qa-arena/blob/main/README.md
  3. OneUptime blog, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
  4. OneUptime blog, "How to Calibrate an LLM-as-a-Judge Against Human Labels with Cohen's Kappa" (2026). https://oneuptime.com/blog/post/2026-08-31-calibrate-llm-judge-cohens-kappa/markdown
  5. arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data