Skip to content

Evaluation and benchmarking datasets

Evaluating RAG on versioned, outdated and conflicting documents

Quick answer

To evaluate RAG on outdated documents, build a corpus that deliberately keeps superseded and conflicting versions next to the current one, label each document with effective dates and supersedes links, and write questions whose correct answer depends on choosing the current authority. Score two things separately: whether retrieval surfaced the current version, and whether the answer states its value without repeating the superseded one. Generic faithfulness scores miss this failure because a stale answer can be perfectly grounded.

By SourceX Editorial · Updated

Why stale-version failures pass ordinary RAG metrics

A stale answer usually passes groundedness checks, because the old policy really is in the retrieved context. Faithfulness metrics ask whether the answer is supported by retrieved text, not whether that text is still in force, so a system that quotes the 2023 travel policy verbatim scores as faithful even when the 2025 revision changed the per-diem. Researchers call this an inter-context knowledge conflict: two retrieved passages disagree, and the model must decide which one governs.

Enterprise corpora make this the default state rather than an edge case. SharePoint libraries keep "Policy_v3_FINAL" next to "Policy_v4", Confluence spaces hold archived pages that still rank well, and contract repositories contain an MSA plus three amendments that each override different clauses. Questions whose answers change over time are a known weak spot for language models, and retrieval does not fix that if the index serves the wrong version.

If your current eval set was built from a clean, deduplicated snapshot, it cannot detect this failure. See faithfulness evaluation sets for the grounded-versus-unsupported layer; this page covers the layer above it: grounded in the wrong version.

Version metadata the corpus must carry

Each document in a version-aware eval corpus needs machine-readable fields that say when it applied and what replaced it. Without them, you cannot write a gold label for "current," and you cannot tell a retrieval miss from a ranking error. Treat the fields below as the minimum; teams commonly add jurisdiction or business-unit scope because two documents can both be current for different audiences.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy the eval needs it
doc_idHR-TRAVEL-POLStable identity across versions
version_idHR-TRAVEL-POL@v4The unit retrieval is scored on
effective_from / effective_to2025-03-01 / nullDefines which version governs an as-of date
supersedesHR-TRAVEL-POL@v3Explicit chain; do not infer from file names
statuscurrent, superseded, draft, withdrawnDrafts and withdrawn documents are distinct distractors
authority_rankpolicy > FAQ > emailResolves conflicts between document types
scopeUS employeesTwo versions can both be current for different scopes
changed_sections§4.2 per-diemLets you target questions at the delta

The supersedes chain is a working hypothesis for most buyers: real document systems rarely record it, so someone has to reconstruct it from approval logs, amendment numbering or change history. Ask how the supplier derived it and what fraction was verified by a person. For corpus-level packaging, a Croissant JSON-LD description can carry these as record-level fields so loaders and eval harnesses read them consistently [5]. The RAG corpus metadata requirements page lists the broader field set a licensed corpus should ship with.

Designing questions that only the current version answers

Good items target the delta between versions, not content that stayed the same. If v3 and v4 share 95% of their text, a question about an unchanged clause tests nothing about version handling. Pull the diff, then write questions about each changed value, added obligation or removed exception.

Use several item types so you can diagnose where the system fails:

  • Value-changed items. "What is the domestic per-diem?" The current version says one number, the superseded version another.
  • As-of items. "What per-diem applied to a trip in June 2024?" Here the superseded version is the correct authority, which catches systems that simply prefer the newest file.
  • Amendment-stack items. "What is the termination notice period under the MSA?" The answer lives in amendment 2, which overrides the base agreement, while amendment 3 is silent.
  • Removed-content items. The current version deleted a benefit. The correct answer says it no longer applies; reciting the old benefit is a failure, and so is claiming it never existed.
  • Draft and withdrawn traps. An unapproved draft with a plausible new value sits in the index. Citing it is a failure even though it is the newest file.
  • Unresolvable conflicts. Two current documents with equal authority disagree. The correct behavior is to surface both and flag the conflict, which overlaps with abstention testing on unanswerable questions.

As-of items matter most for temporal RAG evaluation, because they separate "understands effective dates" from "sorts by modified time." Borrow the false-premise pattern from temporal QA benchmarks too: include questions whose premise is only true under the old version, such as asking how to claim a benefit that was removed.

Near-duplicate distractors and why they are the hard part

Superseded versions are the hardest distractors you can put in a RAG index because they are near-duplicates of the gold document. Recent benchmark work deliberately mixes multiple gold documents with a dense spectrum of semantically similar distractors, so that retrieval has to discriminate fine differences rather than topic [1]. Versioned enterprise documents give you that spectrum naturally: v1 through v4 of the same policy, plus a regional variant and an FAQ that paraphrases v2.

Do not deduplicate these away. Training-data pipelines routinely run MinHash or suffix-array deduplication because common corpora are full of near-duplicates [4], and RAG ingestion pipelines often reuse the same step. For a version-conflict eval, that step deletes the test. Run deduplication only on exact copies with identical version_id, and log what was removed.

Calibrate distractor difficulty. Measure embedding similarity between each gold chunk and its superseded counterpart, then stratify items into bins so you can report accuracy separately for near-identical versions and loosely related ones. The corpus quality page on duplicates and stale versions covers cleaning a production index; the eval corpus should keep the mess on purpose.

Scoring: retrieval of the current version vs answer correctness

Score retrieval of the governing version and answer correctness as separate metrics, because they fail independently. One benchmark design reports evidence access separately from answer quality for exactly this reason [1]: a system can retrieve v4 and still answer from v3 if both are in context, or retrieve only v3 and get lucky.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "travel-perdiem-007",
  "question": "What is the domestic per-diem for US employees?",
  "as_of": "2026-09-30",
  "gold_version_ids": ["HR-TRAVEL-POL@v4"],
  "superseded_version_ids": ["HR-TRAVEL-POL@v3", "HR-TRAVEL-FAQ@2023"],
  "required_claims": ["per-diem is the v4 amount", "cites HR-TRAVEL-POL v4 or its effective date"],
  "forbidden_claims": ["states the v3 amount as current", "cites the 2023 FAQ as authority"],
  "acceptable_behaviors": ["notes the amount changed on 2025-03-01"],
  "item_type": "value_changed",
  "distractor_similarity_bin": "high"
}

Required and forbidden claims are a practical way to encode "do not state the superseded value" when there is no single reference answer [2]. A grader, human or model, checks each claim independently, so an answer that gives the new value but also says "previously X, still applicable for some cases" is caught.

Report at least these metrics per item type and distractor bin:

  1. Current-version recall@k: share of items where any gold_version_ids chunk is in the top k.
  2. Superseded intrusion rate: share where a superseded_version_ids chunk outranks the gold chunk.
  3. Forbidden-claim rate: share of answers that state any forbidden claim, regardless of retrieval.
  4. Correct citation rate: answer cites the governing version, not just any version of the document.
  5. Conflict disclosure rate: on unresolvable items, share that surface both sources.

A system with high recall but high forbidden-claim rate has a generation problem (the model ignores dates in context). Low recall with low intrusion points to chunking or metadata filters dropping the current file. The question-answer-citation triples guide covers the citation layer in more depth.

Sourcing real versioned document sets

Real version histories are hard to synthesize convincingly, which is why buyers look for operational documents with their change history intact. Generated "v2" documents tend to change one obvious number, while real revisions move clauses, rename sections, add exceptions and leave stale cross-references. Public benchmarks also rarely include internal enterprise content; work such as EnterpriseRAG-Bench targets company internal knowledge [3], but your own policy, contract and runbook structures will differ.

Useful source material includes knowledge-base articles with edit history, policy libraries with approval dates, contract sets with amendments, engineering runbooks and API docs across releases, and support macros that were retired. When you evaluate a supplier, ask for:

  • Version chains with dates, and how supersedes was established.
  • Whether drafts and withdrawn documents are included and labeled.
  • How personal details in older versions were handled, since stale documents often carry names and contact details that later versions removed.
  • Whether the license allows evaluation use and retention of the eval set across model versions.
  • Pinned snapshots, so results stay comparable; see pinned corpus snapshots.

SourceX sources operational datasets from US companies on request, including documents, support histories and engineering records, and manages licensing; a request does not guarantee a match. You can describe the version-history characteristics you need on the SourceX buyer page. For document-based RAG use cases more broadly, see RAG evaluation datasets from real company documents and licensing knowledge base articles.

Common construction mistakes

Most failed version-conflict evals fail because the corpus or labels quietly removed the conflict. Watch for these:

  • Using file modified time as effective_from. A re-saved v2 looks newer than v4. Use approval or publication dates.
  • Labeling against today only. Without as_of, the gold label rots when a v5 lands; fix the as-of date per item.
  • Leaking the answer through metadata. If the retriever sees status: superseded but production will not, you are testing a filter, not the model. Run with and without metadata exposed.
  • Asking about unchanged text. Items must target the diff.
  • Single-score reporting. One accuracy number hides whether retrieval or generation failed.
  • Testing only one conflict type. Context-memory conflicts, where the model's training knowledge disagrees with a current document, behave differently from conflicts between two retrieved documents; include both.

For allocating items across rare high-stakes conflicts, such as safety procedures or regulated disclosures, see stratified evaluation sets. The evaluation datasets hub maps the other eval types this set sits beside.

Sourcing version-history documents for RAG evaluation

SourceX finds US businesses that hold the operational documents you describe, assesses data and licensing permissions, and, once a supplier agrees, delivers each dataset under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, and every release is approved by the supplying company. Describe the version chains and conflict types you need on the SourceX buyer page.

Sources

  1. arXiv, "Beyond the Needle’s Illusion: Decoupled Evaluation of Evidence Access and Use under Semantic Interference at 326M-Token Scale (arXiv:2601.20276v1)" (2026). https://arxiv.org/html/2601.20276v1
  2. OneUptime, "How to Build Ground Truth for RAG Evaluation When No Reference Answers Exist" (2026). https://oneuptime.com/blog/post/2026-08-31-build-rag-ground-truth-without-reference-answers/markdown
  3. arXiv, "EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge" (2026). https://arxiv.org/pdf/2605.05253
  4. Lee et al., arXiv / ACL 2022, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  5. Akhtar et al. (MLCommons), arXiv, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data