Skip to content

Retrieval, RAG and grounding data

Superseded versions, drafts and near-duplicates for realistic retrieval testing

Quick answer

The most realistic RAG distractor documents are not generated; they are the drafts, superseded policy versions, forked templates and copy-pasted pages that already sit in a company's document systems. Buy corpora that keep those version chains intact with created, modified, effective and superseded dates, then label which version is authoritative for each test question at its query date. Synthetic distractor injection is a useful baseline, but it rarely reproduces how real documents drift apart over time.

By SourceX Editorial · Updated

Why synthetic distractors underestimate the stale-document problem

Synthetic distractors test whether a retriever can reject noise, but real superseded documents test whether it can reject content that was once correct. Benchmarks such as EnterpriseRAG-Bench deliberately add off-topic threads, half-finished drafts and near-duplicate pages to mimic enterprise noise [1][2], and the RAFT training recipe mixes "oracle" and distractor documents so a model learns to ignore unhelpful context [4]. Those designs are valuable, yet an injected distractor is usually different in a way the generator chose, not in the way an organization's documents actually diverge.

In a real archive, version 3 of a travel policy and version 4 can share almost all of their text and differ only in one reimbursement limit and one effective date. A dense retriever scores both nearly identically, and a BM25 index may rank the older file higher if it happens to repeat the query terms more often. Public collections such as those in BEIR were assembled for heterogeneous zero-shot evaluation [6], so they seldom contain this kind of dated version lineage.

Time also corrupts the test itself. Research on temporal conflict finds that substantial portions of common QA benchmarks contain outdated information [3], and earlier work showed language model performance degrades on text from after the training period [7]. If your gold answers do not carry a date, a retriever that returns the newer, correct document can be scored as wrong.

What a version-aware evaluation corpus has to contain

A usable corpus pairs every document with its lineage and validity period, not just its text. Without lineage, you cannot tell a near-duplicate that is a harmless copy from one that is a superseded rule. Ask suppliers for these elements, and see the metadata a licensed RAG corpus should ship with for the broader field list.

  • Stable document ID and version ID. A doc_family_id shared across all versions plus a version_id per revision, so you can group a chain even after titles change.
  • Lifecycle status. Values such as draft, in review, approved, published, superseded and archived, mapped from the source system (for example SharePoint check-in states, Confluence page history, Google Drive revisions or a document management system's approval workflow).
  • Four dates, not one. created_at, modified_at, effective_from and effective_to (or superseded_at). File-system modified time alone is unreliable because migrations and bulk exports reset it.
  • Supersession links. An explicit supersedes or superseded_by pointer, or a change-log reference, rather than an inferred one.
  • Copy provenance. Where a page was copied into another space, team wiki or email attachment, keep both copies and the path or container each came from.
  • Access scope. Group or role labels if you will also run permission-aware RAG evaluation, since the authoritative version and the visible version can differ by user.

Types of real distractors worth paying for

Not every duplicate is a useful distractor, so specify the classes you need and the share of each. The table below separates the patterns that genuinely stress retrievers from the ones that mostly add storage cost.

Distractor classHow it arises in real archivesWhat it testsTypical failure mode
Superseded versionPolicy, price sheet or SOP replaced by a newer approved revisionRecency and validity filteringRetriever returns the older version with an obsolete limit or step
Unapproved draftDraft left in a shared folder beside the finalStatus-aware rankingGenerator quotes a clause that was struck before approval
Forked templateTeam copies a master template and edits locallyScope and ownershipAnswer cites a regional variant for a global question
Exact or near-exact copySame file attached to emails, wikis and ticketsResult diversity and dedup at query timeTop-k filled with one document, crowding out the second needed source
Conflicting peer documentTwo current documents from different teams disagreeConflict detection and abstentionModel picks one silently instead of flagging the conflict
Topical near-missSame product or customer, different questionPrecision under lexical overlapKeyword match beats semantic match

Keep duplicates in the evaluation corpus, remove them from training

Deduplication policy for evaluation should be the opposite of deduplication policy for training. For training data, near-duplicate removal is established practice: Lee et al. found common language-modeling datasets contain many near-duplicates and that deduplication improves models [5]. For a retrieval test collection, the duplicates are the point, because production indexes contain them.

The practical rule is to run near-duplicate detection, such as the MinHash and LSH methods used for licensed text, to label clusters rather than to delete them. Store a near_dup_cluster_id and a similarity score per document, then report retrieval metrics both with and without cluster-level dedup at query time. If you need a clean production index instead, the separate guide on cleaning duplicates, stale versions and conflicts in a knowledge corpus covers that workflow.

Watch for one procurement trap. A supplier that deduplicates "for quality" before delivery has removed exactly what you are buying, so state in the request that version chains and copies must be delivered intact.

How to label authority for each test question

Every test item needs a query timestamp and a gold document set that is valid at that timestamp. "Which travel policy applies?" has a different correct answer in March and in November if a revision took effect in between. Record the authoritative version, the acceptable alternates (an approved copy in another space) and the explicitly wrong versions (the superseded and draft ones).

That structure lets you compute metrics standard qrels miss. Useful ones include the superseded-retrieval rate (share of queries where a superseded version appears in top-k), the stale-citation rate (share of answers citing a non-authoritative version), and draft leakage (share citing an unapproved draft). Pair these with conventional recall@k and nDCG. Teams building the full collection can follow the private domain retrieval test collection guide and, for human-graded labels, the notes on buying relevance judgments.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "query_id": "q-0412",
  "query_text": "What is the per-night hotel cap for domestic travel?",
  "query_timestamp": "2025-11-03",
  "gold": [
    {"doc_family_id": "pol-travel", "version_id": "v4", "status": "published",
     "effective_from": "2025-07-01", "effective_to": null, "role": "authoritative"}
  ],
  "acceptable_alternates": [
    {"doc_family_id": "pol-travel", "version_id": "v4-copy-wiki", "near_dup_cluster_id": "c-118"}
  ],
  "known_wrong": [
    {"doc_family_id": "pol-travel", "version_id": "v3", "status": "superseded",
     "effective_to": "2025-06-30", "reason": "older cap"},
    {"doc_family_id": "pol-travel", "version_id": "v5-draft", "status": "draft",
     "reason": "proposed cap, not approved"}
  ]
}

Where real version chains come from, and what to ask suppliers

Version chains exist wherever a business revises controlled documents: policy and procedure libraries, product manuals and release notes, contract templates and clause libraries, pricing and rate sheets, engineering design docs and runbooks, and support knowledge bases. Each source system stores history differently, so ask how the export preserved it. A flat export of "latest files" from a shared drive loses the chain entirely.

Request checklist for a version-aware retrieval corpus:

  1. Which systems were exported, and was revision history included or only current files?
  2. How are effective_from and superseded_at derived: from document fields, approval logs or inference?
  3. What share of document families have two or more versions, and what is the median chain length?
  4. Are drafts and unapproved revisions included, and are they labeled as such?
  5. Were copies across spaces, emails and tickets kept, with their container paths?
  6. Was any deduplication or "cleanup" applied before delivery?
  7. How were personal details handled, and does the method preserve the tokens your queries depend on? See retrieval-preserving de-identification.
  8. Is the license scoped for evaluation, for retrieval in production, or both?

Cross-system corpora are the hardest and most realistic case, because the same policy lives in a wiki, a PDF on a drive and a pasted chat message. The guide to multi-source enterprise search corpora covers combining those sources.

Common failure modes when buying versioned corpora

Most problems surface as silent date errors rather than missing files. Migration timestamps that overwrite original modified dates make every version look equally recent. Version numbers that restart after a system move break supersedes chains. Redaction that replaces product names or rate values with placeholders can make v3 and v4 identical, erasing the very difference you want to test.

Sample before you commit. Pull 20 document families, verify that each chain sorts correctly by effective date, diff adjacent versions to confirm the changes are real, and check that at least some test questions have answers that changed across versions. If none did, the corpus has duplicates but not superseded content.

How SourceX approaches version-history corpora

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases; the data kinds include documents, support histories, engineering records and finance and legal workflows. Nothing is held in stock, so a request describing version chains, drafts and validity dates starts a search for US businesses that hold such data, and a request does not guarantee a match. Every release is approved by the supplying company, rights-reviewed, and delivered under a license that defines records, uses, term and delivery. Buyers can review the process and describe requirements on the buyers page, and see related SourceX pages on enterprise document datasets and RAG evaluation datasets from real company documents. The retrieval and RAG data buyer's guide covers the wider cluster.

Request realistic distractor documents for retrieval testing

If your RAG evaluation needs real superseded versions, drafts and near-duplicates with dates, describe the document types and lineage fields you need rather than specific companies. SourceX assesses data and licensing permissions with candidate suppliers, and nothing is contracted until a supplier agrees. Describe your retrieval test corpus on the SourceX buyers page.

Sources

  1. Onyx, "EnterpriseRAG-Bench". https://onyx.app/enterpriserag-bench
  2. arXiv, "EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge" (2026). https://arxiv.org/pdf/2605.05253
  3. arXiv, "Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs" (2025). https://arxiv.org/html/2506.07270v1
  4. arXiv, "RAFT: Adapting Language Model to Domain Specific RAG" (2024). https://arxiv.org/pdf/2403.10131
  5. arXiv (Lee et al., ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  6. arXiv (Thakur et al.), "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
  7. arXiv (Lazaridou et al.), "Mind the Gap: Assessing Temporal Generalization in Neural Language Models" (2021). https://arxiv.org/abs/2102.01951v2

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data