Data quality, coverage and contamination
Knowledge-Corpus Quality for RAG: Duplicates, Stale Versions and Conflicting Documents
Quick answer
RAG knowledge base data quality comes down to three defects that retrieval amplifies: duplicate chunks that crowd the top-k, superseded versions that still rank as current, and documents that contradict each other. Fix them before indexing. Hash and cluster near-duplicates into canonical records, attach effective-date, status and superseded-by metadata so retrieval can prefer the current version, scan topic clusters for conflicting claims, and then measure retrieval per topic cluster against a pinned snapshot so you can see whether the cleanup worked.
By SourceX Editorial · Updated
Why corpus hygiene is a different problem from training-data quality
Corpus hygiene matters more for retrieval than for training because a RAG system quotes one retrieved passage as evidence, so a single stale or contradictory chunk can become the answer. Training-data quality work, covered in the data quality cluster hub, usually asks whether a dataset improves a model on average. A retrieval corpus is judged per query: if the 2022 refund policy and the 2025 refund policy both match "refund window for annual plans", the generator may cite either one.
Research on knowledge conflicts separates context-memory conflicts (retrieved text versus what the model learned in pretraining), inter-context conflicts (retrieved passages that disagree with each other) and intra-memory conflicts [2]. Corpus cleanup mainly targets inter-context conflicts, because those are created by the documents you control. Benchmarks of evolving knowledge show that models handle updated and conflicting facts poorly when old and new statements coexist, which is the normal state of an enterprise wiki [1].
The source systems make this worse. Confluence keeps page history, SharePoint libraries keep major and minor versions, Zendesk Guide articles carry translations and per-locale update times, and policy PDFs are often re-exported under new filenames. A naive crawler ingests all of it as peers.
Duplicate chunks: exact copies, near-duplicates and boilerplate
Deduplicate at both document and chunk level before embedding, keeping one canonical record per cluster and a pointer from every copy. Deletion without a pointer loses provenance and breaks citation links back to the original source.
Work through three layers:
- Exact duplicates. Normalize whitespace, Unicode (NFKC), case and markup, then hash with SHA-256. This catches the same PDF in three folders and the same article exported twice.
- Near-duplicates. Compute MinHash signatures over word shingles and use locality-sensitive hashing to find candidate pairs, then confirm with Jaccard similarity. This is the approach used to deduplicate language-model corpora, where researchers found many near-duplicate documents and long repeated substrings in widely used datasets [4]. The method itself is covered in depth in MinHash and LSH near-duplicate detection.
- Boilerplate chunks. Legal footers, navigation, "Was this article helpful?" widgets and repeated disclaimers produce chunks that match almost every query weakly. Strip them in parsing, or tag them so they are excluded from the index.
The main failure mode is over-merging. Templated knowledge-base articles for two product tiers can be 90% identical and differ only in a limit or a price, which is the part users ask about. Set thresholds per document type, inspect a sample of each similarity band, and never merge across different product, region or locale values. Duplicate clusters that differ only in date are not duplicates; they are a version chain and belong in the next step.
Stale versions: version chains and retrieval-time filtering
Every document that can be superseded needs an explicit status, an effective date and a superseded-by link, and retrieval should filter or down-weight anything that is not current. Without those fields, a vector index has no way to know that version 4 replaced version 3.
Build version chains from source-system history where it exists (page version IDs, document library version labels, article updated_at), and from filename and title patterns where it does not ("_v3_final", "2024 edition"). Assign each chain one current member. Where a document has no reliable date, record that as a defect instead of guessing.
At query time, the practical options are a metadata filter (status = current), a time-decay boost toward newer documents, or routing time-sensitive questions to a designated authoritative source [6]. Do not delete superseded documents outright: auditors, support agents and some users legitimately ask what the policy was on a past date, so keep retired versions in the store with status = superseded and an explicit date range. FreshQA results show why currency matters: models struggle on fast-changing and false-premise questions and improve substantially when given current search evidence [3].
The metadata a licensed RAG corpus should ship with page lists the fields to require from a supplier. This page covers what to do with them once you have them.
Conflicting documents: detecting contradictions across a corpus
Conflict detection works best when you narrow comparisons to documents about the same entity or topic and compare extracted claims, not whole documents. Pairwise comparison of every chunk is too expensive and too noisy.
A workable pipeline:
- Cluster documents by topic, product or policy area using embeddings plus metadata such as
productanddoc_type. - Within each cluster, extract checkable claims: numbers, dates, thresholds, eligibility rules, step orders, contact routes and named owners.
- Compare claims that share a subject with a natural-language inference model or an LLM judge, labeling pairs as consistent, contradictory or unrelated.
- Resolve each contradiction by version (one supersedes the other), by scope (different region, tier or audience, so add the missing metadata) or by escalation to the document owner.
Most "conflicts" in operational knowledge bases turn out to be missing scope metadata: the EU and US returns policies are both correct. True contradictions are usually a version chain that was never closed. Track how many of each you find, because the ratio tells you whether to invest in metadata or in content governance. Evolving-knowledge research also suggests organizing facts by time and source instead of treating the corpus as a flat bag of passages [1].
Expect some residual conflict. Models fine-tuned with distractor documents in context get better at ignoring irrelevant passages [5], but no model reliably picks the right side of a contradiction it cannot see the date of.
Measuring knowledge-base quality before and after cleanup
Measure retrieval quality per topic cluster, not just in aggregate, because a cleanup can raise the overall score while breaking a small but critical area such as billing or safety procedures [7]. Pair corpus-level defect metrics with retrieval metrics on a fixed question set.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Metric | How to compute | What it tells you |
|---|---|---|
| Exact duplicate rate | Documents sharing a normalized SHA-256 hash / total documents | Ingestion and export hygiene |
| Near-duplicate cluster count and max size | MinHash/LSH clusters above the chosen Jaccard threshold | Top-k crowding risk |
| Boilerplate chunk share | Chunks tagged as footer, navigation or disclaimer / total chunks | Parsing quality |
| Dated-document coverage | Documents with a valid effective_date / total | Whether recency filtering is possible |
| Open version chains | Chains with zero or more than one current member | Stale-answer risk |
| Conflict pairs per topic cluster | Contradictory claim pairs after scope resolution | Governance backlog |
| Recall@k and answer accuracy per topic cluster | Fixed question set, pinned corpus snapshot | Whether cleanup helped retrieval |
| Superseded-hit rate | Share of top-k results with status = superseded for current-state questions | Effectiveness of filters and boosts |
Run the comparison on a pinned corpus snapshot for reproducible RAG evaluation, and add deliberately stale and contradictory items to the test set using the methods in evaluating RAG on versioned, outdated and conflicting documents. If you need hard negatives for that test set, see superseded versions, drafts and near-duplicates for retrieval testing.
An audit record for each document in the corpus
Store the outcome of every hygiene check on the document record so retrieval filters, evaluations and owners all read the same state. A per-document audit record also makes the cleanup repeatable on the next ingest.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "kb-refunds-annual-v4",
"source_system": "wiki",
"source_version": "page_version_17",
"doc_type": "policy",
"product": "subscriptions",
"region": "US",
"locale": "en-US",
"status": "current",
"effective_date": "2025-03-01",
"supersedes": "kb-refunds-annual-v3",
"superseded_by": null,
"last_reviewed": "2026-06-15",
"owner_role": "billing-operations",
"content_sha256": "<hash of normalized text>",
"near_dup_cluster": "nd-0412",
"canonical": true,
"boilerplate_chunks_removed": 3,
"conflicts": [
{"with": "kb-faq-billing-v2", "claim": "refund window", "resolution": "scope: FAQ applies to monthly plans"}
],
"deid_applied": false
}
Keep the record in your document store or vector database metadata and export it with each snapshot. If the corpus is licensed, ask the supplier for equivalent fields in a machine-readable form; a format such as Croissant describes dataset files and record structure in JSON-LD, so the audit fields can be declared as part of the record schema [8].
Cleanup order of operations before indexing
Run the hygiene steps in a fixed order, because each step depends on the one before: deduplicating before building version chains merges versions you need to keep apart. A typical sequence:
- Parse and normalize: extract text and structure, strip boilerplate, keep headings and table structure.
- Exact dedup on normalized hashes; record pointers from copies to the canonical record.
- Near-duplicate clustering, with per-type thresholds and a sampled manual check of each band.
- Version-chain resolution: split date-differing clusters into chains, set one
currentmember, recordsuperseded_by. - Scope metadata fill:
product,region,locale,audience, so legitimate differences stop looking like conflicts. - Conflict scan within topic clusters; route unresolved pairs to owners.
- De-identification where needed, using methods that keep retrieval working (see de-identifying a RAG corpus without breaking retrieval).
- Chunk, embed and index, carrying all metadata onto each chunk.
- Regression evaluation per topic cluster on the pinned snapshot; publish a short dataset quality report with defect counts and known gaps.
Repeat steps 2 through 9 on every incremental ingest, not just the first build. Knowledge bases drift weekly, and a clean index decays as soon as the next export lands.
Checks for a licensed knowledge corpus
When the corpus comes from another company rather than your own systems, the same defects arrive with less context, so ask for the evidence up front. Useful questions for a supplier of knowledge-base articles or SOP and knowledge-base datasets:
- Does the delivery include version history, or only current documents? Is each document marked current, draft or retired?
- Are effective dates and last-reviewed dates populated from the source system, and what share are missing?
- Was deduplication run, at what level, and with what thresholds? Is there a pointer table from removed copies to canonical records?
- How were personal details handled, and does the redaction keep entity references consistent enough for retrieval?
- Is there a quality report listing known contradictions or gaps?
SourceX sources operational datasets, including documents and support histories, from US companies on request, and prepares diligence materials on source, rights, preparation and allowed use for each dataset. Personal details such as names, emails and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. For use-case context, see RAG evaluation datasets from real company documents and the knowledge base glossary entry. Buyers can describe the corpus they need on the SourceX buyers page.
Sourcing a knowledge corpus for RAG
If your retrieval system needs real operational documents such as SOPs, policies or knowledge-base articles, describe the data rather than a specific business. SourceX looks for US companies that hold it, reviews rights and consents, and delivers only under a license that defines records, uses, term and delivery, with every release approved by the supplying company; a request does not guarantee a match. Describe the knowledge corpus you need.
Sources
- arXiv, "Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs" (2025). https://arxiv.org/html/2506.07270v1
- arXiv, "Knowledge Conflicts for LLMs: A Survey" (2024). https://arxiv.org/html/2403.08319
- arXiv, "FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation" (2023). https://arxiv.org/pdf/2310.03214
- arXiv (ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
- arXiv, "RAFT: Adapting Language Model to Domain Specific RAG" (2024). https://arxiv.org/pdf/2403.10131
- tianpan.co, "Temporal reasoning failures in production AI systems" (2026). https://tianpan.co/blog/2026/03/16/temporal-reasoning-failures-production-ai-systems
- tianpan.co, "Long-tail query coverage and RAG retrieval gaps" (2026). https://tianpan.co/blog/2026/05/08/long-tail-query-coverage-rag-retrieval-gaps
- arXiv (MLCommons), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.