Data quality, coverage and contamination
Factual Accuracy Checks for Licensed Text and Knowledge Content
Quick answer
To estimate the factual accuracy of training data, sample atomic claims from the licensed text, verify each one against a dated authoritative reference or a domain expert, and classify failures as wrong, outdated, unverifiable or internally contradictory. Report each rate with a confidence interval, separated by document type and age. Outdated content needs its own bucket because version dates and superseded flags can remove it, while wrong content cannot be fixed by metadata.
By SourceX Editorial · Updated
Knowledge-heavy text (KB articles, product documentation, SOPs, support macros and Q&A pairs) is valuable precisely because it asserts facts. That also makes it the place where a quiet error rate turns into a model that answers confidently and incorrectly. This page sits inside the training data quality assessment hub and covers content correctness only: label accuracy, duplicates and freshness metadata have their own pages.
Why factual accuracy is a separate quality check
Factual accuracy measures whether the statements in a text are true for the period and scope they claim, which is a different property from label correctness, fluency or deduplication. A recent data-centric survey of LLM training lists factual accuracy as its own metric-based check next to format, duplication, diversity and fluency, combined with LLM-based scoring and human review of samples [1]. Buyers who only run the generic checks in training data quality metrics will miss it.
Model-based quality raters do not close the gap. LLM-based scoring is now benchmarked as a data-preparation task in its own right [5], but such raters judge what a passage looks like rather than checking it against a reference. QuRating, for example, scores a "facts and trivia" criterion [9] that rates how much factual content a passage carries, not whether the facts are right. A passage packed with confidently stated, superseded API parameters can score high on that dimension and still be wrong.
The distinction from label audits matters too. Label-error studies such as the confident-learning audit of 10 benchmark test sets, which estimated an average error rate of at least 3.3% [6], check whether an assigned class matches the input. A factual check asks whether the input itself is true. For label work, use annotation quality audits; for content, use the protocol below.
How wrong facts hurt SFT, RAG and evaluation differently
The damage from factual errors depends on how the text is used, so set your tolerance per application rather than per dataset. The same 4% error rate is a nuisance in a pretraining mix, a real defect in SFT, and a direct output defect in RAG.
- SFT and instruction tuning. Q&A pairs teach the model what to assert. Gekhman et al. found that models learn fine-tuning examples carrying new knowledge slowly, and that as those examples are eventually fitted, the model's tendency to hallucinate rises [3]. Factually wrong answers are an extreme case of unfamiliar knowledge, so a small share of them can have outsized effects.
- RAG and grounding. A retrieved KB article that is wrong is usually reproduced faithfully. Faithfulness checkers will score the answer as grounded because it matches the source; the error is upstream of any groundedness metric.
- Evaluation. Wrong or outdated gold answers penalize models that are correct today. Research on question answering under temporal conflict shows that answers which were correct when written can become wrong as facts change [2], which means part of what looks like model error can be reference error.
Separating wrong, outdated and unverifiable content
In enterprise knowledge text, many factual defects are outdated rather than invented, and the two need different remedies. An article that was correct for release 4.2 and is now superseded by 5.0 is fixable with metadata; an article that never matched the product is not.
Use four failure codes and keep them apart in every report:
| Code | Definition | Typical signal in KB and doc corpora | Remedy |
|---|---|---|---|
| WRONG | False for the version, date and scope the text claims | Incorrect default value, wrong error-code meaning, wrong step order in an SOP | Exclude, correct or down-weight; cannot be fixed by metadata |
| OUTDATED | True for an earlier version or date, false now | Deprecated endpoint, retired pricing tier, old form number, superseded policy | Keep with valid_to and superseded_by fields, or exclude from current-state SFT |
| UNVERIFIABLE | No authoritative reference or expert can confirm it | Undocumented workaround, tribal knowledge in a support macro | Flag; decide per use case (often acceptable for RAG with attribution, risky for SFT) |
| CONFLICT | Contradicts another record in the same delivery | Two articles give different limits for the same feature | Resolve with version metadata; see stale versions and conflicting documents in RAG corpora |
Outdated content can only be separated from wrong content if the delivery carries dates. Ask for created_at, last_reviewed_at, product_version or applies_to, status (published, archived, deprecated) and any superseded_by link that the source system already holds. Systems such as Zendesk Guide, Confluence, ServiceNow Knowledge and Salesforce Knowledge track article state and versions natively, so these fields usually exist even when they are not exported by default. The metadata fields to require with licensed text corpora page lists the full set.
A claim-sampling protocol for estimating the error rate
The practical method is to sample documents, extract atomic claims, verify a fixed number of claims per document, and estimate rates with intervals that account for clustering. This is a hypothesis-driven audit, not a full fact-check of the corpus, and it fits inside the process framework that ISO/IEC 5259-4 sets out for data quality in training and evaluation [7].
- Stratify. Split the delivery by document type (KB article, API reference, SOP, Q&A pair), by age band from
last_reviewed_at, and by product line. Errors cluster in old, rarely viewed articles, so a simple random sample will under-represent the tail you care about. - Sample documents, then claims. Draw documents per stratum, then extract 3 to 5 atomic claims per document: a single checkable assertion such as "the default timeout is 30 seconds" or "form X is filed within 10 days." Skip opinions, instructions without a factual premise, and boilerplate.
- Fix the reference before verifying. For each stratum, name the authoritative reference: the current product spec or release notes, a dated regulation or standard, an internal source-of-truth table, or a named domain expert. Record the reference version and date on every verdict.
- Verify in two passes. Use an automated pass to triage, then human or expert adjudication. Small fact-checking models such as MiniCheck check a claim against a grounding document at a fraction of large-model cost [4], which makes them suitable for triage when you have a reference document, but not as the final verdict on specialist content.
- Adjudicate disagreements. Have a second reviewer resolve every case where the automated verdict and the first reviewer disagree, and every WRONG verdict on regulated or safety-relevant content. For specialist domains, qualify reviewers as described in verifying domain-expert annotators.
- Estimate with intervals. Report the rate per failure code with a Wilson or similar interval. Because claims within a document are correlated, compute the interval at document level or apply a design-effect correction; claim-level intervals will look tighter than they are.
- Decide. Compare the upper bound, not the point estimate, against your tolerance per application, using the acceptance logic in acceptance sampling for dataset deliveries.
Worked example: a claim-level audit record and result
A compact audit record per claim keeps verdicts reproducible and lets you rerun the estimate when the reference changes.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"claim_id": "kb-48213-c2",
"doc_id": "kb-48213",
"doc_type": "kb_article",
"doc_last_reviewed_at": "2023-03-14",
"doc_applies_to": "Billing v4.x",
"claim_text": "Refunds over 500 USD require a second approver.",
"reference": { "type": "policy_doc", "id": "FIN-POL-12", "version": "2026-01", "date": "2026-01-15" },
"auto_verdict": { "checker": "grounded_fact_checker", "label": "unsupported", "score": 0.21 },
"human_verdict": "OUTDATED",
"evidence_note": "Threshold changed to 250 USD in policy revision 2025-07; prior text matches 2023 version.",
"reviewer_id": "rev-07",
"adjudicated": false
}
Suppose a 400-claim sample drawn from 100 documents (four claims each) returns 14 WRONG and 22 OUTDATED verdicts. The point estimates are 3.5% wrong and 5.5% outdated. A Wilson 95% interval computed naively at claim level gives roughly 2.1% to 5.8% for WRONG and 3.7% to 8.2% for OUTDATED; a document-level calculation will widen both because errors cluster within articles.
The decision then splits by use. For RAG, the OUTDATED share is mostly fixable by filtering on doc_applies_to and last_reviewed_at, so the binding number is the WRONG upper bound. For SFT on current behavior, both buckets count, and a combined claim-level upper bound of about 12% (wider at document level) would normally trigger re-stratified sampling of the oldest age band before acceptance.
Questions to put to the supplier before and after sampling
Ask suppliers how correctness is maintained in the source system, because the answer predicts the error rate better than any sample of a few hundred claims. A team with a review cadence, an owner per article and an archive policy produces a different corpus from one where articles are written once and never revisited.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Question | What a useful answer contains | Red flag |
|---|---|---|
| How are articles reviewed and retired? | Review cycle, owner field, archive or deprecate states exported in the data | "Everything in the export is current" with no status field |
| Which version or date does each record apply to? | applies_to, release tags, effective dates | Only created_at, overwritten on every edit |
| Are superseded records included? | Yes, with superseded_by links, or excluded by a stated rule | Unknown, or silently mixed |
| What is the source of truth for disputed facts? | Named spec, policy register or system of record | Individual agents' notes |
| Were any answers generated or rewritten by a model? | Flag per record and the date range affected | No tracking; see detecting model-generated content |
| Can a sample of reference documents be shared for verification? | Policy or spec extracts under the same agreement | Verification only possible by trusting the supplier |
Record the answers in the dataset's documentation. Data Cards and similar formats were designed to capture upstream sources, collection methods and the decisions that affect model performance [8], and an accuracy audit summary belongs in the same place.
Failure modes that inflate or hide the error rate
Factual audits fail in predictable ways, and most of them make a corpus look cleaner than it is. Check for these before trusting a number.
- Reference drift. Verifying 2023 articles against a 2026 spec without dates labels every once-true fact as WRONG. Always pair the verdict with the reference version.
- Checker agreement mistaken for truth. An LLM verifier and an LLM-generated corpus can share the same misconception. Keep human adjudication on a fixed fraction of verdicts.
- Easy-claim bias. Extractors prefer short numeric claims. Procedural claims in SOPs (step order, prerequisites) are harder to check and often wronger; sample them deliberately.
- Duplicate inflation. One wrong macro copied into 300 tickets counts as 300 errors or one, depending on whether you deduplicated first. Run near-duplicate detection before sampling and report both views.
- Survivorship in exports. If the supplier filtered to "published" articles, archived content is missing and the outdated rate looks low, but so does coverage of older product versions.
How SourceX handles knowledge content requests
SourceX sources operational datasets, including documents and support histories, from US companies on request; categories are not inventory and a request does not guarantee a match. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Personal details such as names, emails and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked. If you are scoping knowledge content, describe the document types, versions and date fields your fact-check needs on the buyer request page, and see the knowledge base article datasets page and the knowledge base glossary entry for background.
Sourcing knowledge text you can verify
SourceX looks for US businesses that hold the knowledge content you describe, and every release is approved by the supplying company. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set per deal in a license. Describe the articles, documentation or Q&A data you need, including the version and review fields your accuracy audit depends on, at sourcex.si/buyers.
Sources
- arXiv (2603.14712), "Towards Next-Generation LLM Training: From the Data-Centric Perspective" (2026). https://arxiv.org/pdf/2603.14712
- arXiv (2506.07270), "Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs" (2025). https://arxiv.org/html/2506.07270v1
- Gekhman et al., arXiv / EMNLP 2024, "Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?" (2024). https://arxiv.org/abs/2405.05904v3
- Tang, Laban, Durrett (UT Austin, Salesforce AI Research), arXiv, "MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents" (2024). https://arxiv.org/html/2404.10774v2
- arXiv (2607.20465), "DataPrep-Bench: Benchmarking LLMs as Training Data Preparators" (2026). https://arxiv.org/pdf/2607.20465
- Northcutt, Athalye, Mueller, arXiv / NeurIPS 2021, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and ML, Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
- Pushkarna, Zaldivar, Kjartansson (Google Research), FAccT 2022, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Wettig et al., "QuRating: Selecting High-Quality Data for Training Language Models" (2024). https://arxiv.org/abs/2402.09739
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.