Skip to content

Data quality, coverage and contamination

Training Data Quality Assessment: How AI Teams Judge Dataset Quality, Coverage and Contamination

Quick answer

Training data quality assessment is the set of measurements a buyer runs before a licensed dataset enters training or evaluation. It covers five families of checks: intrinsic record quality, label quality, coverage of the deployment distribution, duplication and benchmark contamination, and residual personal data or secrets. ISO/IEC 5259 supplies the shared terminology and measures [1] [2]. In practice, each family is measured on a random sample of the actual delivery, against thresholds written down before the sample arrives.

By SourceX Editorial · Updated

The five check families and what each one catches

A licensed dataset can fail in five independent ways, and passing one family says nothing about the others. Clean schemas can carry wrong labels, and accurate labels can sit in records that duplicate a public benchmark.

Check familyQuestion it answersWhat to measureFailure it preventsDeep dive
Intrinsic record qualityAre records correct, complete and internally consistent?Schema conformance, missingness by field, out-of-range values, referential integrity, timestamp orderTraining on truncated threads, null outcomes or mixed unitsAccuracy, completeness and timeliness metrics
Label qualityDo labels and outcome fields mean what the supplier says?Error rate on an adjudicated audit sample, inter-annotator agreement, per-class confusionLearning annotator noise or auto-set status fieldsAnnotation quality audit
Coverage and representativenessDoes the data span the inputs the model will see?Records per deployment slice, long-tail frequency, date gaps, subgroup countsGood aggregate scores that hide failing slicesCoverage gap analysis
Duplication and contaminationIs the data new, and is it free of your test sets?Exact and near-duplicate rates, overlap with owned data, n-gram and semantic benchmark overlapMemorization, paying twice, inflated eval scoresNear-duplicate detection with MinHash and LSH
Safety and privacy residueWhat personal data, secrets or toxic text survived preparation?Personally identifiable information (PII) hits by type, credential and API-key matches, toxicity scores, language-ID mismatchesMemorized personal data and leaked credentialsPII scanning before fine-tuning

The grouping is a working structure for buyers, not a standard; ISO/IEC 5259 frames the same ground as quality characteristics.

ISO/IEC 5259, Article 10 and other dataset quality standards

ISO/IEC 5259 is the reference standard series for data quality in analytics and machine learning, and its terms are the safest basis for quality thresholds in a data request or license annex. Part 1 (2024) gives the overview and terminology, Parts 2 to 4 (2024) cover measures, management requirements and a process framework [1], Part 5 (2025) adds a governance framework [22], and Technical Report 5259-6 on visualization followed in 2026 [23]. Part 2 defines a data quality model, measurable characteristics and guidance on reporting, and builds on ISO/IEC 25012 and ISO 8000 [2].

A published summary of the series lists characteristics such as accuracy, completeness, consistency, timeliness and representativeness [3]. The standards are paid, so quote the clause your team adopts, not a summary.

For high-risk AI systems in the EU, Article 10(3) of the AI Act requires training, validation and testing data sets to be relevant, sufficiently representative and, to the best extent possible, free of errors and complete in view of the intended purpose; Article 10(2) adds examination of possible biases and identification of data gaps [4]. As of October 2026, the AI Act has been amended by Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026. Secondary sources report that the amendment moved high-risk application dates to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems [5]. The amending regulation also amends Article 10 [5]. The AI training data compliance hub covers who these obligations reach.

Documentation expectations are rising. Datasheets for Datasets proposed that every dataset document its motivation, composition, collection process and recommended uses [6]. NeurIPS 2026 requires Croissant-RAI-based responsible-AI metadata, including limitations, potential biases and intended use, for its Evaluations and Datasets Track [7]. Ask suppliers for measured values behind a datasheet for datasets, not prose; what a dataset quality report should contain lists the fields.

Training data quality vs quantity: what the evidence supports

Published results show that a smaller, cleaner dataset often matches or beats a larger noisy one, though the effect depends on the training stage and how quality was defined. For buyers, record counts are a weak proxy for value; the useful number is how many records survive your own filters.

  • Instruction tuning. LIMA fine-tuned a 65B-parameter LLaMA model on 1,000 curated prompt-response pairs, and its authors concluded that most knowledge comes from pretraining while a small high-quality set teaches output format [8]. AlpaGasus had ChatGPT grade Alpaca's 52,000 examples, trained on about 9,000 high scorers, and outperformed the original Alpaca in GPT-4 evaluations [9].
  • Tool use. Models trained on smaller sets of validated synthetic tool-use data outperformed models trained on larger unvalidated sets [10].
  • Deduplication. Lee et al. found a single 61-word sentence repeated over 60,000 times in C4. Models trained on deduplicated data emitted memorized text about ten times less often and reached the same or better accuracy in fewer training steps [11].
  • Filtering side effects. Across 28 pretrained 1.5B-parameter models, quality filtering raised downstream performance and toxic generation even though it removed more than 10% of the data, while toxicity filtering reduced toxic generation at some cost to generalization. The effects were not predictable from domain characteristics [12].

LIMA's authors note that curation is labor-intensive and that LIMA is less robust than product-grade systems [8]. "Less is more" is not a reason to buy less. It is a reason to ask for the expected yield after deduplication, filtering and label audit, and to run an ablation first; see estimating what a dataset adds before you buy it.

Data quality for LLM training by stage: which checks come first

The same dataset needs different checks depending on whether it feeds pre-training, supervised fine-tuning, preference tuning, evaluation or retrieval. How procurement differs by training stage covers the commercial side.

UseChecks to run firstThreshold to state in the request
Pre-training or continued pre-trainingNear-duplicate rate, overlap with public web crawls, language ID, toxicity, PII residueMaximum near-duplicate share at a named similarity method and cutoff
Supervised fine-tuning (SFT)Response correctness, format conformance, instruction diversity, share of templated repliesMaximum error rate on an expert-audited sample
Preference tuning (RLHF or direct preference optimization)Rater agreement on overlapping pairs, tie handling, rater calibration, position biasMinimum agreement statistic and the share of pairs double-rated
Evaluation setsGold-label accuracy, contamination against training corpora, leak-free splitsNo verbatim benchmark overlap; adjudicated gold labels
Retrieval-augmented generation (RAG) corporaDuplicate and superseded document versions, conflicting documents, freshnessVersion and effective-date fields on every document

Deep dives: preference data noise and agreement, instruction-tuning data filtering, pretraining-scale text filtering and knowledge-corpus quality for RAG. Eval set design belongs to the LLM evaluation datasets hub. If a model scores your records, validate it first with the LLM-judge guide.

How to evaluate a dataset before buying it: a nine-step checklist

Evaluate a dataset before buying it by testing a random sample drawn from the exact segment you will license, with the same pipeline and thresholds you will apply at delivery. One 2026 data-centric paper describes an assessment stack that combines metric checks (format, duplication, diversity, fluency, factual accuracy), model-based scoring and human review of samples [13].

  1. Fix thresholds first. Write a pass line for each check family before the sample arrives, so the sample cannot set its own bar.
  2. Get a random sample, not a showcase. Ask how it was drawn, from which source systems and over which date range; checking a sample against the full dataset explains the tests.
  3. Validate structure. Check schema, types, missingness by field, value ranges, foreign keys and timestamp order.
  4. Deduplicate inward and outward. Measure duplicates within the sample, against data you already own and against public pretraining corpora.
  5. Scan for contamination against every public benchmark and private eval set you report on.
  6. Audit labels on an adjudicated subsample sized for the error rate you need to detect; see sample sizes for estimating error rates.
  7. Scan for residue: PII by type, credentials and secrets in tickets and logs, and model-generated text in purchased human data.
  8. Map coverage against your deployment taxonomy and record empty or thin slices.
  9. Write a quality report the supplier can see, so any dispute is about numbers.

Illustrative example: invented to show structure; it does not describe an available dataset.

quality_report:
  dataset: support_ticket_threads_sample      # supplier sample, not the full delivery
  sample: {method: stratified_random, strata: [product_line, quarter], records: 2000}
  intrinsic:
    schema_conformance: 0.997
    missing_resolution_code: 0.043            # threshold <= 0.05
    truncated_threads: 0.012                  # threshold <= 0.02
  labels:
    audited_records: 400
    adjudicated_error_rate: 0.031             # threshold <= 0.05
    krippendorff_alpha_resolution_code: 0.78
  coverage:
    taxonomy_slices: 42
    slices_under_50_records: 6
  duplication_contamination:
    near_duplicate_rate_minhash: 0.064        # method and cutoff named in the request
    overlap_with_owned_data: 0.009
    benchmark_13gram_hits: 0
  residue:
    pii_hits: {email: 3, phone: 1, account_number: 0}
    credential_pattern_hits: 2                # threshold = 0
  verdict: fail
  open_issues:
    - "Auto-closed tickets carry resolution_code=resolved with no agent reply"
    - "Credential hits must be scrubbed and the sample re-sent"

Measurement is separate from contract acceptance. Thresholds become enforceable only when they appear in acceptance criteria for licensed training data with an inspection plan. ANSI/ASQ Z1.4 attribute sampling plans, with switching between normal, tightened and reduced inspection [14], were built for manufactured lots; acceptance sampling for dataset deliveries adapts them to records and labels. When SourceX sources a dataset, it checks the data and the supplier's licensing permissions and prepares diligence materials on source, rights, preparation and allowed use for the buyer's review; your quality report sits beside them, and you can state the thresholds a dataset must meet in your data request.

Contamination and overlap: where standard checks fall short

N-gram overlap is the predominant contamination check and catches verbatim leakage cheaply, but it misses paraphrased, translated and synthetic copies of test items. A survey of benchmark contamination reports thresholds that differ across labs, from 13-gram matches for GPT-3 to other approaches in later work [15]. LMSYS researchers showed the gap: a 13B model trained on rephrased MMLU test items scored 85.9 on MMLU while n-gram overlap failed to flag it [16].

Contamination can also retire a benchmark. In a February 2026 post, OpenAI said SWE-bench Verified had become increasingly contaminated, stopped reporting it and recommended SWE-bench Pro in the interim [17].

Lee et al. also found train-test overlap affecting more than 4% of validation examples in standard benchmarks [11]. For licensed data, test overlap with data you already own, which means paying twice, and with public web crawls, which erodes the novelty you are paying for. Decontaminating a training set against benchmarks covers thresholds and reports; the guide to contamination checks for licensed evaluation data covers eval-set hygiene, and benchmark contamination has the definition.

Label and outcome-field quality in operational records

Labels in licensed data are rarely as clean as their documentation implies, so audit them on a sample instead of accepting a supplier-reported accuracy figure. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets, including at least 6% of the ImageNet validation set, and showed that such errors can reorder model rankings [18].

Agreement statistics measure consistency, not correctness. In Krippendorff's alpha, 1 indicates perfect reliability and 0 indicates agreement no better than chance [19], but high agreement can still mean annotators share a systematic error. That is why audits use golden datasets and adjudication alongside inter-annotator agreement metrics.

Operational business records add failure modes that research datasets rarely show. Support tickets carry macros and canned replies that inflate apparent diversity. A "resolved" status may be set by an auto-close rule rather than by an outcome. Email threads arrive truncated at export limits, and CRM migrations change what a field means partway through the history.

Each needs its own check: templates and boilerplate in business records, completeness of case and ticket records and outcome fields as ground truth.

Mistakes that let a weak dataset through

Weak datasets get through when checks run on the wrong sample, at the wrong level of aggregation, or only once.

  • Trusting the supplier's accuracy number. Re-measure on your own adjudicated sample.
  • Treating "representative" as one property. A dataset can mirror a population or cover the input space. Coverage is more robust to distribution shift and narrows accuracy gaps between subgroups, while population mirroring serves population-level inference [20]. Say which you need; assessing representativeness shows how to test each.
  • Reporting only aggregates. A strong overall error rate can hide one product line or customer segment where most labels are wrong. Report per slice, and check long-tail and edge-case coverage directly.
  • Deduplicating only exact matches. Templated records differing only by a ticket number or name pass exact hashing.
  • Assuming de-identified means clean. Carlini et al. extracted hundreds of verbatim training sequences from GPT-2, including contact details [21]. SourceX removes or replaces names, emails, phone numbers and account numbers before delivery and records the method used, but no de-identification method is perfect, so run your own scan. The de-identified data hub covers methods.
  • Measuring once. Re-run the checks on every delivery and refresh; supplier systems and export jobs change between them.

To compare suppliers rather than measure a delivery, use the guide to evaluating data supplier quality. The AI data hub maps the other buyer topics, including data provenance.

Need licensed data that has to pass these checks?

Describe the data you need and the quality thresholds it must meet. SourceX looks for US businesses that hold it, checks the data and the supplier's licensing permissions, and manages a license that defines which records are included and what they can be used for; a request does not guarantee a matching dataset. Describe your dataset and quality requirements.

Guides in this section

Sources

  1. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-1:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 1: Overview, terminology, and examples" (2024). https://www.iso.org/standard/81088.html
  2. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  3. Management Solutions, "ISO/IEC 5259 Artificial intelligence: Data quality for analytics and machine learning (ML)". https://www.managementsolutions.com/en/node/4456
  4. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  5. European Parliament and Council of the European Union (Official Journal, via EUR-Lex), "Regulation (EU) 2026/1744 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  6. Gebru et al. (arXiv; Communications of the ACM), "Datasheets for Datasets" (2018; CACM 2021). https://arxiv.org/pdf/1803.09010
  7. NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  8. Zhou et al. (arXiv:2305.11206; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  9. Chen et al. (arXiv:2307.08701), "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
  10. arXiv:2409.16341 (EMNLP 2024), "Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs" (2024). https://arxiv.org/abs/2409.16341v1
  11. Lee et al. (arXiv:2107.06499; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  12. Longpre et al. (arXiv:2305.13169; NAACL 2024), "A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity" (2023). https://arxiv.org/pdf/2305.13169
  13. arXiv:2603.14712, "Towards Next-Generation LLM Training: From the Data-Centric Perspective" (2026). https://arxiv.org/pdf/2603.14712
  14. ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables" (2003, reaffirmed 2018). https://asq.org/quality-press/display-item?item=T1164
  15. arXiv:2406.04244, "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
  16. LMSYS Org, "Catch me if you can! How to beat GPT-4 with a 13B model (LLM Decontaminator, 2023-11-14)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  17. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  18. Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  19. Klaus Krippendorff, Annenberg School for Communication, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  20. arXiv:2203.04706, "Data Representativity for Machine Learning and AI Systems" (2022). https://arxiv.org/pdf/2203.04706
  21. Carlini et al. (USENIX Security 2021), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  22. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-5:2025 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 5: Data quality governance framework" (2025). https://www.iso.org/standard/5259-5
  23. ISO/IEC JTC 1/SC 42, "ISO/IEC TR 5259-6:2026 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 6: Visualization framework for data quality" (2026). https://www.iso.org/standard/5259-6

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data