Data quality, coverage and contamination
Training Data Quality Metrics: How to Measure Accuracy, Completeness, Consistency and Timeliness
Quick answer
Training data quality metrics are ratios you can recompute: a numerator of records or values that pass a stated rule, divided by the records or values checked. Report completeness, validity, consistency, uniqueness, label accuracy, timeliness and representativeness separately, never as one blended score. Pull definitions from ISO/IEC 5259-2 [1] and set the pass thresholds yourself. When a metric comes from an audit sample rather than a full scan, report the sample size and a confidence interval with it.
By SourceX Editorial · Updated
This page is the formula reference for the data quality cluster. For what the finished report should contain, see what a dataset quality report should contain.
Which standards define the quality dimensions
ISO/IEC 5259-2:2024 provides the data quality model and measures for analytics and ML data, and it builds on ISO/IEC 25012 and ISO 8000 [1]. Published summaries of the 5259 series list characteristics including accuracy, completeness, consistency, timeliness and representativeness [2]. The standard is paywalled, so cite the clause you use for each measure in your internal spec and keep a licensed copy on file.
ISO 8000-8 adds a useful split between syntactic quality (does the value conform to the format), semantic quality (does it correctly describe the real thing) and pragmatic quality (is it fit for the intended use) [3]. Neither standard sets your business thresholds [3]. A 98% completeness floor can be right for a ticket resolution_code field and wrong for an optional customer_tier field. Set a threshold per measured object (field, record type or label) and write it down before the data arrives.
Core formulas for completeness, validity and consistency
Completeness, validity and consistency are deterministic checks you can run on every record, so report them as exact counts rather than estimates. Each needs a declared scope: which fields are required, which code sets apply and which cross-field rules hold.
- Completeness = populated required fields / (records × required fields). Count empty strings, whitespace-only values and sentinel values such as
N/A,0000-00-00or-1as missing, and list the sentinels you treated that way. - Record completeness = records with every required field populated / total records. This is often far lower than field completeness, and it is what a training pipeline that drops incomplete rows actually experiences.
- Validity = values passing schema, type, range and code-set rules / values checked. Examples: ISO 8601 timestamps parse, currency codes fall in ISO 4217,
priorityis in{P1,P2,P3,P4}. - Consistency = records passing all cross-field and cross-table rules / records checked. Typical rules:
closed_at >= opened_at,status = closedimplies non-nullresolution_code, line items sum to the invoice total, and everyticket_idin a message table exists in the ticket table. - Format conformance for file deliveries = lines that parse / total lines. For JSON Lines, the spec requires UTF-8 without a byte order mark and a valid JSON value on every line; a blank line is invalid [8].
For tabular deliveries, the full check list (schema drift, null patterns, referential integrity) lives in validation checks for structured dataset deliveries.
Duplicate and near-duplicate rates
Duplicate rate is the share of records that are redundant copies of another record, and you should report exact and near-duplicate rates as two numbers. Exact duplicate rate = (records − distinct normalized records) / records, after a declared normalization (lowercasing, whitespace collapse, stripping signatures and tracking IDs). Near-duplicate rate = records in a near-duplicate cluster, minus one keeper per cluster, / records, at a stated similarity threshold such as MinHash Jaccard of 0.8 on 5-gram shingles.
Always publish the threshold and shingle size, because the rate moves sharply with both. In business records, canned replies and email templates inflate near-duplicate rates without being errors; report them, then decide. The method details are in exact and near-duplicate detection with MinHash and LSH.
Label accuracy and agreement from an audit sample
Label accuracy is the share of audited examples whose delivered label matches an adjudicated reference label, and it is almost always an estimate from a sample. Formula: label accuracy = audited examples where delivered label = adjudicated label / audited examples. Draw the sample randomly (stratified by label if classes are imbalanced), have two reviewers label blind, adjudicate disagreements, and report inter-reviewer agreement (Cohen's kappa or Krippendorff's alpha) alongside accuracy so readers can judge the reference itself.
Model-assisted methods help you choose where to look. Confident learning estimates the joint distribution of noisy and true labels and flags examples likely to be mislabeled [6]. Use flagged counts as a triage signal, not as the accuracy figure, because flagged examples are not a random sample. Audit design is covered in how to audit annotation quality, and operational outcome fields used as labels have their own traps, covered in verifying ground truth in operational records.
Text-specific metrics for pre-training and fine-tuning data
Text corpora need metrics on top of field checks: format, duplication, diversity, fluency and factual accuracy are common metric-based checks used before LLM training [4]. Make each one computable:
- Diversity: distinct n-gram ratio (distinct 3-grams / total 3-grams), topic or intent entropy over a classifier's labels, and the share of tokens contributed by the top 1% of sources or authors.
- Fluency: share of documents with perplexity below a cutoff under a named reference model (lower perplexity means more fluent text), plus language-ID confidence for the target language.
- Length and truncation: token-length distribution (p5, p50, p95) and share of documents ending mid-sentence or cut at a fixed byte limit.
- Factual accuracy: share of sampled verifiable claims judged correct against a stated reference, reported with sample size. See factual accuracy checks for licensed text.
These rule-based metrics differ from model-based per-example scores. QuRating-style scoring rates each document on dimensions such as writing style, required expertise, facts and trivia, and educational value [5]. Per-example scores are useful for filtering and ranking, but they depend on the scoring model and prompt, so record both and never compare scores produced by different scorers. The side effects of filtering on such scores are discussed in quality filtering for pretraining-scale text. Recent work recommends combining metric checks, model scoring and human review of samples rather than relying on any one of them [4].
Operational-record metrics for case and workflow data
Operational records such as support tickets, claims and work orders need process metrics, because a record can pass every field check and still be useless as training data. Three ratios catch most problems:
- Outcome coverage = cases with a recorded terminal outcome (resolution, decision, disposition) / cases in scope.
- Step timestamp coverage = cases with a timestamp on every expected workflow step / cases in scope. Report per step too; a missing
escalated_atis common and changes what a routing model can learn. - Valid reason-code rate = cases whose reason or resolution code is in the current code set and is not a catch-all (
OTHER,MISC) / cases with a code.
Truncated threads and absent outcomes are covered in depth in checking completeness of case and ticket records.
Timeliness and representativeness
Timeliness measures how current the data is relative to the use it will serve, so define it against a reference date. Two useful forms: record age = reference date − event timestamp, reported as a distribution, and currency = records with an event date inside the required window / records. Also report the latency between event and extraction if the supplier's export lags source systems.
Representativeness compares the delivered distribution with a target distribution you specify. Compute it as a divergence (Jensen-Shannon or population stability index) or as per-stratum coverage ratios, for example delivered share of each product line, region or case type divided by target share. A dataset with perfect completeness and validity can still fail here, which is why ISO/IEC 5259 treats it as a separate characteristic [2].
Reporting audit-sample metrics with confidence intervals
Any metric computed on a sample must carry its sample size and an interval, or two datasets cannot be compared. For proportions, use a Wilson score or exact (Clopper-Pearson) interval rather than the normal approximation, which gives intervals that are too narrow on small samples; a 2025 ICML position paper makes this point for evaluation sets with fewer than a few hundred items [7]. How many records to check for a target margin is covered in sample sizes for estimating a dataset's error rate.
Worked arithmetic: 12 label errors in a random audit of 400 gives an observed error rate of 3.0%, and a 95% Wilson interval of roughly 1.7% to 5.2%. If your acceptance threshold is "error rate below 5%", this sample does not establish it, because the upper bound exceeds 5%. Either enlarge the sample or state the result as inconclusive.
A metric record to attach to every delivery
A per-delivery metric record makes results comparable across suppliers and refreshes. Keep one row per metric, with the rule and scope explicit enough that someone else can rerun it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| metric | scope | formula | numerator | denominator | method | value | 95% CI | threshold | pass |
|---|---|---|---|---|---|---|---|---|---|
| field_completeness | tickets: 6 required fields | populated / (records × fields) | 1,178,040 | 1,200,000 | full scan | 98.2% | n/a | ≥ 97% | yes |
| validity_priority | tickets.priority | values in {P1..P4} / values | 199,410 | 200,000 | full scan | 99.7% | n/a | ≥ 99.5% | yes |
| consistency_closed_has_code | status=closed | closed with resolution_code / closed | 171,300 | 180,000 | full scan | 95.2% | n/a | ≥ 98% | no |
| near_dup_rate | message bodies | MinHash J ≥ 0.8, 5-gram shingles | 14,600 | 200,000 | full scan | 7.3% | n/a | report only | n/a |
| label_accuracy | intent label | matches adjudicated / audited | 388 | 400 | random audit, 2 blind reviewers | 97.0% | 94.8–98.3% | ≥ 95% (lower bound) | no |
| outcome_coverage | tickets in scope | with terminal outcome / in scope | 189,000 | 200,000 | full scan | 94.5% | n/a | ≥ 90% | yes |
Note the label-accuracy row: the point estimate passes, but the threshold was written against the lower bound, so it fails. Writing thresholds that way before delivery prevents arguments after it. A scoring template is available in the dataset quality scorecard, and supplier-level questions are in evaluating data supplier quality.
Common measurement failures
Most disputed quality numbers come from undeclared scope, not from bad arithmetic. Watch for these:
- Denominators that silently exclude rows dropped during parsing, which inflates every downstream ratio.
- Completeness computed over all fields, which hides a 60%-populated critical field among many full optional ones.
- Deduplication run after train/eval splitting, which leaves near-duplicates across splits and contaminates evaluation.
- Label accuracy reported from supplier-selected "gold" items rather than a random sample.
- Per-example quality scores compared across different scoring models or prompt versions.
Getting licensed data you can measure
Metrics only help if the supplier can tell you what the fields mean and how the data was prepared. SourceX sources operational datasets from US companies on request rather than from stock, and prepares diligence materials covering source, rights, preparation and allowed use for each dataset. You can bring your metric definitions and thresholds to a buyer request on SourceX and describe the data you need.
Bring your quality bar to a training data request
SourceX looks for US businesses that hold the data you describe, and each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Describe the dataset and quality bar you need.
Sources
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Management Solutions, "ISO/IEC 5259 Artificial intelligence: Data quality for analytics and machine learning (ML)". https://www.managementsolutions.com/en/node/4456
- ISO via ANSI Webstore, "ISO 8000-8:2015 Data quality - Part 8: Information and data quality: Concepts and measuring (preview)" (2015). https://webstore.ansi.org/preview-pages/ISO/preview_ISO+8000-8-2015.pdf
- arXiv, "Towards Next-Generation LLM Training: From the Data-Centric Perspective" (2026). https://arxiv.org/pdf/2603.14712
- arXiv, "DataPrep-Bench: Benchmarking LLMs as Training Data Preparators" (2026). https://arxiv.org/pdf/2607.20465
- arXiv, "Confident Learning: Estimating Uncertainty in Dataset Labels" (2019). https://arxiv.org/pdf/1911.00068
- ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- jsonlines.org, "JSON Lines". https://jsonlines.org/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.