Data quality, coverage and contamination
What a Dataset Quality Report Should Contain: Measured Metrics, Methods and Known Defects
Quick answer
A dataset quality report is the measured counterpart to a dataset card: it states, for one specific delivery version, how many records arrived per stratum, how complete and valid each field is, the exact and near-duplicate rates, label audit results, residual personal-data findings and the defects the supplier already knows about. Every number should carry its method: sample size, sampling frame, auditor and tool. Require one with each delivery, not once per relationship.
By SourceX Editorial · Updated
How a quality report differs from a dataset card
A dataset card describes what the data is; a quality report proves how good one shipped version of it is. Datasheets for datasets ask about motivation, composition, collection process and recommended uses [5], and Data Cards summarize upstream sources, collection and annotation methods and intended use [6]. Those answers rarely change between deliveries, so they belong in documentation such as our guide to dataset cards for licensed enterprise data and the datasheet for datasets glossary entry.
A quality report changes every time the supplier re-extracts, re-labels or appends data. It is the document your ML data engineer uses to accept or reject a delivery and the one procurement attaches to the acceptance record. If the two are merged, the measured numbers tend to drift out of date while the descriptive text stays fixed, and nobody can tell which version the figures refer to.
Anchoring the report structure to ISO/IEC 5259
ISO/IEC 5259-2 is the closest published standard for structuring a dataset quality report, because it defines a data quality model, a set of quality measures and guidance on reporting data quality for analytics and ML [1][2]. It builds on the older ISO/IEC 25012 data quality model and ISO 8000 [1]. Part 1 supplies the shared terminology [3], and Part 4 covers process approaches for training and evaluation data, including labeling for supervised learning [4].
You do not need a supplier to certify against the series, and the standards are paywalled, so many vendors will not have read them. The practical ask is narrower: name each reported characteristic using 5259 vocabulary (such as completeness, accuracy, consistency and currentness), state which measure was computed and give the formula. That makes reports from different suppliers comparable and gives your reviewers a neutral reference when a supplier invents its own metric names. For the definitions behind each metric, see training data quality metrics.
The ten sections a supplier report should include
A complete report has ten sections, and the first two identify exactly what was measured. Without a version identifier and a record count that reconciles to the delivery manifest, every later number is unanchored. Compare the counts against the file-level inventory described in the sample manifest.
- Scope and version. Dataset name, delivery ID, extraction date range, source systems (for example Zendesk tickets, Salesforce Opportunity history, Jira issues), schema version and file hashes.
- Record counts by stratum. Counts per stratum you specified in the request: product line, region, year, ticket channel, document type. Report the expected count next to the delivered count.
- Field-level completeness. Null, empty-string and placeholder rates per field ("N/A", "unknown", "1900-01-01"), split by stratum where rates differ.
- Field-level validity. Share of values passing type, format, range and enumeration checks: ISO 8601 timestamps, currency codes, status values against the source system's picklist, foreign keys that resolve.
- Consistency checks. Cross-field rules, such as closed_at later than created_at, invoice totals equal to line-item sums, resolution codes present only on closed tickets.
- Duplicate rate. Exact-duplicate rate by content hash and near-duplicate rate with the method and threshold (MinHash Jaccard 0.8 on 5-gram shingles, for example). See near-duplicate detection with MinHash and LSH.
- Label audit results. For labeled fields: audited sample size, agreement statistic, estimated label error rate with a confidence interval, and confusion between the most-confused classes.
- PII residue test. What de-identification was applied, which detectors were run afterward, how many records were reviewed by hand, and how many residual identifiers were found per 10,000 records.
- Coverage and distribution notes. How the delivered distribution compares with the population it was drawn from, and where it is known to be thin.
- Known defects and limitations. A dated list of issues the supplier is aware of but has not fixed, with affected record counts.
Method notes: why every number needs its provenance
A quality figure without its method cannot be checked, compared or relied on in an acceptance decision. A completeness rate of 98% means something different when computed on all 2.1 million records than on a 500-record sample drawn only from the most recent month. Require a method block for each metric.
The minimum method block has six fields: population or sample size, sampling frame and selection procedure (simple random, stratified, systematic), who measured it (supplier staff, a third-party annotation vendor, an automated check), the tool and version (for example Great Expectations suites, Deequ checks, a SQL script with its hash), the date run, and the threshold that defines pass or fail. Ask for the check code itself where the supplier will share it; re-running it on your side is the cheapest verification there is.
Treat a report that gives only dataset-level averages and no field-level statistics as a red flag. Buyer guides commonly recommend using structured review and acceptance workflows to identify incomplete or unusable submissions before a delivery is finalized [11]. You cannot apply those thresholds if the supplier reports one blended score.
Label audits and the label error rate
The label audit section should report an estimated error rate with a confidence interval, not an assurance that labels were reviewed. Research on popular benchmark test sets found label errors widespread enough to change which model appeared to perform best [8], so even well-known data is not exempt. Ask how the audit sample was drawn, who adjudicated disagreements and whether auditors saw the original label (which biases them toward agreement).
Model-assisted methods such as confident learning flag likely mislabeled examples from predicted probabilities [9], and suppliers increasingly report counts of flagged items. A flag count is not an error rate: ask what fraction of flagged items a human confirmed as wrong. For outcome fields taken from operational systems (a "resolved" status, a won/lost flag), the audit should test whether the field reflects what actually happened, as covered in verifying outcome labels in operational records. For the audit procedure itself, see how to audit annotation quality.
PII residue testing after de-identification
The report should state what personal data remained after de-identification, measured on a sample, rather than only what method was applied. Detectors such as Microsoft Presidio or cloud DLP services miss identifiers embedded in free text, attachments, file names and email signatures, which are common in support and sales histories. Ask for the residue rate per identifier type, the sample size and whether reviewers checked free-text fields specifically.
For health data, the report should name the HIPAA pathway: Safe Harbor removes 18 listed identifiers, while Expert Determination relies on a qualified expert's documented analysis [10]. Neither method eliminates re-identification risk, so the report should never claim zero residue. For the risk side of this question, see re-identification risk assessment for licensed datasets.
Coverage limits and the known-defects section
The known-limitations section is where a supplier states how the delivered data departs from the population you will deploy against. Researchers have recommended that dataset documentation record the known limits of coverage and distribution match and add explicit representativity questions [7]. In operational data, typical gaps are a single region, a period before a product launch, one customer tier, or tickets from one channel because the others ran on a different system.
A useful known-defects list is specific and dated: "Tickets from 3 to 17 March missing priority field due to a migration; 4,212 records affected" is useful, while "some fields may be incomplete" is not. Compare each listed gap with your own deployment distribution using a coverage gap analysis. An empty known-defects section on a large operational extract is itself a reason to ask more questions.
Illustrative report skeleton for a support-ticket delivery
The skeleton below shows the level of detail to request; adapt the fields to your modality.
Illustrative example: invented to show structure; it does not describe an available dataset.
quality_report:
delivery_id: DLV-2026-07-B
schema_version: tickets_v3
extraction_window: 2023-01-01/2025-12-31
files: {count: 48, manifest_sha256: "<hash>"}
strata:
- {key: "channel=email", expected: 410000, delivered: 408912}
- {key: "channel=chat", expected: 220000, delivered: 219004}
fields:
- name: resolution_code
completeness: {rate: 0.962, n: 627916, method: "full scan, SQL check v1.4"}
validity: {rate: 0.991, rule: "value in source picklist (37 codes)"}
duplicates:
exact: {rate: 0.004, method: "SHA-256 of normalized body"}
near: {rate: 0.031, method: "MinHash, 128 perms, Jaccard >= 0.8, 5-gram"}
label_audit:
field: intent_label
sample: {n: 1200, frame: "stratified by channel and year"}
auditors: "2 independent reviewers, blind to original label"
error_rate: {estimate: 0.047, ci95: [0.036, 0.060]}
pii_residue:
method: "pattern + NER detection, then manual review"
sample: {n: 2000, free_text_reviewed: true}
residual_per_10k: {email: 1.5, phone: 0.5, person_name: 6.0}
known_defects:
- {id: KD-3, desc: "priority null for migration window", records: 4212}
coverage_notes: "No tickets from enterprise tier; chat channel starts 2024-02."
Turning the report into acceptance criteria
The report becomes enforceable only when its fields map to thresholds written into the order or license. Agree, before delivery, which metrics are pass/fail, which are informational and what happens on failure: re-delivery, replacement of defective records or a price adjustment. Pair the report with your own re-measurement on a sample, using an acceptance sampling plan for dataset deliveries.
Counsel should check that the license warranties reference the report rather than generic "industry standard quality" language; see data warranties in AI training licenses. Keep each signed-off report with the delivery ID so later model audits can trace which version trained which model. More measurement topics sit in our training data quality assessment hub.
How SourceX handles quality evidence
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked; no method is perfect. Buyers can describe the quality evidence they need as part of a request through the SourceX buyer page.
Request a dataset with documented quality evidence
SourceX sources operational datasets from US companies on request and delivers each under a license that defines the records, uses, term and delivery. Every dataset is rights-reviewed, and nothing is contracted until the supplying company agrees. Describe the data and the quality evidence you need at sourcex.si/buyers.
Frequently asked questions
Should the quality report come from the supplier or a third party?
Either can work if the method block is complete. A supplier-produced report is acceptable when you receive the check code or can re-run the checks on a sample; independent review matters most for label audits, where the supplier's annotators should not grade their own work.
Is one report enough for a recurring data feed?
No. Each delivery or refresh needs its own report tied to a delivery ID, because extraction logic, source-system schemas and labeling staff change over time. Track the key metrics across deliveries so drift is visible.
What if a supplier refuses to share field-level statistics?
Treat it as a material gap and ask why. Some suppliers worry that the statistics reveal business volumes; aggregated rates by field and stratum, without raw counts, usually address that concern while still letting you apply thresholds.
Sources
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Standards Council of Canada, "ISO/IEC 5259-2:2024 (catalogue listing)" (2024). https://scc-ccn.ca/standardsdb/standards/8187108
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-1:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 1: Overview, terminology, and examples" (2024). https://www.iso.org/standard/81088.html
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
- Gebru et al., "Datasheets for Datasets" (arXiv:1803.09010). https://arxiv.org/pdf/1803.09010
- Pushkarna, Zaldivar, Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Clemmensen, Kjærsgaard, "Data Representativity for Machine Learning and AI Systems" (2022). https://arxiv.org/pdf/2203.04706
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/pdf/2103.14749
- Northcutt, Jiang, Chuang, "Confident Learning: Estimating Uncertainty in Dataset Labels" (2019). https://arxiv.org/pdf/1911.00068
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- CloudPano, "Selecting the Best AI Training Data Provider: A Practical Buyer's Guide". https://www.cloudpano.com/blog/selecting-best-ai-training-data-provider-buyers-guide
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.