Data quality, coverage and contamination
How Many Records to Check: Sample Sizes for Estimating a Dataset's Error Rate
Quick answer
To estimate a dataset's error rate, draw a simple random sample, count errors, and report a binomial confidence interval rather than a point rate. About 400 records gives roughly ±2 percentage points at 95% confidence when the true rate is near 5%; about 1,800 gives ±1 point. If you find zero errors, the 95% upper bound is close to 3/n, so 300 clean records bounds the rate near 1%. Stratify by source, period and class when those segments matter.
By SourceX Editorial · Updated
What the audit sample actually estimates
An audit sample estimates one proportion: the share of records in a defined population that fail a written error definition. Everything downstream depends on that definition, so fix it before drawing a single record. "Error" might mean a wrong class label, a resolution code that contradicts the ticket thread, a mis-transcribed amount in an invoice extraction, or a missing outcome field. Mixing these into one count produces a number nobody can act on.
Write the population, unit and defect taxonomy down. The population is the exact file set (for example, every row in the delivered Parquet partitions for 2023-2025), the unit is a record or a labeled span, and each defect type gets an ID, a definition and a severity. Our guide on training data quality metrics covers how error rate sits beside completeness, consistency and timeliness.
The reviewer is part of the measurement. If two reviewers disagree on 8% of sampled records, your estimate carries that disagreement too, so double-review a slice and adjudicate, and track agreement with the methods in inter-annotator agreement for dataset buyers.
Choosing a confidence interval for the observed error rate
Report a Wilson or Jeffreys interval, not the textbook Wald interval, because the Wald interval undercovers badly for small samples and rates near zero. Brown, Cai and DasGupta showed the Wald interval's coverage is erratic even at sample sizes many practitioners treat as safe, and recommended Wilson or Jeffreys for small n and Agresti-Coull for larger n [6]. The same problem shows up in model evaluation, where intervals that lean on the central limit theorem are too narrow on samples of a few hundred items [5].
The Wilson score interval for x errors in n records, with p̂ = x/n and z = 1.96 for 95%, is:
(p̂ + z²/2n ± z·sqrt(p̂(1−p̂)/n + z²/4n²)) / (1 + z²/n)
Clopper-Pearson, the "exact" interval built from beta quantiles, always delivers at least nominal coverage but is conservative and wider. Use it when a contract or reviewer demands assured coverage; otherwise Wilson is the practical default. In Python, statsmodels.stats.proportion.proportion_confint(x, n, method="wilson") returns it, and method="beta" gives Clopper-Pearson.
Illustrative example: invented to show structure; it does not describe an available dataset.
Worked example: a reviewer audits 400 randomly drawn support tickets and finds 12 with a resolution code that contradicts the thread. The point estimate is 3.0%. The Wald interval is about 1.3% to 4.7%; the Wilson interval is about 1.7% to 5.2%. Note the asymmetry: Wilson correctly allows more room above the estimate than below it, which is the side that matters when you are deciding whether the labels are usable.
Sample size versus margin of error
The sample size needed for a target margin E at 95% confidence is approximately n = 1.96² · p(1−p) / E², where p is your best prior guess of the error rate. When you have no guess, p = 0.5 is the worst case and gives 385 for ±5 points. For realistic label error rates, the requirement depends heavily on how tight you need the bound to be.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Expected error rate | n = 100 | n = 400 | n = 1,000 | n = 2,000 | n for ±1 point |
|---|---|---|---|---|---|
| 2% | ±2.7 pts | ±1.4 pts | ±0.9 pts | ±0.6 pts | ~753 |
| 5% | ±4.3 pts | ±2.1 pts | ±1.4 pts | ±1.0 pts | ~1,825 |
| 10% | ±5.9 pts | ±2.9 pts | ±1.9 pts | ±1.3 pts | ~3,458 |
Half-widths use the normal approximation to show scale; compute the final reported interval with Wilson. Three practical consequences follow. Margins shrink with the square root of n, so halving a margin costs four times the review effort. The absolute size of the dataset barely matters once it is much larger than the sample. And when the population is small, apply the finite population correction n′ = n / (1 + (n − 1)/N): a 385-record plan against a 2,000-record delivery drops to about 323.
For context on why these numbers matter, Northcutt and colleagues estimated an average label error rate of at least 3.3% across the test sets of 10 widely used benchmark datasets [3]. A 100-record spot check cannot distinguish 3% from 7%, which is the difference between a usable eval set and one that distorts model rankings.
When you find zero errors: the rule of three
If n randomly sampled records contain zero errors, the one-sided 95% upper confidence bound on the true error rate is approximately 3/n. The rule comes from Hanley and Lippman-Hand's analysis of zero numerators and follows from solving (1 − p)ⁿ = 0.05, which gives p = 1 − 0.05^(1/n) ≈ 3/n. It is accurate once n is above roughly 30 [7].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Target upper bound (95%) | Exact zero-error n | Rule of three n | 90% confidence (≈2.3/n) | 99% confidence (≈4.6/n) |
|---|---|---|---|---|
| 5% | 59 | 60 | 46 | 92 |
| 1% | 299 | 300 | 230 | 460 |
| 0.5% | 598 | 600 | 460 | 920 |
| 0.1% | 2,995 | 3,000 | 2,300 | 4,600 |
The rule answers "how bad could it be?", not "how good is it?". Zero errors in 300 records is consistent with a true rate anywhere from 0% to about 1%. It also covers only the defect types your reviewers were checking, so a clean sample on label correctness says nothing about PII leakage or duplicates unless those were in the checklist.
Stratified audit sampling and per-stratum intervals
Stratify the audit when error rates plausibly differ by segment, and report an interval per stratum as well as a weighted overall rate. In operational data, the usual strata are source system (two CRMs merged after an acquisition), period (before and after a taxonomy change), class or label value, annotator or vendor team, and language. A single pooled 2% can hide one class at 15%.
Allocate the sample deliberately. Proportional allocation mirrors the population and usually estimates the overall rate at least as precisely as a simple random sample of the same size; a fixed minimum per stratum (say 100 to 300 records) lets you bound each stratum's rate, which is what matters for rare classes. The overall estimate is Σ W_h · p̂_h, where W_h is each stratum's population share, with variance Σ W_h² · p̂_h(1 − p̂_h)/n_h.
Watch for clustering. Records from the same ticket thread, annotator session or batch file tend to share errors, so sampling 400 records from 20 batches carries less information than 400 independent records. Inflate n by the design effect 1 + (m − 1)ρ, where m is records per cluster and ρ is the within-cluster correlation, or sample clusters first and records within them. Related checks on rare segments are in long-tail and edge-case coverage.
Random samples versus model-flagged samples
Use a random sample to estimate the rate and a model-flagged sample to find and fix errors; never report the flagged hit rate as the dataset's error rate. Confident learning, implemented in the open-source cleanlab library, ranks examples by how likely their given label is wrong based on estimated class-conditional noise [4]. In the Northcutt test-set study, algorithmically flagged candidates were then validated by human reviewers [3].
That workflow is efficient for cleaning but biased for estimation, because flagged records are enriched for errors by design. A sound audit runs both: a simple random or stratified sample for the headline interval, and a ranked review queue for remediation. Document which records came from which draw, and store the random seed and the sampling frame (for example, a SHA-256 hash of the sorted record IDs) so a counterparty can reproduce the draw.
Illustrative example: invented to show structure; it does not describe an available dataset.
Audit plan checklist:
- Population: file manifest, row count, record ID field, snapshot date.
- Defect taxonomy: ID, definition, severity, examples of pass and fail.
- Design: simple random or stratified; strata, weights and per-stratum n; seed.
- Target: margin or zero-error bound, confidence level, interval method (Wilson or Clopper-Pearson).
- Review protocol: blind review, double-review share, adjudication rule.
- Output: x, n and interval per stratum and overall; defect examples by ID.
Estimation versus accept/reject decisions
Estimating an error rate and deciding whether to accept a delivery are different questions, and the second uses a sampling plan with an operating characteristic curve. A single sampling plan is a pair (n, c): inspect n units and reject the lot if more than c are defective [1]. The OC curve then shows the probability of accepting the lot at each true defect rate, which is how buyers and suppliers agree on the risk each side carries [2].
The two connect directly. A zero-acceptance plan with n = 300 and c = 0 has roughly a 5% chance of accepting a delivery whose true error rate is 1%, which is the rule of three read from the other side. If your contract defines acceptance thresholds, build the plan from the OC curve as described in acceptance sampling for dataset deliveries, and set the threshold itself using acceptable label noise by use. For how to choose which records represent a business dataset in the first place, see selecting a representative sample.
Where sourced operational data fits
Sample-size math applies the same way to records you license from another company. SourceX sources operational datasets such as support and sales histories, engineering records, documents, and finance and legal workflows from US companies on request, and manages the licensing process; data is not held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with diligence materials prepared per dataset. Buyers who already know the error tolerance and audit design they need can describe the dataset on the SourceX buyers page.
More quality methods, from annotation audits to overlap checks, are collected in the data quality, coverage and contamination hub and the broader AI data buyer guides. For label-level review procedures, see how to audit annotation quality.
Sourcing data you will audit by sample
If you are planning an error-rate audit for operational training or evaluation data, SourceX can look for US businesses that hold the records you describe and manage the license through Find, Assess, Agree, Transact and Manage. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Start by describing your data requirements to SourceX.
Sources
- NIST, "NIST/SEMATECH e-Handbook of Statistical Methods: Lot acceptance sampling, single sampling plans". https://itl.nist.gov/div898/handbook/pmc/section2/pmc23.htm
- Open Exam Prep (CQE study guide), "4.5 Acceptance Sampling Plans & Standards". https://open-exam-prep.com/study-guides/cqe/product-and-process-control/acceptance-sampling-plans
- arXiv / NeurIPS 2021 Datasets and Benchmarks (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (Northcutt, Jiang, Chuang), "Confident Learning: Estimating Uncertainty in Dataset Labels". https://arxiv.org/pdf/1911.00068
- ICML 2025, "Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints" (2025). https://icml.cc/virtual/2025/poster/40132
- Brown, Cai and DasGupta (Statistical Science), "Interval Estimation for a Binomial Proportion" (2001). https://projecteuclid.org/journals/statistical-science/volume-16/issue-2/Interval-Estimation-for-a-Binomial-Proportion/10.1214/ss/1009213286.full
- Hanley and Lippman-Hand (JAMA), "If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators" (1983). https://jamanetwork.com/journals/jama/article-abstract/387532
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.