Data quality, coverage and contamination
How Much Label Noise Is Acceptable? Setting Error Tolerances by Use
Quick answer
There is no single acceptable label error rate. Tolerance depends on what the labels do: large pretraining or weakly supervised corpora absorb a few percent of random noise, supervised fine-tuning sets need far less, and reward-model, preference and evaluation data need the tightest bounds because every wrong label directly moves the optimization target or the score you report. Set a separate, lower cap for systematic errors, and write both into the acceptance spec before delivery.
By SourceX Editorial · Updated
Why one error rate cannot fit every use
The cost of a wrong label scales with how much leverage that label has over the final model, so tolerance must be set per use, not per dataset. In a corpus of millions of records, a randomly flipped label is one weak gradient signal among many; in a 1,000-example SFT set it is 0.1% of everything the model learns about the format; in an evaluation set it is a direct error in the number your team ships decisions on.
Measured evidence supports treating noise as a real cost rather than a nuisance. The confident learning work shows label errors can be estimated from model predictions, and that pruning them before training moderately increases accuracy [1]. The same group audited ten widely used benchmark test sets and estimated an average label error rate of at least 3.3%, including at least 6% of the ImageNet validation set [2]. Those benchmarks were curated by experienced teams, which is a useful prior for any operational dataset you license.
Volume does not fully compensate. Generalization error improves with more data along a power law [11], but that curve assumes the labels mean the same thing throughout. Extra records with a consistent bias shift the target rather than average out.
Random noise versus systematic error
Systematic errors matter more than random errors at the same headline rate, so measure and cap them separately. Random noise (a tired annotator occasionally clicking the wrong class) spreads across classes and tends to cost some accuracy; systematic noise (one queue that always codes "billing dispute" as "refund request", or a CRM field that defaults to "closed-won" when left blank) teaches the model a wrong rule.
Aggregate accuracy hides the damage. Slice-discovery research shows models that score well overall can fail consistently on coherent subsets of the data, and those subsets are usually unlabeled [6]. A 2% overall error rate concentrated in one product line, region or time window can mean a 30% error rate where it matters most.
Common sources of systematic error in operational records include:
- Default values: dropdowns pre-filled to the first option, or status fields auto-set by workflow rules rather than humans.
- Taxonomy drift: a category split or renamed mid-period, so "Tier 2" means different things before and after a migration date.
- Proxy outcomes: "ticket closed" used as "issue resolved", or "invoice paid" used as "dispute settled", when the field records a process step, not the outcome. See verifying ground truth in operational records.
- Single-annotator queues: one labeler with a consistent misreading of the guideline, invisible unless you break agreement down by annotator ID.
Tolerances by training use
Tolerance should tighten as the dataset gets smaller and as each label gains more direct influence over the objective. The ranges below are starting points for negotiation and internal sign-off, not published norms; adjust them after a pilot fine-tune or ablation on your own task.
Pretraining and continued pretraining. Labels often matter little here because the objective is next-token prediction over text. What matters is metadata that drives filtering or mixing (language, domain, document type, date). A few percent random error in those tags is usually tolerable; a mislabeled language or domain tag that pulls a whole slice into the wrong mixture bucket is not.
Supervised fine-tuning. Small SFT sets are sensitive. LIMA fine-tuned a 65B LLaMA model on only 1,000 carefully curated prompt-response pairs, and the authors stress that the curation was the work [7]. A response that is wrong, off-policy or written in the wrong format is learned as a target. Examples that introduce facts the model does not already know are learned slowly during fine-tuning and are associated with more hallucination [10]; a factually wrong target is worse still, so "factually incorrect target" deserves its own defect class. For classification fine-tunes, see labeled outcome data for classification fine-tuning.
Reward modeling and preference optimization. Preference data is where noise compounds. In the InstructGPT pipeline, human rankings train a reward model that then steers reinforcement learning [8]; in DPO the preference pairs act on the policy directly with no separate reward model to smooth them [9]. Recent work reports that inconsistent human preference labels can distort reward models and degrade the aligned model [3]. Ambiguous pairs, where both responses are acceptable, are not errors; flipped pairs are. Separate the two using the methods in preference data quality and agreement.
Evaluation and test sets. Evaluation data has the lowest tolerance because errors bias the metric itself. Test-set label errors were large enough that, in the benchmark audit, corrected labels could change which model ranked higher [2]. If two candidate models differ by 1.5 points and your gold labels are 3% wrong, the comparison is not decided by the data. Gold-label adjudication for delivered eval sets is covered in accepting a delivered eval set.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Use of the data | Random error cap (starting point) | Systematic error cap | Critical defect (reject at any rate in sample) | How to confirm |
|---|---|---|---|---|
| Pretraining / continued pretraining (metadata tags) | 5% on filter-driving tags | 1% per slice | Wrong license or source tag | Stratified sample by tag value |
| SFT demonstrations | 2% | 0.5% per slice | Factually wrong target; PII left in target text | Expert review of 200 to 400 pairs |
| Classification fine-tuning (outcome labels) | 3% | 1% per class or slice | Label leaks the outcome via another field | Per-class confusion review |
| Reward model / DPO preference pairs | 2% flipped pairs | 0.5% per prompt category | Pair where chosen response is unsafe | Re-annotation by 3 raters |
| Evaluation / test set | 0.5% to 1% | 0% tolerated once found | Item also present in training data | Full adjudication of disputed items |
Framing tolerance as an acceptance quality limit
Treat the tolerance like an acceptance quality limit: a worst acceptable average, not a quality target. In manufacturing inspection, an AQL is the poorest process average that will still be accepted most of the time, and suppliers are expected to run well below it [4]. Writing "2% label error" in a spec without that framing invites deliveries that sit exactly at 2%.
Attribute sampling standards such as ASQ/ANSI Z1.4 add two ideas that carry over to recurring data deliveries: separate classes for critical, major and minor defects, and switching rules that move from normal to tightened inspection after repeated lot rejections and, under stated conditions, to reduced inspection after a sustained run of accepted lots [5]. Applied to labels, a critical defect is anything that breaks the use (an eval item duplicated into training, PII in a target), a major defect is a wrong label, and a minor defect is a formatting or casing issue a script can fix.
The mechanics of choosing sample sizes and acceptance numbers are on separate pages: how many records to check and adapting AQL plans to dataset deliveries. The decision on this page comes first: what rate, per defect class and per use, you are willing to sign off.
Writing tolerances into a dataset acceptance spec
A tolerance is only enforceable if the spec defines what counts as an error, who decides, and how the rate is estimated. Ambiguity in any of those three turns every borderline label into a dispute after delivery.
Illustrative example: invented to show structure; it does not describe an available dataset.
acceptance_spec:
dataset: support_ticket_resolution_labels
intended_use: classification_fine_tuning
label_field: resolution_code # 14-class taxonomy, v3 (2025-03 onward)
error_definition:
major: "resolution_code disagrees with adjudicated label per guideline v3"
critical:
- "label derivable from free text of closing agent note (leakage)"
- "unredacted customer name, email, phone or account number"
minor: "trailing whitespace, casing, deprecated code alias"
not_an_error: "ambiguous ticket where guideline permits two codes"
adjudication: "two domain reviewers; third breaks ties; decisions logged"
sampling:
method: stratified_by_class_and_month
sample_size: 400
tolerances:
major_overall_max: 0.03
major_per_class_max: 0.06 # catches concentrated systematic error
critical_max_in_sample: 0
on_fail: "re-label affected strata and re-sample; not a price adjustment"
Two details in that spec do most of the work. The not_an_error line keeps honest ambiguity out of the error count, which matters because inter-annotator disagreement and label error are different quantities; see choosing agreement metrics. The per-class cap catches systematic error that an overall rate would average away.
When noisy data is still worth licensing
Noisy data is often still useful if the noise is measured, random and correctable, and if it is not destined for evaluation. A dataset with 5% estimated random error and strong coverage of a rare workflow can beat a clean dataset that never shows that workflow. The decision is about expected cost of cleaning against the value of coverage.
Ask three questions before rejecting:
- Can errors be found cheaply? Confident learning and similar methods rank likely label errors so reviewers inspect the top few percent instead of the whole set [1]. A walkthrough is in finding label errors with confident learning.
- Is the noise random? If audit errors cluster by annotator, time window or source system, treat the affected slice as a separate, lower-grade dataset.
- Can you split the use? Use the noisy bulk for training and carve a small, fully adjudicated subset for evaluation, never the other way round.
For how quality problems are handled before data reaches a buyer, see SourceX's notes on what happens when data is low quality and what happens when data contains errors. The wider framework of accuracy, completeness and consistency metrics sits in the data quality hub. If you would rather have operational records sourced with diligence materials you can check against a spec like this, tell SourceX what labeled data you need.
Setting label-error tolerances for data you license
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and finance and legal workflows, and manages the licensing process. Each dataset is rights-reviewed and comes with diligence materials on its source and preparation, which your team can review when setting and checking tolerances like the ones above. If you have a labeled or outcome-coded dataset in mind, describe the data you need.
Frequently asked questions
Is 1% label noise acceptable for an evaluation set?
It depends on the gap you need to resolve. If candidate models differ by several points, 1% may be workable; if they differ by about 1 point, 1% gold-label error can decide the ranking. Adjudicate every item where model and gold label disagree before trusting a close comparison [2].
Should label noise tolerance be stated in a data license?
Many buyers put the acceptance spec in a schedule or statement of work referenced by the license, with the error definition, sampling method and remediation path. Have counsel decide where it sits; the technical content is the same either way.
Does more data make up for noisier labels?
Only partly. More data helps against random noise, but systematic errors scale with the data, so the model learns the wrong rule more confidently. Cap systematic error per slice regardless of volume [6] [11].
Sources
- arXiv (Northcutt, Jiang, Chuang), "Confident Learning: Estimating Uncertainty in Dataset Labels" (2019). https://arxiv.org/pdf/1911.00068
- arXiv (Northcutt, Athalye, Mueller); NeurIPS 2021 Datasets and Benchmarks, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- International Conference on Machine Learning (ICML 2026), "Reliability-Aware LLM Alignment from Inconsistent Human Feedback" (2026). https://icml.cc/virtual/2026/poster/66787
- Open Exam Prep (CQE study guide), "4.5 Acceptance Sampling Plans & Standards". https://open-exam-prep.com/study-guides/cqe/product-and-process-control/acceptance-sampling-plans
- ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes" (2018). https://asq.org/quality-press/display-item?item=T1164
- arXiv (Eyuboglu et al.); ICLR 2022, "Domino: Discovering Systematic Errors with Cross-Modal Embeddings" (2022). https://arxiv.org/abs/2203.14960v2
- arXiv (Zhou et al.); NeurIPS 2023, "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- arXiv (Ouyang et al., OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- arXiv (Rafailov et al.); NeurIPS 2023, "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- arXiv (Gekhman et al.), "Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?" (2024). https://arxiv.org/pdf/2405.05904
- arXiv (Hestness et al., Baidu Research), "Deep Learning Scaling is Predictable, Empirically" (2017). https://arxiv.org/abs/1712.00409v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.