Skip to content

Data quality, coverage and contamination

Finding Label Errors in a Licensed Dataset with Confident Learning

Quick answer

To find label errors in a dataset, train a classifier with K-fold cross-validation so every record gets an out-of-sample predicted probability, then apply confident learning: compute per-class confidence thresholds, count where the given label and the confidently predicted label disagree, estimate the joint distribution of given versus true labels, and rank records by label quality [1][2]. The open-source cleanlab package implements this [4]. The output is a ranked list of candidates for human review, not a list of confirmed errors.

By SourceX Editorial · Updated

What confident learning actually estimates

Confident learning estimates how often each given label corresponds to each true label, then uses that estimate to decide how many records to flag and which ones. The method assumes class-conditional noise: the chance that a record labeled "billing_dispute" is really "refund_request" depends on the pair of classes, not on the individual record [1]. That assumption fits operational labels well, because mislabels in support dispositions, claim outcome codes or document types usually come from a few confusable class pairs, not uniform randomness.

The core object is the confident joint, a K×K count matrix whose rows are given labels and columns are likely true labels [1]. A record is counted in cell (i, j) when its given label is i and its predicted probability for class j meets or exceeds class j's threshold, which is the average predicted probability of j among records labeled j [1]. Calibrating those counts to the dataset size gives an estimate of the joint distribution; the off-diagonal mass is the estimated noise, and its row sums tell you roughly how many records per class to flag [1][2].

Three principles drive the ranking, as the paper frames them: prune noisy data, count with probabilistic thresholds, and rank examples to train with confidence [1]. The per-class threshold matters because a model that is systematically underconfident on a rare class would otherwise flag most of it. The authors reported finding numerous label issues even in heavily used datasets such as ImageNet and CIFAR [1][2].

Why the probabilities must be out-of-sample

Confident learning only works if the predicted probabilities come from a model that never trained on the record being scored. A model fit on all the data memorizes noisy labels, so the mislabeled records receive high probability for their wrong class and disappear from the flags. Use stratified K-fold (five folds is a common default) and keep the held-out pred_probs matrix, shape N×K, with rows summing to 1.

Fold design is where operational data breaks naive setups. Split by group, not by row: all messages in one ticket thread, all pages of one loan file, all line items of one invoice and all records from one customer account belong in the same fold (scikit-learn's StratifiedGroupKFold handles this). Otherwise near-duplicates leak across folds and the model "remembers" labels through a sibling record; run near-duplicate detection first, as described in our guide to MinHash and LSH deduplication.

Model choice is secondary to calibration and coverage. For text dispositions, a fine-tuned encoder or a logistic regression over sentence embeddings is usually enough; for tabular category codes, gradient-boosted trees work. A weak model inflates flags in hard classes, so record per-class held-out accuracy alongside the flag counts.

Running the screen with cleanlab

In cleanlab 2.x, the screen is two inputs and one call: integer-encoded labels and out-of-sample pred_probs passed to cleanlab.filter.find_label_issues, which returns a boolean mask or, with return_indices_ranked_by="self_confidence", indices sorted from most to least suspicious [4][5]. cleanlab.rank.get_label_quality_scores gives each record a 0–1 score you can store as a column, and cleanlab.count.compute_confident_joint exposes the K×K matrix for reporting. The package also detects outliers, which are a different problem from mislabels and should go to a separate queue [4].

A practical sequence for a large delivery:

  1. Freeze the delivery: hash the files, record the label taxonomy version and the record ID field.
  2. Encode labels against the licensed taxonomy; quarantine values outside it (these are schema defects, not noise).
  3. Build group-aware folds and produce pred_probs for every record.
  4. Run find_label_issues with a pruning method such as prune_by_noise_rate, and save the confident joint.
  5. Join flags back to source IDs, timestamps, annotator or system-of-origin fields.
  6. Push the top-ranked candidates into a review tool; the Rubrix tutorial shows this pattern with cleanlab flags loaded into a labeling interface [5].

Reading the confident joint for real-world noise

Real-world label noise is sparse: most of the off-diagonal mass sits in a handful of class pairs, so read the confident joint as a confusion map before reading any single flag [1]. In a support-ticket taxonomy you might see "account_access" and "password_reset" exchanging labels in both directions while every other cell is near zero. That pattern usually indicates a taxonomy problem (overlapping definitions or a merged code), which relabeling records will not fix.

Asymmetric cells tell a different story. If many "escalated" records look like "resolved_first_contact" but not the reverse, check whether the source system defaulted the field, whether a workflow auto-closed tickets, or whether the code changed meaning after a CRM migration. Slice the flag rate by month, team, form version and source system; a flag rate that jumps on one date is a process change, and our page on verifying outcome labels in operational records covers that audit.

Hierarchical and multi-label codes need care. Run the screen at the level the model will train on (for HTS or UNSPSC codes, often a parent level first), and use cleanlab's multi-label mode instead of forcing one label per record when records legitimately carry several tags.

Flags are candidates, not confirmed errors

Every flagged record should be treated as a hypothesis until a qualified reviewer confirms it. The test-set audit by Northcutt, Athalye and Mueller used confident learning to find candidates and then had crowdsourced humans validate them; they estimated an average label error rate of at least 3.3% across ten widely used test sets, and at least 6% of the ImageNet validation set [3]. Note the two steps: algorithmic ranking, then human confirmation.

Reviewers should see the record, the given label, the model's suggested label and the taxonomy definition, and choose among: given label correct, suggested label correct, a third label, ambiguous (both defensible) or out of scope. Track the confirmation rate by rank bucket; it should fall as you move down the list, which tells you where to stop reviewing. If the top bucket confirms below roughly half, revisit the model or the folds before blaming the supplier.

Where the delivery includes several annotators per record, cleanlab's multi-annotator functions estimate a consensus label, a consensus quality score and per-annotator quality from the same pred_probs [4]. That complements agreement statistics such as Krippendorff's alpha, covered in inter-annotator agreement metrics, and the disagreement-handling methods in annotation adjudication.

Comparing label-noise detection methods

Confident learning is one of several approaches, and the right one depends on what metadata the delivery carries. The table below compares the options an ML data engineer typically weighs.

MethodNeedsStrengthMain failure mode
Confident learning (cleanlab)Out-of-sample pred_probs, single labelsEstimates noise per class pair; ranks recordsLeaky folds or weak model hide or inflate errors
Loss or margin rankingAny trained modelSimple to computeNo principled cutoff; biased toward hard classes
Multi-annotator consensusSeveral labels per recordSeparates annotator quality from item difficultyAbsent when operational labels have one source
Random audit samplingReviewer timeUnbiased error-rate estimateInefficient at finding rare errors
Rule checksField logic (dates, codes, status)Catches impossible labels cheaplyMisses plausible-but-wrong labels

Pair confident learning with a random sample: the ranked list finds errors efficiently, while a random sample gives the unbiased rate that a contract can reference. Our guide to acceptance sampling for dataset deliveries explains how to size that sample, and how much label noise is acceptable covers tolerance by use (training, SFT, evaluation).

Turning flags into a remediation request

A remediation request should give the supplier record IDs, evidence and a clear ask, not a model score alone. Share the flagged IDs with the given and suggested labels, the confirmed rate from human review, the class pairs involved and the taxonomy version; then ask for relabeling of confirmed records, a corrected codebook if the confusion is definitional, or a commercial adjustment if your license allows one. Whether a credit, replacement or relabel applies depends on the terms you negotiated, so check the delivery and acceptance clauses first.

Illustrative example: invented to show structure; it does not describe an available dataset.

remediation_request:
  delivery_id: "DLV-2026-0412"            # hash-pinned delivery
  taxonomy_version: "dispositions_v3.2"
  method: "confident_learning; 5-fold StratifiedGroupKFold by thread_id"
  records_screened: 120000
  flagged_by_cl: 4310
  human_reviewed: 1200                     # top 800 ranked + 400 random
  confirmed_errors_in_reviewed: 512
  random_sample_error_rate: "2.1% (95% CI 0.9-4.1%)"
  top_confusions:
    - { given: "password_reset", likely_true: "account_access", confirmed: 188 }
    - { given: "escalated", likely_true: "resolved_first_contact", confirmed: 97 }
  attachments: ["flags_ranked.parquet", "confident_joint.csv", "review_log.csv"]
  ask:
    - "Relabel confirmed records and return a diff keyed on record_id"
    - "Clarify codebook definitions for the two top confusions"
    - "Confirm whether 'escalated' was auto-populated before 2024-03"

Keep the review log: record ID, reviewer, decision, timestamp and rationale. For high-risk systems under the EU AI Act, As amended by Regulation (EU) 2026/1744, Article 10 requires training, validation and testing data to meet quality criteria under documented data governance practices; as of October 2026 the high-risk application dates have reportedly moved to 2 December 2027 for Annex III systems [6][8]. Even outside that scope, a Data Card or similar record of annotation methods and known label issues helps downstream teams interpret model performance [7].

Where confident learning stops helping

Confident learning cannot fix labels that are consistently wrong in the same direction across the whole dataset, because the model learns the bias as truth. Historical decisions in lending, hiring or claims records are the classic case; see historical decision bias in operational labels. It also struggles with free-text targets for SFT, where there is no fixed class set; use rubric-based review or the filtering approaches in instruction-tuning data quality filtering.

For evaluation sets, screen before you score anything, because mislabeled test items distort model rankings [3]. Do not "clean" an eval set with the same model family you plan to evaluate, and record which items were changed. The wider assessment framework sits in our training data quality hub, and the data annotation glossary entry defines the labeling terms used here.

Sourcing labeled operational data you can screen

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows; nothing is held in stock, and a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery, with diligence materials prepared per dataset, so you can describe the labels and taxonomy you need on the SourceX buyer page.

Request labeled data for your classifier or eval set

If you need classification labels, dispositions or category codes from real business workflows, describe the data, label fields and intended use rather than specific companies. SourceX looks for US businesses that hold it, and nothing is contracted until a supplier agrees. Start a buyer request.

Sources

  1. arXiv (Northcutt, Jiang, Chuang), "Confident Learning: Estimating Uncertainty in Dataset Labels" (2019). https://arxiv.org/pdf/1911.00068
  2. Journal of Artificial Intelligence Research, "Confident Learning: Estimating Uncertainty in Dataset Labels" (2021). https://www.jair.org/index.php/jair/article/view/12125
  3. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  4. Python Package Index, "cleanlab 2.2.0". https://pypi.org/project/cleanlab/2.2.0
  5. Rubrix documentation, "Find label errors with cleanlab". https://rubrix.readthedocs.io/en/v0.4.1/tutorials/07-find_label_errors.html
  6. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  7. arXiv (Pushkarna, Zaldivar, Kjartansson), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  8. EUR-Lex, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data