Skip to content

Data quality, coverage and contamination

Acceptance Sampling for Dataset Deliveries: Adapting AQL Plans to Record and Label Defects

Quick answer

Acceptance sampling for data labeling treats each dataset delivery as a lot, each record or label as a unit, and a pre-agreed defect definition as the pass/fail test. You pick a plan (sample size n, acceptance number c) from ANSI/ASQ Z1.4 or compute one from its operating characteristic (OC) curve, inspect a random sample, and accept or reject the whole delivery. Switching rules then tighten or relax inspection as the supplier's history accumulates [1][2].

By SourceX Editorial · Updated

Mapping manufacturing lots onto dataset deliveries

The mapping works when you define the lot, the unit and the defect before the first delivery arrives, not after you see the data. Manufacturing plans assume a lot is a homogeneous batch from one process [3]; in data, the closest analogue is one delivery drop (a weekly Parquet partition, a monthly export of support tickets, a labeled batch from one annotation run).

Manufacturing termDataset equivalentPractical note
LotOne delivery or partition (e.g., delivery_id = 2026-10-W2)Split a delivery into separate lots if it mixes source systems, labeling vendors or guideline versions
UnitOne record, one label, or one annotated itemChoose one per plan; do not mix record-level and label-level defects in the same count
Defective unitA unit with at least one critical or major defectCount defectives (attribute plan), not total defects, unless you deliberately use a defects-per-hundred-units plan
AQLWorst process average you will tolerate over many deliveriesNot a quality target [2]
LTPD / LQDefect rate you want rejected almost every timeSet from model impact, not convenience
Inspection levelHow much sample per lot sizeGeneral Level II is the usual default in Z1.4 [2]

Heterogeneity is the main failure mode. If 90% of a delivery came from one annotation team and 10% from a new one, a random sample of 200 holds only about 20 of the new team's units, often too few to reject the lot even if that team's work is poor. Stratify, or treat each source as its own lot.

Writing defect definitions that an inspector can apply

A defect definition must be binary, testable on a single unit, and written into the delivery specification before inspection. Vague definitions ("label is wrong") produce reviewer disagreement that looks like supplier error; measure reviewer agreement first using inter-annotator agreement metrics.

Classify defects by severity, as manufacturing does with critical, major and minor classes, and give each class its own AQL:

  • Critical (record): residual direct identifier in free text, record outside the licensed scope, duplicate primary key, record from a source system not named in the license. For personal-data residue, use the dedicated plan in auditing residual PII with sampling rather than a general AQL.
  • Major (label): class label contradicts the guideline, bounding box IoU below the agreed floor, transcript segment with wrong speaker, outcome field inconsistent with the event history (see verifying outcome labels in operational records).
  • Minor (format): timestamp in the wrong timezone, trailing whitespace in a categorical field, non-normalized currency code.

Many structural defects (schema drift, nulls in required columns, referential integrity breaks) should be caught by 100% automated checks, not sampling. Sampling is for defects that need human judgment; run structured dataset validation checks first and sample only what passes them.

Reading a Z1.4 plan: code letters, n, Ac and Re

A Z1.4 single sampling plan is read in two steps: lot size and inspection level give a code letter, and the code letter plus AQL give the sample size n with an acceptance number (Ac) and rejection number (Re) [1][2]. If the sample contains Ac or fewer defectives you accept the lot; Re or more, you reject it. This is the same (n, c) plan the NIST handbook describes, where the lot is rejected when the sample holds more than c defectives [4].

The tables descend from MIL-STD-105E, whose code-letter and master tables are reproduced online [5]. Z1.4:2003 (R2018) is the edition ASQ lists as current as of October 2026 [1]; confirm the edition in your agreement, because table values and switching details should be quoted from the edition you actually reference.

Illustrative example: invented to show structure; it does not describe an available dataset.

DeliveryLot size (labels)Code letter (Level II)nAQL (major)AcReDefectives foundDecision
D-148,400L2001.0563Accept
D-159,100L2001.0567Reject
D-167,950L2001.0564Accept

Check the letter and Ac/Re values against your copy of the standard before use. Sample with a seeded random generator over stable record IDs (for example, ORDER BY hash(record_id || seed) and take the first n) and log the seed, so the supplier can reproduce exactly which units were inspected.

Reading the OC curve before you accept a plan

The OC curve tells you the probability of accepting a delivery at each true defect rate, and it is the most direct way to judge what a plan actually protects against [3]. Acceptance probability for a single plan follows the binomial distribution for large lots [4]; you can compute it in a spreadsheet, Python or the AQLSchemes R package, which also retrieves standard schemes [6].

Illustrative example: invented to show structure; it does not describe an available dataset. Binomial acceptance probabilities computed for three plans:

True defect raten=200, c=5n=200, c=3n=80, c=2
0.5%99.9%98.1%99.2%
1.0%98.4%85.8%95.3%
2.0%78.7%43.1%78.4%
3.0%44.3%14.7%56.8%
4.0%18.6%4.0%37.5%
5.0%6.2%0.9%23.1%

Read it as a buyer. With n=200 and c=5, a delivery that is truly 3% defective still passes about 44% of the time. If 3% major label error is unacceptable for your use, that plan does not protect you; tighten c, raise n, or lower the AQL. The small n=80 plan accepts a 5% defective lot nearly one time in four, which is why small samples feel reassuring and are not. Set the LTPD from what the error does to your model, using acceptable label noise by use, then pick the cheapest plan whose OC curve rejects that rate with the confidence you need.

Why AQL is not a quality target

AQL is the worst process average you are willing to tolerate over a long series of lots, not a promise that any single delivery meets it [2]. A plan at AQL 1.0 is designed to accept most lots from a process running at 1% defective; it says little about a single lot at 2.5%. Do not write "the dataset will be 99% accurate" into a contract and point to an AQL plan as proof.

Pair the AQL with an explicit per-delivery limit (LTPD) and a separate estimate of the overall error rate when you need a quantified quality claim. For one-off estimation with confidence intervals, a full annotation quality audit is the better tool; for a reported number from the supplier, compare against what a dataset quality report should contain.

Switching rules for recurring supply

Switching rules convert a static plan into a scheme that reacts to supplier history, and they are where Z1.4 adds the most value for ongoing deliveries [1]. As summarized in the study guide, inspection moves from normal to tightened when 2 of 5 consecutive lots are rejected, and back to normal after 5 consecutive tightened lots are accepted [2]. Reduced inspection requires a run of accepted lots plus stable production and approval from the responsible authority; confirm the exact conditions (including the switching score in the 2003 edition) against your copy [1][2].

Translate those states into delivery operations:

  1. Normal: default plan; record decision, defectives and sample seed per delivery_id.
  2. Tightened: typically a smaller Ac for the same code letter, taken from the tightened table; notify the supplier, request root cause (guideline change, new annotators, source system migration).
  3. Discontinue: if tightened inspection persists, the standard provides for stopping acceptance until corrective action; mirror this with a contractual hold on further deliveries.
  4. Reduced: smaller n once history is strong; re-enter normal on any rejection or on a process change such as a schema change across recurring deliveries.

A process change should reset the history even without a rejection. New labeling guidelines, a different extraction query, or a switch from incremental to full-refresh deliveries means the old acceptance record no longer describes the current process.

Variables plans for continuous measures

When the quality measure is continuous, such as segmentation IoU, word error rate on a transcript, or timestamp offset, a variables plan usually needs fewer units than converting the measure to pass/fail. ANSI/ASQ Z1.9 is the companion standard for inspection by variables; it assumes the measure is roughly normally distributed, which IoU (bounded, often skewed) and WER (heavy-tailed on noisy audio) frequently are not. Check the distribution on early deliveries before relying on it; if it is skewed, a pass/fail threshold with an attribute plan is more defensible.

What to agree with the supplier before the first delivery

The plan only resolves disputes if both parties agreed to it in advance; write it into the delivery specification that sits alongside the license. Use this as a starting checklist.

Illustrative example: invented to show structure; it does not describe an available dataset.

acceptance_sampling_spec:
  standard: "ANSI/ASQ Z1.4-2003 (R2018)"
  lot_definition: "one delivery_id; split by source_system and guideline_version"
  unit: "label"            # or "record"; one per plan
  inspection_level: "General II"
  defect_classes:
    critical: { aql: 0.065, definition_ref: "spec section 4.1" }
    major:    { aql: 1.0,   definition_ref: "spec section 4.2" }
    minor:    { aql: 4.0,   definition_ref: "spec section 4.3" }
  ltpd_major: 0.03          # rate the plan must usually reject
  sampling: "seeded hash over record_id; seed logged per lot"
  reviewers: "two independent; disagreements adjudicated"
  switching_rules: "per standard; process change resets history"
  rejected_lot_remedy: "replace, relabel, or 100% screen; agree in contract"
  records_kept: [delivery_id, n, defectives_by_class, decision, seed]

Remedies for rejected lots (relabel, replace, screen every unit, price adjustment) and how acceptance links to payment belong in the commercial terms; see buyer acceptance and getting paid and data supplier SLAs. Teams sourcing recurring operational data through SourceX can describe the delivery they need and bring their acceptance spec into the license discussion. How refresh cadence and remedies are structured across a contract is covered in ongoing data supply agreements. Recording these acceptance records also supports a data quality management system of the kind ISO/IEC 5259-3 describes, which sets requirements without prescribing specific metrics [7].

Common ways acceptance sampling fails on data

Most failures come from treating the plan as math alone. Watch for these:

  • Non-random samples: inspecting the first 200 rows of a file sorted by date or by annotator.
  • Reviewer drift: your own reviewers become stricter over time, so the supplier appears to degrade; re-score a fixed calibration set each quarter.
  • Rejected lots resubmitted unchanged: require evidence of correction and inspect resubmissions under tightened rules.
  • Correlated defects: one bad guideline produces thousands of identical errors; a cluster sample by annotation task catches these better than unit sampling.
  • Deletion and correction churn: records withdrawn after acceptance still need tracking; see propagating deletions and corrections.

For a broader view of how acceptance sampling fits alongside coverage, bias and contamination checks, start at the training data quality hub or the AI data guide index.

Sourcing recurring operational data you can inspect

SourceX sources operational datasets, such as support histories, engineering records and finance workflows, from US companies on request and manages the licensing and ongoing purchases, with every release approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, which gives your acceptance plan a fixed reference point. If you need a recurring dataset, describe it to SourceX; a request does not guarantee a match.

Sources

  1. ASQ Quality Press, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes". https://asq.org/quality-press/display-item?item=T1164
  2. Open Exam Prep (CQE study guide), "4.5 Acceptance Sampling Plans & Standards". https://open-exam-prep.com/study-guides/cqe/product-and-process-control/acceptance-sampling-plans
  3. NIST, "NIST/SEMATECH e-Handbook of Statistical Methods: Lot Acceptance Sampling Plans". https://itl.nist.gov/div898/handbook/pmc/section2/pmc23.htm
  4. NIST, "NIST/SEMATECH e-Handbook of Statistical Methods: Single Sampling Plans". https://itl.nist.gov/div898/handbook/pmc/section2/pmc231.htm
  5. SQC Online, "MIL-STD-105E sampling tables by attributes". https://www.sqconline.com/node/2
  6. CRAN, "AQLSchemes package vignette". https://mirror.csclub.uwaterloo.ca/CRAN/web/packages/AQLSchemes/vignettes/AQLSchemes.pdf
  7. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Data quality for analytics and ML, Part 3: Data quality management requirements and guidelines" (2024). https://www.iso.org/standard/81092.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data