Image data
Normal-Only Image Sets for Unsupervised Anomaly Detection
Quick answer
An unsupervised anomaly detection project needs two different datasets: a large training set of verified defect-free (nominal) images that spans every acceptable variation on the line, and a smaller, carefully labeled test set mixing nominal and anomalous parts, ideally with pixel masks. Public benchmarks such as MVTec AD use exactly this split [1][2]. The training set teaches what "good" looks like; the test set is the only proof the model catches real defects without flagging benign variation.
By SourceX Editorial · Updated
Why one-class training changes the data specification
One-class training removes the need for thousands of labeled defects, but it moves the burden onto nominal coverage and test-set quality. Defects are low-probability events in production, so supervised defect classifiers starve for positive examples and overfit the few they get [3]. One-class methods instead model normal products, for example with autoencoders or memory banks of features from nominal images, and score a new image by how far it deviates from that normal model [2].
The consequence for buyers: anything absent from the nominal set looks anomalous at inference. A lot of resin from a second supplier, a recalibrated ring light or a new fixture can all trigger false rejects if the training images never showed them. Reconstruction-based and feature-memory methods share this dependency, because both learn only the normal distribution.
If you are still deciding between this approach and supervised defect detection, compare it with our guide to industrial defect image datasets and the options for rare defect coverage with real, pooled and synthetic images.
What the nominal training set must cover
A nominal set is fit for purpose when it contains every source of acceptable variation the deployed camera will see, sampled across time. Count images per variation stratum, not in total. A few hundred images from one shift on one day can produce a clean benchmark score and a noisy production model.
Variation axes to stratify:
- Product variants: SKUs, colors, revisions, and parts that share a station.
- Within-tolerance geometry: dimensional spread, acceptable cosmetic marks, mold cavity numbers, and print registration drift that QA accepts.
- Material lots and suppliers: resin, coil, PCB fabricator, and finish batches, keyed by lot ID.
- Imaging conditions: lens, exposure, lighting angle and aging, camera position after maintenance, and dust on the enclosure. See our guide to camera, lens and lighting diversity.
- Pose and fixturing: part rotation, conveyor position, and partial occlusion by grippers.
- Time: seasons, line restarts, and the period after preventive maintenance.
The acoustic anomaly detection community documents the same failure: detectors trained on normal machine sounds degrade when operating and environmental conditions shift between training and deployment [6]. Our page on industrial machine sound data for anomaly detection covers that domain-shift design for audio.
How many good images you need
There is no universal count; the honest answer is "enough per stratum that held-out nominal images score low." Feature-memory methods such as PatchCore subsample their nominal bank (coreset selection) to keep inference fast as it grows, so raw volume is rarely the limit. The binding constraint is usually variation coverage, not raw volume.
A practical way to size the request:
- List the variation strata above for one inspection station.
- Ask for an initial nominal pull covering every stratum, sampled across at least several production weeks rather than consecutively.
- Hold out a nominal validation slice per stratum and plot the false-positive rate as you add training images.
- Stop adding images to a stratum when its false-positive curve flattens; add more only where it stays high.
This learning-curve approach turns "how many good images" into a measured question and gives you a defensible reason to request more data from specific lots or conditions.
Contaminated "normal" sets and how to detect them
A nominal set that quietly contains defective parts teaches the model that those defects are normal. This is a common hidden failure in one-class inspection data, because "passed inspection" is not the same as "defect-free": legacy automated optical inspection (AOI) systems and human inspectors both have escape rates.
Ask the supplier how nominal status was established, and prefer images with one of these provenance signals:
- The part passed downstream functional test or end-of-line test, not only the camera station.
- The part shipped and received no field return or warranty claim within a defined window.
- A second human reviewer re-inspected a random sample of the nominal pull and recorded the escape rate.
On receipt, run your own contamination screen: train a first model, score the training set itself, and manually review the highest-scoring nominal images. Label errors are not a niche concern. An audit of widely used benchmark test sets estimated an average label error rate of at least 3.3%, enough to change which model ranks first [5].
Designing the anomaly test set
The test set is a held-out mix of nominal and anomalous images, labeled at image level and, where possible, with pixel masks so you can measure localization as well as detection [1][2]. Common reporting combines image-level AUROC with pixel-level AUROC and the per-region overlap (PRO) metric, which needs region masks to compute.
Requirements that matter more than size:
- Defect taxonomy coverage: include every defect class your quality team tracks (scratch, dent, contamination, missing component, misprint, crack), and include subtle, near-tolerance cases, not only obvious ones. Real-world defect types vary widely across products [2].
- Nominal negatives from the same period: test nominal images must come from the same lots and conditions as the anomalies, or the model learns to detect the lot rather than the defect.
- Strict separation: no part serial, lot, or capture session should appear in both training and test splits.
- Independent labeling: masks drawn by someone other than the person who curated the training set, with a second-review sample.
- Realistic prevalence reporting: test sets are enriched with defects, so report precision at the line's true defect rate separately.
Benchmark saturation is a reason to build your own test set. Leading methods already exceed 99% AUROC on MVTec AD, so it no longer separates them; Real-IAD was introduced with multi-view real-world captures partly for that reason [4]. Your private test set is the benchmark that predicts your line. Our page on AI evaluation datasets built from real business work explains why held-out operational data makes the stronger test.
Request specification: nominal set plus anomaly test set
The specification below is the artifact to send to a data supplier or internal QA owner. It separates the two deliverables, because they often come from different systems and carry different sensitivity.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Nominal training set | Anomaly test set |
|---|---|---|
| Purpose | One-class training (normal only) | Held-out evaluation (normal and anomalous) |
| Station and camera | Station ID, camera model, lens, lighting config ID | Same stations as training |
| Image format | Lossless PNG or TIFF at native sensor resolution, no resizing | Same |
| Nominal verification | Passed end-of-line test plus a stated re-inspection sample | Nominal negatives verified the same way |
| Strata | Variant, lot, shift, lighting config, capture week | Same strata, distinct lots and sessions |
| Labels | status=nominal only | Image label plus defect class, PNG pixel mask, severity |
| Metadata per image | image_id, part_serial_hash, lot_id, station_id, capture_ts, light_config, qa_result | Same plus defect_class, mask_path, reviewer_id |
| Split rule | No serial, lot or session overlap with test | Group-held-out by lot and session |
| Exclusions | Images with operators, badges or screens in frame | Same |
| Documentation | Collection method, inspection escape-rate estimate, known gaps | Labeling guide, reviewer agreement on a sample |
A JSONL manifest row for the test set might look like this:
{"image_id": "st04-000812", "split": "test", "status": "anomalous", "defect_class": "scratch", "mask_path": "masks/st04-000812.png", "lot_id": "L-2291", "station_id": "ST04", "capture_ts": "2026-03-14T09:12:05Z", "light_config": "ring-v2", "reviewer_id": "R2"}
Strip or review camera EXIF and embedded timestamps against your privacy and confidentiality needs; our guide to EXIF metadata in image training data covers what to keep.
Sourcing both sets from operational records
Most of the value sits in plant systems that already exist: AOI image archives, vision-system reject bins, MES lot genealogy, and nonconformance reports. Nominal images are plentiful because inline cameras capture every part; the scarce assets are verified nominal status and labeled anomalies with masks. That makes quality records, such as NCRs, rework tickets and end-of-line test results, as important as the pixels. See manufacturing quality datasets for the record types that usually travel with inspection images.
Points to settle with a supplier before delivery:
- Rights: confirm the company owns the images and that customer contracts or NDAs do not restrict sharing images of customer-designed parts.
- Confidentiality: product geometry can reveal trade secrets; agree on which stations and variants are in scope.
- Use terms: if a supplier is more comfortable releasing anomaly images for evaluation than for training, write that distinction into the license explicitly.
- Refresh: decide whether you need new nominal pulls after line changes, since that is when one-class models drift.
For electronics specifically, our page on PCB and electronics inspection images covers AOI-specific formats and defect classes. For the broader category map, start at the image data hub.
SourceX sources operational datasets, including engineering records and new recordings of hands-on work, from US companies and manages the licensing process. Images are sourced on request rather than held in stock, and every release is approved by the supplying company. Buyers can describe the nominal and test images they need without naming specific businesses.
Governance when the inspection model is high-risk
If the inspection system is a high-risk AI system under the EU AI Act, for example a safety component of machinery or another product covered by the Union harmonisation legislation in Annex I, Article 10 requires training, validation and testing data sets that meet quality criteria and sit under documented data governance [7]. As of October 2026, Regulation (EU) 2026/1744 has reportedly moved the high-risk application dates to 2 December 2027 for Annex III and 2 August 2028 for Annex I and also amends Article 10 [8]. The documentation in the specification above (provenance, nominal verification method, strata, known gaps and labeling guide) maps directly onto that governance record. See EU AI Act Article 10 for licensed training data for detail.
Request nominal and anomaly inspection images
SourceX looks for US businesses that hold the inspection data you describe, reviews ownership and consents for each dataset, and delivers it under a license that defines records, uses, term and delivery. Nothing is contracted until the supplier agrees, and a request does not guarantee a match. Describe your normal-only training set and anomaly test set to SourceX.
Sources
- arXiv, "A Review of Benchmarks for Visual Defect Detection in the Manufacturing Industry" (2023). https://arxiv.org/pdf/2305.13261
- arXiv, "VISION Datasets: A Benchmark for Vision-based InduStrial InspectiON" (2023). https://arxiv.org/html/2306.07890v1
- arXiv, "TL-SDD: A Transfer Learning-Based Method for Surface Defect Detection with Few Samples" (2021). https://arxiv.org/pdf/2108.06939
- arXiv (CVPR 2024), "Real-IAD: A Real-World Multi-View Dataset for Benchmarking Versatile Industrial Anomaly Detection" (2024). https://arxiv.org/pdf/2403.12580
- arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (Tanabe et al., Hitachi), "MIMII DUE: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection with Domain Shifts" (2021). https://arxiv.org/pdf/2105.02702
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- European Parliament and Council of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.