Skip to content

Image data

Industrial Defect Image Datasets for Visual Inspection Models

Quick answer

An industrial defect detection dataset for a production model needs images from your kind of line, with your defect taxonomy, labels at the right granularity, and capture metadata. Public benchmarks such as MVTec AD, DAGM, VISION and Real-IAD are useful for method selection and pretraining, but they are small, partly synthetic or captured under lab conditions [1][2]. Teams that cannot collect enough of their own defects often license labeled production images from manufacturers and test them against the gaps described below.

By SourceX Editorial · Updated

Where public defect benchmarks stop representing production

Public benchmarks are good for comparing algorithms and weak as a proxy for a specific line. A 2023 review of visual defect benchmarks found that DAGM is synthetic, holds at most one defect per image, and shows statistical bias compared with real production data [1]. The same review describes MVTec AD as 5,354 images across 15 categories, with a training split that contains only normal images [1]. That design fits unsupervised anomaly detection, not supervised classification of named defect types.

Leaderboard headroom is the second problem. Published anomaly detection methods report very high image-level scores on MVTec AD, so small differences between methods say little about your parts. Newer multi-view collections such as Real-IAD were built partly in response, but they are still lab-captured objects. A model that tops a crowded leaderboard can still miss the low-contrast scratch that your customer rejects.

Scale is the third. VISION pulls together 14 industrial inspection datasets with about 18k images and 44 defect types under instance segmentation [2], which is helpful breadth but modest next to the image volume of a single high-throughput line. Process-specific sets, such as weld surface images with pixel-level ground truth on IEEE DataPort [4], exist because general benchmarks rarely cover one process in depth. For weld and electronics work, see weld inspection image datasets and PCB and electronics inspection images.

MVTec AD, DAGM, VISION and Real-IAD compared

Each benchmark answers a different question, so pick by the task you need to validate. The comparison below uses only figures reported in the cited papers; read each dataset's own license file before any commercial training use.

BenchmarkWhat it containsUseful forGap versus production
DAGMSynthetic textured surfaces, at most one defect per image [1]Quick sanity checks of segmentation codeSynthetic statistics; no multi-defect parts [1]
MVTec AD5,354 images, 15 categories, normal-only training split [1]Unsupervised anomaly detection baselinesNo defect examples to train on; single object per category [1]
VISION14 datasets, about 18k images, 44 defect types, instance masks [2]Breadth across materials and defect classesSmall per category relative to production volumes [2]
Real-IAD [7]Multi-view, higher-resolution captures of many object types (check the paper for counts)Multi-view evaluationStill lab-captured objects, not your line or part revisions

Whichever benchmark you use, test the assumption that training folders are clean. Production "good" folders are rarely clean, because escapes, mis-sorts and borderline parts end up there, and normal-only methods learn whatever contamination they contain. Evaluate with a deliberately contaminated normal set before you trust an anomaly score threshold on the line.

What real production defect data adds

Production images carry the conditions your model will actually face. The production gaps worth testing for are several defects on one part, line-specific lighting and camera drift, part revisions that change geometry, and defect taxonomies that change when quality engineering splits or merges codes. Defects are a small-probability event on a healthy line, which makes defect classes scarce and pushes models toward overfitting [3]; see rare defect coverage and class imbalance for pooling and synthetic options.

The most valuable addition is business meaning. When each image links to its nonconformance report (NCR) and material review board (MRB) disposition, such as scrap, rework or use-as-is, the label reflects what the plant actually decided, not just what an annotator saw. That linkage usually lives in manufacturing quality records and inspection reports, not in the image store.

Labels also need auditing on arrival. Even widely used benchmark test sets show an average label error rate of at least 3.3% [5], and borderline cosmetic defects are where inspector decisions are most likely to disagree. Budget a relabel pass on a stratified sample before you trust any supplier's labels as ground truth.

How to specify a defect image request

A usable request names the part, the defect taxonomy, the label type and the capture conditions. Ambiguous requests ("surface defect images") produce mismatched deliveries. State whether you need normal-only images for anomaly detection (see normal-only image sets), image-level class labels, bounding boxes, or pixel masks.

Illustrative example: invented to show structure; it does not describe an available dataset.

request: production defect images for supervised segmentation
part_family: machined aluminum housings (2 part revisions)
material_and_finish: 6061 aluminum, anodized
defect_taxonomy: [scratch, dent, porosity, burr, anodize_discoloration, other]
taxonomy_version_history: required (code splits/merges with dates)
labels:
  image_level: defect_class (multi-label allowed)
  pixel_level: polygon or PNG mask per defect instance
normal_images: included, ratio to defective noted
per_image_metadata:
  - line_id and station_id
  - camera_model, lens, resolution, bit_depth
  - lighting_setup (ring, dome, coaxial, backlight)
  - part_number and revision
  - capture_timestamp (UTC)
  - inspector_decision (pass / fail / hold)
  - ncr_id and mrb_disposition (scrap / rework / use-as-is)
file_formats: lossless PNG or TIFF; labels in COCO JSON
split_guidance: by time window and line, not random

Split by time and by line, not randomly. Random splits put near-duplicate frames of the same part in train and test and inflate scores. For variation across optics, see camera, lens and lighting diversity.

Rights and review issues specific to factory images

Factory images raise rights questions that generic photo licenses do not cover. Parts made to a customer's drawing can reveal that customer's design, so the supplier's customer contracts, NDAs and IP clauses need review before release. Frames can also capture operators' faces or badges, screen contents, and part labels with serial or customer numbers, which should be cropped, masked or removed.

If your inspection model will be a safety component of a regulated product, data governance duties may apply. Article 10 of the EU AI Act requires high-risk systems to be built on training, validation and testing data that meet quality and governance criteria [6]. As of October 2026, Article 10 was amended by Regulation (EU) 2026/1744, which reportedly moved high-risk application dates to December 2027 and August 2028. Keep the per-image metadata and taxonomy history from your request; it is the provenance record reviewers will ask for.

How SourceX sources production defect images

SourceX sources operational datasets from US companies, including new recordings of hands-on work, and manages the commercial process through licensing and ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match. You describe the data you need; SourceX looks for US businesses that hold it, and every release is approved by the supplying company.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded, and a sample is checked, although no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. SourceX does not source scraped web content or generic CCTV or photos, and it does not train models. You can start from the SourceX buyer page, browse the image data hub, or see inspection photo licensing and manufacturing buyers.

Request labeled production defect images

Describe the part, defect taxonomy, label type and capture metadata you need, and SourceX will look for US manufacturers that hold matching data. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set in a per-deal license. Submit a defect image data request.

Sources

  1. arXiv, "A Review of Benchmarks for Visual Defect Detection in the Manufacturing Industry" (2023). https://arxiv.org/pdf/2305.13261
  2. arXiv, "VISION Datasets: A Benchmark for Vision-based InduStrial InspectiON" (2023). https://arxiv.org/html/2306.07890v1
  3. arXiv, "TL-SDD: A Transfer Learning-Based Method for Surface Defect Detection with Few Samples" (2021). https://arxiv.org/pdf/2108.06939
  4. IEEE DataPort, "Weld surface defect dataset". https://ieee-dataport.org/documents/weld
  5. arXiv (NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  7. arXiv (Wang et al., CVPR 2024), "Real-IAD: A Real-World Industrial Anomaly Detection Dataset" (2024). https://arxiv.org/abs/2403.07604

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data