Skip to content

Image data

Rare Defect Coverage: Real, Pooled and Synthetic Defect Images for Imbalanced Inspection Data

Quick answer

Rare defect classes fail because the model has seen too few real examples of them, not because the architecture is wrong. Set a per-class coverage target (a working floor is a few hundred labeled images per defect class across severities and part variants), measure per-class recall on a held-out set of real defects, and close the gap in this order: longer historical archives, pooled defects from other plants or suppliers, then synthetic defects for geometric and texture variation. Algorithmic fixes help, but they do not create missing failure modes.

By SourceX Editorial · Updated

Why rare defect classes collapse while overall accuracy looks fine

Rare classes collapse because defects are low-probability events, so the loss is dominated by good parts and common defects. A published weld study of 31,193 real images found defects in only about 4% of them, and the authors applied hybrid oversampling to address the imbalance [1]. Inside that 4%, the distribution is itself skewed: porosity or spatter may be common while lack of fusion or burn-through appears a handful of times.

The same pattern shows up across inspection domains. Work on drone inspection of aircraft fuselage describes the rarest, most valuable samples as a few among thousands of annotated objects [3], and research on civil infrastructure inspection reports the same long-tail imbalance outside manufacturing [4]. The TL-SDD authors frame rare defect classes as the core failure mode of surface defect detection, noting that defect scarcity leads easily to overfitting [2].

The practical consequence is that headline metrics hide the problem. A model can report 98% image-level accuracy on a line where 96% of parts are good while missing most cracks. Report per-class recall, per-class precision and a confusion matrix by defect type, and track escapes (missed defects) separately from overkill (false rejects), because the cost of each differs by product.

What algorithmic fixes can and cannot do for few defect samples

Algorithmic remedies raise rare-class performance at the margin but cannot invent a failure mode the data never contains. Transfer learning from a source domain improved results on rare defect classes by up to 11.98% in TL-SDD [2], oversampling is a common response to skewed defect data [1], and focal or class-weighted loss and one-class anomaly detection are also reasonable first steps.

Each has a known failure mode on inspection data:

  • Oversampling and class weighting replay the same 15 crack images many times, which teaches the model those 15 cracks, including their lighting, fixture position and background, rather than cracks in general.
  • One-class or normal-only anomaly detection flags "different from good" but does not tell a cosmetic scratch from a structural crack, so it rarely satisfies a disposition rule that depends on defect type. See normal-only image sets for anomaly detection for where that approach fits.
  • Few-shot and transfer methods still need a real evaluation set per class; without one, you cannot tell whether the gain is real.
  • Active learning prioritizes which unlabeled images to label [4], which helps only if the rare defects already exist somewhere in your unlabeled archive.

When the bottleneck is examples rather than method, the answer is more real defects, more varied defects, or both.

How many images per defect class an inspection model needs

There is no universal number, but you can set a defensible floor from your evaluation needs first, then your training needs. Per-class recall measured on 20 real examples moves 5 percentage points with each miss, so a test set that small cannot distinguish 90% from 95% recall. Label noise compounds this: even well-known benchmark test sets carry an average label error rate of at least 3.3% [6], and a few mislabeled rare-defect images can swing a small class.

A workable planning rule is to reserve roughly 50 to 100 real, independently verified examples per critical defect class for evaluation, then target several hundred training examples per class spread across severities, part numbers, lines and camera setups. For general sizing logic across tasks, see how many images you need to train a computer vision model. Treat these as starting hypotheses to test with learning curves, not as thresholds.

The arithmetic usually explains why a single plant cannot get there. If a defect class occurs on 1 in 5,000 inspected parts and you need 300 examples, you need roughly 1.5 million inspected images that were retained and are labelable. Many lines keep only rejected images, overwrite good-part images after days, or store crops without the full frame, so the archive is smaller than production volume suggests.

Coverage spec for rare defect classes

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample entryWhy it matters
defect_classweld.lack_of_fusionMust map to your taxonomy, not the supplier's free text
min_train_images300Training floor, tested with learning curves
min_eval_images80, real only, from held-out lines or datesLets per-class recall be estimated meaningfully
severity_rangeminor / reject / critical, each at least 20% of classPrevents a model that only finds obvious defects
part_variantsat least 4 part numbers or geometriesAvoids learning one fixture's background
capture_setupcamera model, lens, lighting, resolution, modality (visible, X-ray, IR)Domain shift is the usual cause of pooled-data failure
time_spanat least 12 production monthsCaptures tool wear, material lots, seasonal process drift
label_typebounding box or polygon mask plus disposition (accept, rework, scrap)Ties pixels to the decision the model supports
label_provenanceinspector ID role, re-inspection result, destructive test confirmation where availableSeparates confirmed defects from suspected ones
negativesgood parts from the same lines and datesKeeps precision honest

Pooling real defects across plants and suppliers

Pooling real defects from several plants or companies is usually the fastest way to reach rare-class floors, but it only works if taxonomies, capture conditions and license terms are harmonized before training. Two suppliers may call the same flaw "cold lap," "lack of fusion" and "LOF," or split one of your classes into three. Build a crosswalk table from each source's codes to your taxonomy and have a domain reviewer resolve ambiguous mappings on sample images, not on code names alone.

Capture conditions matter as much as labels. Pooled sets mix sensor resolution, bit depth (8-bit JPEG versus 16-bit TIFF or DICONDE radiographs), lighting geometry and compression, and models readily learn "plant B means defect" if one source contributed mostly rejects. Stratify by source in your splits, report per-source recall, and check that each source contributes good parts as well as defects. ISO/IEC 5259-2 gives a vocabulary of measurable data quality characteristics that is useful for writing these acceptance checks into a contract [7].

Licensing is per source. Each contributing company sets its own allowed uses, term and delivery conditions, and those terms can differ even when the images look interchangeable. Ask every supplier the questions in evaluating data supplier quality, and for domain-specific sourcing see industrial defect image datasets, weld inspection image datasets and PCB and electronics inspection images.

Synthetic defect images vs real: where generated defects help and where they mislead

Synthetic defects are useful for expanding geometric and texture variation of defects you already understand, and risky as the only source for process-specific failure modes. Copy-paste of real defect crops onto new good parts, procedural texture generation, rendered CAD scenes and diffusion or GAN inpainting can all add position, scale, orientation and background variety at low cost.

The evidence for caution is concrete. DAGM, one of the most widely used defect benchmarks, is synthetic: small grayscale images with at most one defect each, and a benchmark review found statistical bias relative to real production data [5]. A model tuned on such images can score well and still miss the multiple, overlapping, low-contrast defects real lines produce.

Generated defects also inherit the generator's assumptions. Inpainting models trained on a handful of real cracks reproduce those cracks' morphology; rendered porosity follows the noise model the artist chose, not the gas chemistry of your process. Three rules keep synthetic data honest:

  1. Never put synthetic images in the evaluation set; measure every model on held-out real defects only.
  2. Track the synthetic ratio per class and run an ablation (real only versus real plus synthetic) for each rare class.
  3. Seed generation from real examples of each class, and stop adding synthetic data when real-defect recall stops improving.

For the broader trade-off, see combining licensed and synthetic data and licensed vs synthetic vs scraped AI training data.

Decision table: real, pooled or synthetic for each gap

Illustrative example: invented to show structure; it does not describe an available dataset.

Gap you observeFirst moveSecond moveAvoid
Class has under 30 real examples anywhere in your archiveExtend archive retrieval to older months and other linesPool real examples from other plants or suppliersTraining on synthetic only
Class exists but only at one severityRequest severity-stratified real examplesSynthetic variation of size and contrast, seeded from realOversampling the same images
Good recall on one line, poor on anotherPool data with the failing line's camera and lightingDomain-specific fine-tuningTreating it as a class-imbalance problem
Defect is geometric (scratch, dent, missing component)Synthetic placement on real good partsReal hard negativesSkipping a real-only eval
Defect is process-driven (porosity, delamination, cold solder)Real examples from comparable processesPooled data with documented process parametersRendered or generated-only data

What to ask for when sourcing rare defect images

Ask for historical archives with confirmed dispositions, not just image dumps. The most valuable rare-defect data usually sits in quality systems that link an image to a nonconformance report, MRB (material review board) decision, rework record or destructive test result, which is why manufacturing quality records often matter as much as the images themselves.

A useful request specifies the defect classes and severity mix you need, the modality and minimum resolution, the time span, whether full frames and good parts are included, the label format (COCO JSON, Pascal VOC XML, YOLO txt or segmentation masks), and how labels were confirmed. Ask how images with operator faces, badge numbers or customer markings were handled, and whether the supplier can provide a sample before agreement. If you plan to label raw images yourself, compare costs in pre-labeled vs raw image datasets; for measuring tail coverage across modalities, see long-tail and edge-case coverage. The image data hub and AI data hub list related image categories.

SourceX sources operational datasets, such as engineering records and new recordings of hands-on work, from US companies on request; it does not hold stock, and a request does not guarantee a match. Buyers describe the defect data they need through SourceX for buyers, and every release is approved by the supplying company. Related category detail is on images and inspection photos.

Request real rare-defect inspection images

SourceX looks for US businesses that hold the defect images and quality records you describe, reviews ownership and consents, and delivers each dataset under a license that defines records, uses, term and delivery. Nothing is contracted until a supplier agrees, and terms are set per deal. Describe your defect classes and coverage spec at https://sourcex.si/buyers.

Sources

  1. Alexandria Engineering Journal (via DOAJ), "Welding defect image study with hybrid oversampling" (2025). https://doaj.org/article/a7f91d9ea879427fa513ad25ff017d0e
  2. arXiv, "TL-SDD: A Transfer Learning-Based Method for Surface Defect Detection with Few Samples" (2021). https://arxiv.org/pdf/2108.06939
  3. ETH Zurich, "Drone-image inspection of aircraft fuselage" (Research Collection). https://www.research-collection.ethz.ch/items/592ca7f6-bc65-41aa-be82-04f12187ae84
  4. arXiv, "Active learning for imbalanced civil infrastructure inspection data" (2022). https://arxiv.org/pdf/2210.10586
  5. arXiv, "A Review of Benchmarks for Visual Defect Detection in the Manufacturing Industry" (2023). https://arxiv.org/pdf/2305.13261
  6. arXiv, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data