Skip to content

Image data

Weld Inspection Image Datasets: Visual and Radiographic Defect Data

Quick answer

A usable weld defect image dataset pairs images from one declared modality (visual surface, radiographic film or digital, or ultrasonic imaging) with labels mapped to a recognized imperfection taxonomy such as ISO 6520-1, plus the acceptance decision against a quality level like ISO 5817 B, C or D. Public sets are good for prototyping but are small, imbalanced or research-only. Production models usually need licensed images from fabricators, pipeline operators or inspection programs.

By SourceX Editorial · Updated

Weld inspection is not generic defect detection

Weld inspection needs its own data specification because modality, defect taxonomy and acceptance code all change what a label means. A crack on a cap-pass photo, a linear indication on a radiograph and an ultrasonic echo describe different physics, and a model trained on one rarely transfers to another. If you are scoping broader surface inspection, start with industrial defect image datasets; this page covers what is specific to welds.

Acceptance also depends on context. ISO 5817 sets three quality levels, B, C and D, with B the strictest, for fusion-welded joints in steel, nickel, titanium and their alloys. The same porosity cluster can pass at level D and fail at level B, so an "accept/reject" label without the governing standard and level is ambiguous training signal.

What public weld datasets cover, and where they stop

Public weld datasets are useful for baselines but rarely cover a production line's materials, joints and license needs. Three patterns recur:

  • Surface image sets with segmentation masks. The WELD dataset on IEEE DataPort provides weld surface images with pixel-level ground truth for several typical weld defect types [1]. That suits segmentation baselines but not radiographic or in-process monitoring.
  • Radiographic research collections. University NDT archives of weld X-ray images exist, often with bounding boxes or image-level labels in text files, but some restrict use to research and education and prohibit redistribution or commercial use. Read the terms before treating one as anything more than a benchmark.
  • Industrial studies with private data. A welding study on copper filter driers used 31,193 images in which roughly 4% were defective, and compared ResNet, EfficientNet, YOLOv8 and Swin [2]. That imbalance is typical of real lines, and the images themselves are usually not released.

Reviews of manufacturing defect benchmarks also note that some widely used sets, such as DAGM, are synthetic, and that many public images contain at most one defect each [3]. Real welds often carry several imperfections at once.

Define the defect taxonomy before you request images

A defensible weld dataset maps every label to a named imperfection class, and ISO 6520-1 is the common reference for fusion welding. It defines geometric imperfections with explanations and illustrations, and excludes metallurgical imperfections. Ask suppliers which taxonomy their inspectors used, and whether their labels are free text ("PO", "porosity", "gas") that you must normalize.

Typical classes buyers request are porosity and pore clusters, cracks (longitudinal, transverse, crater), lack of fusion, incomplete root penetration, slag inclusions, undercut, excess reinforcement, spatter and burn-through. Decide early whether you need class-level labels, bounding boxes or masks. For segmentation, COCO JSON with categories, images and polygon or RLE annotations is widely supported in tools such as CVAT [4].

Modality and capture metadata that make images trainable

Every image needs capture metadata, or you cannot stratify by the conditions that cause domain shift. For visual images that means camera, lens, lighting, distance and whether the weld was cleaned or ground. For radiographs it means film versus computed or digital radiography, source type, exposure technique, image quality indicator reading, bit depth and file format (16-bit TIFF and DICONDE are common in NDT archives).

Also record base material, thickness, joint type (butt, fillet, lap), welding process (GMAW, GTAW, SMAW, SAW, laser) and position. Our guide to camera, lens and lighting diversity covers how to test whether capture variety is real, and EXIF metadata handling covers which fields to keep.

Request specification for weld defect images

A precise request spec lets suppliers confirm quickly whether they hold matching records.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample specificationWhy it matters
ModalityDigital radiography, 16-bit TIFF, single-wall exposureModels do not transfer across modalities
Material and thicknessCarbon steel pipe, 6-19 mm wallContrast and defect appearance change with thickness
Joint and processGirth butt welds, GMAW root plus FCAW fillDefect frequencies differ by process
TaxonomyISO 6520-1 classes, mapped from inspector codesComparable labels across sources
Label typeBounding boxes for all indications; masks for cracks and porosityMatches detection and segmentation heads
Acceptance contextGoverning standard and level per image, accept/reject decisionSeparates "defect present" from "rejectable"
OutcomeRepair performed, re-inspection resultTies visual evidence to real consequences
Class balance targetMinimum count per defect class; all clean images retainedPrevents a 96%-clean set from training a "pass" model
ProvenanceAsset owner, inspection company, consent to licenseConfirms who can grant rights

Pairing each image with its accept/reject decision and repair outcome, where available, is often what separates a useful dataset from a pile of labeled pictures.

Plan for class imbalance and label audits

Expect defective images to be a small minority, so plan coverage per class rather than total volume. With about 4% defective in one real welding dataset [2], a random sample of 10,000 images might yield only a few dozen examples of rare classes such as transverse cracks. Ask suppliers for per-class counts before pricing, keep clean images for false-positive calibration, and see rare defect coverage for pooling and synthetic options.

Audit labels with a sampling plan rather than spot checks. Attribute sampling schemes such as ANSI/ASQ Z1.4 define AQL-based normal, tightened and reduced plans [5]. The standard was not written for labels, but the same logic can be adapted to annotation audits by having a qualified inspector re-read a sample per lot. Track disagreement by class, because porosity versus slag and lack of fusion versus incomplete penetration are common confusions even among human readers.

Rights and chain of title for inspection images

Inspection images often belong to the asset owner, not to the inspection contractor who captured them, so confirm chain of title before you evaluate quality. A third-party NDT firm may hold thousands of radiographs that its service agreements do not let it license. Ask who commissioned the inspection, what the contract says about data reuse, and whether images show customer markings, serial numbers or site identifiers that need redaction.

Check public data terms just as carefully: research-only or non-commercial terms on public weld collections generally rule out training a commercial model. For workers who appear in visual weld photos, see faces in image training data.

How SourceX sources weld inspection data

SourceX sources operational datasets from US companies on request, including documents and new recordings of hands-on work, and manages licensing and ongoing purchases. Nothing is held in stock and a request does not guarantee a match. Buyers describe the data, not the businesses, and SourceX looks for US businesses that hold it. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, through private, access-controlled workflows after an executed agreement and supplier approval.

Personal details such as names and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. Related owner pages cover inspection photos for AI training, inspection reports and manufacturing quality records, which often hold the accept/reject decisions that make weld images useful. You can submit a weld image data request using the spec above.

More image sourcing guides are in the image data hub, alongside corrosion detection image datasets for asset integrity programs.

Request weld defect image data

Describe the modality, materials, joint types, defect classes and acceptance context you need. SourceX looks for US businesses that hold matching data, prepares diligence materials per dataset, and nothing is contracted until a supplier agrees. Describe the weld inspection images you need.

Sources

  1. IEEE DataPort, "WELD dataset". https://ieee-dataport.org/documents/weld
  2. Alexandria Engineering Journal via DOAJ, "Deep learning defect detection on copper filter drier welding images" (2025). https://doaj.org/article/a7f91d9ea879427fa513ad25ff017d0e
  3. arXiv, "A Review of Benchmarks for Visual Defect Detection in the Manufacturing Industry" (2023). https://arxiv.org/pdf/2305.13261
  4. CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
  5. ASQ, "ASQ/ANSI Z1.4:2003 (R2018): Sampling Procedures and Tables for Inspection by Attributes". https://asq.org/quality-press/display-item?item=T1164

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data