Skip to content

Image data

Concrete Crack and Civil Infrastructure Defect Image Datasets

Quick answer

A useful concrete crack dataset for production models pairs varied field imagery of bridges, decks, pillars, tunnels and retaining walls with pixel-level masks for cracks, spalling, exposed rebar, efflorescence and rust staining. It should also carry capture metadata and, ideally, the inspection record each photo came from. Public sets such as CODEBRIM are good for benchmarking. Commercial-grade training data usually comes from inspection consultancies' and asset owners' archives, so ownership and licensing need checking before labels.

By SourceX Editorial · Updated

What separates civil defect imagery from factory defect data

Civil infrastructure defect images are field data, not line-scan data. Factory sets covered in our industrial defect image datasets guide are typically shot under controlled lighting at a fixed distance. Bridge and structure imagery comes from handheld cameras, rope-access crews, snooper trucks and drones, across seasons, wet and dry surfaces, shadows under decks and very different standoff distances.

That variability is the point, not noise. The authors of a 2026 dataset built to challenge vision foundation models deliberately kept differences in resolution, viewpoint, illumination and weather so it would test generalization [2]. Public concrete defect benchmarks such as CODEBRIM were also captured with multiple cameras, at varying scales and in changing weather, partly by UAV where bridge elements are hard to reach.

The defect taxonomy also differs. Concrete defect classes usually include crack, spallation, exposed reinforcement bar, efflorescence and corrosion stains, and they are not mutually exclusive: one image can show a crack, rust bleed and spalling at once. Models that assume one label per tile can underperform on real decks, where defects co-occur along the same joint or drainage path.

Label formats that matter for crack segmentation

Crack segmentation needs pixel masks, not boxes. Cracks are thin, long and often only a few pixels wide at survey resolution, so bounding boxes around them are mostly background. Ask for one of these, and know which you are getting:

  • Binary crack masks (PNG, one channel) for crack-versus-background segmentation. Check the stated mask width policy: some annotators trace the visible crack, others dilate it to a fixed width, which changes IoU.
  • Multi-class semantic masks for crack, spall, exposed rebar, efflorescence, rust stain and sealed or repaired areas, with an explicit "repaired" class so patched cracks are not counted as live defects.
  • Instance polygons (COCO JSON) when the model must count spalls or measure individual defect areas.
  • Image-level multi-label tags, as in CODEBRIM, which are cheaper and adequate for triage classifiers but not for measurement.

For crack width estimation, a mask is not enough. You also need a scale reference: ground sample distance from drone altitude and sensor specs, a target or tape in frame, or a laser rangefinder value. Without it, a 0.3 mm hairline and a 3 mm structural crack can look identical at the pixel level. Our guide to pre-labeled versus raw image datasets covers when to buy masks and when to buy raw imagery and annotate it in-house.

Class imbalance in drone bridge surveys

Most images in a structure survey contain no defect at all. One drone campaign over concrete bridge pillars produced more than 22,000 high-resolution images, with cracks and rust a small minority, and the authors used active learning to cope with the imbalance [1]. A raw survey dump is therefore a poor training set until it has been triaged.

Ask a supplier for the defect-positive rate per structure and per class, not just total image count. Spalling with exposed rebar and active corrosion are far rarer than hairline cracking, and efflorescence is easy to confuse with bird droppings, paint and water staining. Defect-free images still have value as hard negatives and for normal-only anomaly detection sets. Strategies for very rare classes, including pooling and synthesis, are covered in rare defect coverage and class imbalance.

Linking images to inspection records and element condition

The highest-value civil defect data ties each photo to the structure, element and condition rating an inspector recorded. In the US, public highway bridges are inspected under the National Bridge Inspection Standards (23 CFR 650 Subpart C), with inventory and condition data reported to FHWA for the National Bridge Inventory. Many agencies also record element-level condition states following the AASHTO Manual for Bridge Element Inspection.

That means inspection photos often sit next to structured fields: structure number, element (deck, girder, pier cap, bearing, joint), condition state quantities and inspector notes. A photo labeled "pier 3, cap, condition state 3 spall with exposed rebar" is weak supervision you can use for retrieval, captioning or evaluation. See domain captions from work records for using inspection notes as image text, and the inspection reports page for the reports themselves.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "image_id": "img_000418",
  "file": "pier03_cap_east_0418.jpg",
  "capture": {"platform": "uav", "camera": "1-inch sensor, 20 MP", "gsd_mm_per_px": 0.8, "date": "2025-06-14", "surface": "wet"},
  "structure": {"structure_type": "concrete girder bridge", "structure_id": "hashed", "element": "pier cap", "span_or_pier": "pier 3"},
  "inspection": {"condition_state": 3, "inspector_note": "spall approx 0.3 m2 with exposed rebar, rust staining below"},
  "labels": {"mask": "masks/img_000418.png", "classes": ["spall", "exposed_rebar", "rust_stain", "crack"], "annotation_tool": "polygon, reviewed"},
  "exif_gps": "removed; county-level location retained"
}

Who owns bridge inspection photos

Inspection photos frequently belong to the asset owner, not the firm that took them. A state DOT, municipality, port authority or private owner commissions the inspection, and the consultant's contract often assigns photos and reports to that owner; verify this deal by deal. Before licensing a consultancy archive, ask who owns the deliverables, whether the contract allows reuse, and whether any owner restricts release of critical infrastructure details.

Location is the other sensitive field. GPS in EXIF, structure numbers and recognizable landmarks can identify a specific bridge, which may matter for security-sensitive assets. Decide what to keep: county or climate zone is usually enough for generalization testing. Our guide on EXIF metadata in image training data covers what to strip. Photos taken from roads and walkways can include passers-by and vehicle plates, so check the face data consent and anonymization guide as well.

Buyer checklist for concrete defect image data

Use this checklist before you sign for a civil infrastructure defect set.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat to askWhy it matters
Structure and element mixBridges, culverts, tunnels, parking decks, walls; deck, soffit, pier, abutmentModels trained on decks miss soffit and pier defects
Capture platformsHandheld, snooper, rope access, UAV; sensor and GSDDetermines crack width visibility and domain gap
ConditionsWet and dry, shadow, season, efflorescence-stained surfacesGeneralization, as shown in recent benchmarks [2]
Label type and policyMasks vs boxes vs tags; crack width policy; repaired classChanges IoU and measurement accuracy
Defect-positive ratePer class, per structureSurveys are mostly defect-free [1]
Inspection linkageElement, condition state, notesEnables evaluation against inspector judgment
OwnershipOwner vs consultant; contract reuse termsControls who can license the images
Location handlingEXIF GPS, structure IDs, landmarksSecurity and privacy of specific assets
Split designHold out whole structures, not random tilesTiles from one pier leak across splits

The last row is a common failure mode: random tile splits put near-duplicate crops of the same crack in train and test, inflating scores. Hold out entire structures, and ideally entire regions or capture campaigns, for evaluation.

How SourceX helps with infrastructure defect imagery

SourceX sources operational datasets from US companies on request, including documents, engineering records and new recordings of hands-on work, and manages licensing. You describe the imagery, labels and linked records you need, not the businesses; SourceX looks for US companies that hold that data, and every release is approved by the supplying company. Nothing is held in stock and a request does not guarantee a match. Describe your crack and defect image requirements to start.

Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows under a license defining records, uses, term and delivery. Engineering firms can also see how engineering consultancies work with SourceX and the images and inspection photos page. For the broader category, start at the image data hub, compare corrosion detection image datasets, or return to the AI data hub.

Request concrete crack and bridge defect images

SourceX looks for US businesses that hold the defect imagery and inspection records you describe, then runs the Find, Assess, Agree, Transact and Manage process; nothing is contracted until a supplier agrees. Each dataset is rights-reviewed and delivered under a license that sets records, uses, term and delivery. Tell SourceX what concrete defect data you need.

Sources

  1. arXiv, "Active Learning for Imbalanced Civil Infrastructure Data" (2022). https://arxiv.org/pdf/2210.10586
  2. arXiv, "Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models" (2026). https://arxiv.org/pdf/2605.18413

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data