Skip to content

Image data

Construction Safety Image Datasets for PPE and Hazard Detection

Quick answer

A usable construction PPE detection dataset needs more than boxes around hard hats. Buyers should specify a class taxonomy that includes negative classes (no hard hat, no vest), capture diversity across sites, trades, cameras, lighting and weather, documented label QA, and a license that covers commercial training. Community sets are a fast baseline, but they vary in license and annotation quality, and identifiable workers bring employee-notice and biometric-law questions that the supplier must answer before delivery.

By SourceX Editorial · Updated

What public construction PPE datasets cover, and where they stop

Public sets are good for a first detector and weak for production. Roboflow Universe hosts several community "construction site safety" projects with classes such as hardhat, safety vest and person, each carrying its own license that you must read per project [1]. One academic comparison of YOLOv8, EfficientDet and MobileNet-SSD for PPE detection used a Kaggle set whose classes include hard hats, masks, vests, machinery and cones [2]. That class list is narrow for real jobsites, where the failures that matter are missing harnesses at height, unguarded edges and people inside equipment swing radius.

Capture conditions are the bigger gap. One published personnel dataset was collected from four static cameras on a single site [3], which is honest about scope but shows how easily a set encodes one camera height, one background and one season. A model trained on that distribution will look strong on a random split and degrade when moved to a new site.

Common failure modes to plan for:

  • Mixed or unclear licensing. Forks and re-uploads on community hubs can strip provenance, so the license on the page may not match the origin of the frames [1].
  • Positive-only labels. Many sets label "hardhat" but not "head without hardhat", so violation recall cannot be measured.
  • Color shortcuts. Models learn "yellow blob on head" and fire on yellow buckets or hair; hi-vis vests are confused with orange cones and barrier mesh.
  • Near-duplicate frames. Video-derived sets leak adjacent frames across train and test splits, inflating mAP.

Define the label set before you source images

The label taxonomy decides what your model can report, so write it first and send it with every request. Pair each PPE item with an explicit absence class, and separate PPE compliance from scene hazards, because they are labeled at different granularity (person-attached boxes versus scene regions or relations). Keep progress-monitoring labels such as installed elements or percent complete out of this set; they belong to construction progress photo datasets.

Illustrative example: invented to show structure; it does not describe an available dataset.

GroupClassGeometryAttribute fieldsCommon confusions
Head PPEhard_hat, head_no_hard_hatboxcolor, worn_correctly (true/false)caps, hoods, buckets
Torso PPEhi_vis_vest, torso_no_vestboxANSI-style class if visible, worn_openorange mesh, cones
Fall protectionharness, lanyard_attachedbox, keypointstie_off_visiblebackpacks, tool belts
Eye/face/handsafety_glasses, gloves, face_shieldboxoccludedsunglasses
Person contextworkerbox, keypointstrade (if known), elevation_mmannequins, posters
Equipmentexcavator, crane, forklift, skid_steerbox or polygonstate (moving/idle)parked vehicles
Scene hazardunprotected_edge, open_excavation, ladder_misusepolygon or relationseverityshadows, tarps
Site controlcone, barricade, guardrailboxcomplete (true/false)debris

Export in COCO JSON so categories, images and annotations stay in one documented structure that tools such as CVAT can import and export [4]. Put capture metadata (site ID, camera ID, timestamp, weather, lighting) in the image records rather than file names, because split logic depends on it.

Specify site, camera and condition diversity explicitly

Diversity has to be written as quotas, or a supplier will deliver whatever their cameras captured. Ask for images from multiple independent sites and contractors, mixed project types (vertical building, civil, interior fit-out, utilities), and at least three capture modes: fixed pole or crane cameras, handheld or helmet-mounted phones, and drone oblique shots. Each mode changes the apparent size of a hard hat several-fold, and small, distant heads are a common source of missed detections.

Condition coverage should include dawn and dusk, artificial night lighting, rain, dust and heavy occlusion by rebar or scaffolding. Our guide to camera, lens and lighting diversity covers how to state these as metadata fields and per-bucket minimums. Then hold out whole sites, not random frames, for evaluation; a site-level split is the only honest test of the generalization problem shown by single-site sets [3].

Audit label quality before you trust a benchmark number

Label error rates in PPE sets are rarely published, so measure them yourself on a sample. Even standard benchmark test sets carry label errors large enough to reorder model rankings [5], and small community sets have less review. Request the annotation guideline, the reviewer workflow and agreement statistics, and check whether the supplier used consensus labeling or seeded gold items to score annotators [6].

A practical acceptance test: draw a few hundred images stratified by site and camera, have your own reviewer relabel absence classes only, and compare. Absence classes (head without hard hat) are where tired annotators skip boxes, and they are also the classes your safety team will act on. The broader method is in our training data quality assessment guide, and the build-versus-buy tradeoff for labels is covered in pre-labeled versus raw image datasets.

Pair images with safety records where you can

The highest-value construction safety data links frames to what happened next. Daily safety observations, toolbox-talk notes, near-miss logs and incident reports let you label not just "vest missing" but "observation raised, corrected within shift", which supports ranking alerts by consequence rather than by detector confidence. These records sit in contractor systems alongside the construction project records used for document and scheduling models.

Ask suppliers whether each image can carry an observation ID, a hazard category code from their own safety program, and an outcome field. Inspection notes written by safety managers can also become domain captions; see domain captions from work records. Expect free-text records to need the same personal-data review as images before release.

Handle worker privacy, notice and biometrics up front

Jobsite images show identifiable employees, so privacy review is part of the sourcing spec, not a cleanup step. Several US states require employers to give written notice before electronically monitoring employees, including New York [7] and Connecticut [8]; ask the supplier how cameras that produced the images were disclosed to workers. Notice for site safety monitoring does not automatically cover licensing frames for third-party AI training, so confirm the supplier's rights to release them.

If any pipeline step computes face geometry, the data may fall under Illinois BIPA, where a 2024 amendment treats repeated collection of the same identifier from the same person by the same method as a single violation [9]. PPE detection rarely needs identity, so the simpler path is to blur or mask faces, name badges, hard-hat stickers with names and vehicle plates before delivery. Our guides to face data consent and anonymization and biometric data rules for buyers cover the tradeoffs, including the accuracy cost of blurring heads in a head-PPE task.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Buyer request checklist for construction safety images

A complete request lets a supplier say yes or no quickly and lets your reviewers approve the license without rework.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Task: PPE compliance detection plus scene-hazard segmentation; intended use (training, evaluation or both).
  • Classes: the taxonomy table above, with absence classes and attribute fields.
  • Diversity quotas: number of distinct sites, project types, capture modes, night and weather share, US regions.
  • Linked records: observation IDs, hazard codes and outcomes where available.
  • Format: COCO JSON with capture metadata per image; original resolution; EXIF policy (see EXIF metadata in image training data).
  • Privacy: face, badge and plate masking method, written description of method and the sample check performed.
  • Rights: confirmation the company owns the footage, how workers were notified, and permitted uses, term and delivery written into the license.
  • Splits: site-level holdout IDs supplied with the data.

How SourceX sources construction safety image data

SourceX sources operational datasets from US companies on request and manages the licensing process, including recordings of hands-on work, and every release is approved by the supplying company. Categories are not stock, so a request does not guarantee a match, and SourceX does not source scraped web content or generic CCTV. Each dataset is rights-reviewed, personal details are removed or replaced with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows after an executed agreement. Construction teams can review the construction buyer overview or start from the image data hub, and you can describe the data you need.

Request construction PPE and hazard detection data

Describe the classes, site diversity and linked safety records you need, and SourceX will look for US businesses that hold that data and can approve its release. Nothing is contracted until a supplier agrees, and terms are agreed per deal. Start a construction safety data request.

Sources

  1. Roboflow Universe, "Construction Site Safety Computer Vision Dataset". https://universe.roboflow.com/trafik-nesneleri/construction-site-safety-eipvn
  2. National College of Ireland (NORMA repository), "PPE detection on construction sites: comparison of YOLOv8, EfficientDet and MobileNet-SSD (MSc thesis)". https://norma.ncirl.ie/9777/1/mohitgummarajkishore.pdf
  3. PubMed Central (NCBI), "Manually classified dataset of leaning and standing personnel images for construction site monitoring". https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11993151/
  4. CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
  5. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. Dataloop, "Creating Consensus, Honeypot, and Qualification Tasks". https://developers.dataloop.ai/tutorials/task_workflows/quality_control/chapter
  7. New York Laws (public.law), "N.Y. Civil Rights Law Section 52-C". https://newyork.public.law/laws/n.y._civil_rights_law_section_52-c*2
  8. Connecticut Department of Labor, "Electronic Monitoring of Employees (Conn. Gen. Stat. 31-48d notice)". https://portal.ct.gov/dol/-/media/DOL/2022-New-Design-System/Divisions/wage-and-workplace-standards/ElectronicMonitoring.pdf
  9. Illinois General Assembly, "SB 2979 (103rd General Assembly) - AN ACT concerning civil law (BIPA amendment), engrossed text" (2024). https://www.ilga.gov/documents/legislation/103/SB/PDF/10300SB2979eng.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data