Skip to content

Image data

Vehicle Damage Photo Datasets for Claims and Repair AI

Quick answer

A usable vehicle damage dataset for production claims or repair models is not a folder of crash pictures. It is a set of real first-notice-of-loss, appraiser and repair-shop photos with per-image part and damage-type labels, severity, and a link to what actually happened on the claim: repaired, supplemented or declared a total loss. Public sets are small and often scraped, so serious teams license operational photo archives from insurers, appraisal firms and collision repairers, with plates, VINs, faces and GPS removed before delivery.

By SourceX Editorial · Updated

Why public car damage datasets stall at the prototype stage

Public car damage datasets are good for proving a pipeline and poor for training a claims-grade model. A widely used tutorial set holds about 1,610 images at 480x640, labeled only "damaged" or "whole" [3], which cannot teach a model to separate a scuffed bumper cover from a cracked absorber behind it. Academic work confirms the gap: one thesis found the available public data insufficient and assembled its own set from web search results [1], which leaves the rights status of every image unclear.

Segmentation-grade research sets do exist. VehiDE was published in 2024 specifically for damage detection and segmentation [2], and it is a useful benchmark for architecture choices. But research licenses, stock-photo framing, and the absence of claim outcomes limit how far these sets carry a commercial model. Commercial labeling vendors also market vehicle damage data to insurers [6]; treat their scale claims as self-reported until you see a sample and the underlying rights chain.

The recurring failure modes when teams move from public data to live claims photos:

  • Domain shift in capture. Web images are well-lit, centered and often staged; policyholder phone photos are taken at night, in parking garages, through rain, with fingers in frame.
  • Label granularity. Binary or three-class labels cannot drive repair-or-replace decisions or parts estimates.
  • No ground truth on cost. Without the final estimate or total-loss decision, severity labels are an annotator's guess.
  • Rights uncertainty. Scraped images carry no license for training or commercial deployment.

Which capture source fits which model

Capture source determines what your model will see in production, so specify it before anything else. The same dent looks different in a customer's guided-capture app, an independent appraiser's walkaround and a body shop's teardown photos after the bumper cover is removed.

Capture sourceTypical images per claimStrengthsGaps to plan forBest fit
Policyholder self-service (FNOL app, photo estimate)Few, uneven anglesMatches straight-through-processing inputs; real lighting and device diversityBlur, glare, missing angles, reflections mistaken for damageTriage, photo estimating, fraud flags
Staff or independent appraiser walkaroundStructured set, all four corners plus close-upsConsistent angles, odometer and VIN shots, linked to estimate linesFewer extreme low-quality images; may over-represent drivable vehiclesDamage segmentation, estimate line prediction
Repair shop intake and teardownMany, including disassembled panelsHidden damage visible; supplements explain what photos missedShop-specific backgrounds; parts photos lack whole-vehicle contextHidden-damage prediction, supplement forecasting
Salvage and total-loss yardWhole-vehicle setsSevere and structural damage well representedSurvivorship bias toward totals; staged lotsTotal-loss classifiers

For most claims automation programs, a blend of policyholder and appraiser photos with a smaller teardown slice gives the best coverage. Our guide to camera, lens and lighting diversity in image datasets covers how to set device and condition quotas, and rare defect coverage for imbalanced inspection data applies directly to uncommon damage such as airbag deployment or frame rail buckling.

The label schema that makes damage photos trainable

A claims-grade schema labels the part, the damage type, the extent and the outcome, at both image and claim level. Image-level boxes or polygons tell the model where; claim-level fields tell it what the damage cost and whether the estimate held.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "claim_ref": "hashed-7f3a9c",
  "vehicle": { "body_style": "midsize_sedan", "model_year_band": "2019-2021", "paint_finish": "metallic" },
  "loss": { "type": "rear_end_collision", "point_of_impact": "rear_center", "loss_month": "2025-11" },
  "image": {
    "image_id": "img-000412",
    "capture_source": "policyholder_app",
    "view": "rear_left_45",
    "quality_flags": ["glare", "wet_surface"],
    "exif_stripped": true,
    "plate_redacted": true
  },
  "annotations": [
    { "part": "rear_bumper_cover", "damage_type": "crack", "severity": "replace", "polygon": "..." },
    { "part": "tail_lamp_left", "damage_type": "broken_glass", "severity": "replace", "polygon": "..." },
    { "part": "trunk_lid", "damage_type": "dent", "severity": "repair", "polygon": "..." }
  ],
  "outcome": {
    "disposition": "repaired",
    "initial_estimate_band": "B",
    "supplement_count": 1,
    "hidden_damage_found": ["rear_bumper_reinforcement"],
    "estimate_parts_lines": 7
  }
}

Agree a closed part taxonomy (bumper cover, fender, quarter panel, door shell, hood, headlamp and so on) and a damage-type list (dent, scratch, crack, broken glass, tear, misalignment) before annotation starts. If you plan to annotate in-house, compare costs in our note on pre-labeled versus raw image datasets. Adjuster comments and estimate line descriptions can also become captions; see domain captions from work records.

Linking photos to claim outcomes

Outcome linkage is what separates an operational dataset from a picture collection. A photo labeled "severe" by an annotator is an opinion; a photo whose claim ended as a total loss, or whose estimate later needed two supplements, is ground truth your model can be scored against.

Ask the supplier which outcome fields survive in their claim or shop management records: disposition (repair, total loss, cash settlement), supplement history, final parts list, and whether hidden damage was found at teardown. Claim files in the US are shaped by state unfair claims practices rules; the NAIC model regulation includes a section on file and record documentation [7], so ask how long the supplier retains claim files and whether photos are stored with the estimate. If you need the full claim record rather than images, the insurance claims workflow datasets page covers that scope, and claims adjudication decisions with coverage reasoning covers decision data for agents.

Watch for leakage when outcome fields ride along. If the photo set includes the appraiser's final estimate screenshot or a "TOTAL LOSS" placard on the windshield, the model learns the label from the image text, not the damage.

Privacy and rights checks before you license car damage images

Vehicle damage photos carry more personal data than they appear to. License plates, VIN plates and door-jamb stickers, faces of drivers and bystanders, house numbers in driveways, registration and insurance cards on dashboards, and EXIF GPS coordinates can all tie an image back to a person and an address.

EXIF is easy to forget. An audit of a large web-scraped training set found Exif tags with timestamps, geolocation and personal details, and noted that download tooling carried that metadata forward [5]. Decide which fields to keep (capture device model can be useful for domain analysis) and which to strip; our guide to EXIF metadata in image training data sets out a field list.

Redaction can affect accuracy. Research on anonymizing people in training images compared traditional blurring with realistic anonymization and measured the effect on downstream vision tasks [4]; see whether face and license plate blurring hurts training for a practical summary. For damage models, plates and faces rarely overlap the damage region, so targeted redaction usually costs little, but test it on your own validation set.

On rights, ask who owns the photos. Policyholder uploads, appraiser photos and shop photos may sit under different terms of service, vendor contracts or estimating-platform agreements, and the holder must have the right to license them for model training.

Buyer checklist for a vehicle damage dataset request

A good request describes the data you need in enough detail that a holder can say yes or no quickly.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldWhat to specify
Capture sourcePolicyholder app, appraiser, repair shop, salvage, or a target mix
Vehicle scopePassenger cars, light trucks, EVs; model-year range; exclusions (motorcycles, commercial fleet)
Loss typesCollision, hail, vandalism, glass-only, flood, fire
Labels neededPart taxonomy, damage types, severity, polygons or boxes, or raw for in-house annotation
Outcome fieldsDisposition, supplement count, final parts lines, hidden damage flags
Image requirementsMinimum resolution, views per claim, formats (JPEG, HEIC), quality-flag tolerance
Privacy handlingPlate, VIN, face and document redaction; EXIF policy; claim ID hashing
Volume and cadenceInitial backfile plus whether you want ongoing monthly deliveries
Intended useTraining, evaluation holdout, or both; internal or deployed product

Related image work for property claims is covered in property inspection photos linked to findings and roof condition and hail damage imagery. The image data hub lists the full cluster, and the AI data guides cover other modalities, including service, warranty and claims videos.

How SourceX sources vehicle damage photos

SourceX sources operational datasets from US companies on request; vehicle damage photos are not held in stock, and a request does not guarantee a match. You describe the photos, labels and outcome fields you need, and SourceX looks for US businesses that hold that data. Every release is approved by the supplying company, and SourceX does not source scraped web content or generic CCTV or photos.

The process runs Find, Assess (the data and its licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Each dataset is rights-reviewed for ownership and consents, personal details such as names, phone numbers and account numbers are removed or replaced before delivery with the method recorded and a sample checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement. Insurance teams can also review buyers in insurance, buyers in claims administration and how to license images and inspection photos, then start a buyer request.

Request vehicle damage photo data

Describe the capture source, labels and claim outcome fields your damage model needs, and SourceX will look for US companies that hold matching photos and manage licensing if a supplier agrees. SourceX serves AI teams wherever they are based and agrees terms per deal. Describe your vehicle damage dataset request.

Sources

  1. National College of Ireland (NORMA repository), "Car damage detection thesis (MSc, National College of Ireland)". https://norma.ncirl.ie/6101/1/shubhamsarjeraochaudhari.pdf
  2. Journal of Information and Telecommunication (via DOAJ), "VehiDE: Vehicle Damage Detection dataset" (2024). https://doaj.org/article/6e0eb600059440b390431a42dfe8d138
  3. Labellerr, "ML Beginner's Guide to Build Car Damage Detection AI Model". https://labellerr.com/blog/ml-beginners-guide-to-build-car-damage-detection-ai-model
  4. Hukkelas et al., CVF Open Access, "Does Image Anonymization Impact Computer Vision Training? (CVPR 2023 Workshops)" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
  5. arXiv, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
  6. Shaip, "Automotive Insurance: Vehicle Damage Assessment". https://shaip.com/solutions/vehicle-damage-assessment
  7. National Association of Insurance Commissioners, "Unfair Property/Casualty Claims Settlement Practices Model Regulation (Model 902)". https://content.naic.org:443/sites/default/files/model-law-902.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data