Image data
Property Inspection Photos Linked to Findings for Underwriting and Claims Models
Quick answer
A useful property inspection photos dataset is not a folder of house pictures. It is ground-level interior and exterior photos where each image is joined to the inspector's finding: the component, the condition or hazard code, severity, and the recommended action. Buyers should source these pairs from inspection firms, insurers' inspection vendors or adjusting operations, insist on finding-level linkage rather than report-level linkage, define in-home redaction rules up front, and separate underwriting (hazard) inspections from claims (loss) inspections.
By SourceX Editorial · Updated
This page covers the photo-plus-finding pair. If you mainly need the narrative report text, see inspection report data for AI training; for the wider image category, start at the image data hub.
Why photo-finding pairs matter more than raw property images
Paired data matters because the finding is the label, and an unlabeled inspection photo teaches a model almost nothing about condition. Research on roof assessment frames inspection as a slow, costly and sometimes hazardous human task that insurers want to automate from imagery [1]. That automation only works when the training signal says what the inspector concluded about the specific thing in the frame.
Three model families consume these pairs. Condition and hazard classifiers (for example, "double-tapped breaker", "missing GFCI near sink", "active roof leak staining") need image-level or region-level labels. Vision-language models fine-tuned to draft findings need the finding text as the target caption, which is covered in more depth in using inspection notes and adjuster comments as image text. Underwriting automation models need the finding rolled up to a property-level decision, such as refer, require repair or accept.
Overhead imagery is a different product. If your target is roof geometry or hail from aerial or drone captures, the roof condition and hail damage imagery guide is the better fit; ground-level photos see what aircraft cannot, such as the electrical panel, water heater venting and under-sink plumbing.
Underwriting inspections versus claims inspections
Underwriting and claims inspections produce different labels, so treat them as two datasets even when one vendor holds both. Mixing them silently corrupts the target variable.
- Underwriting inspections (new business or renewal) look for hazards and maintenance: roof age and covering, wiring type (knob-and-tube, aluminum branch circuits), panel brand, trampoline or pool fencing, wood stoves, handrails. Findings are usually hazard codes with a required-repair flag and due date.
- Claims inspections document a loss: water intrusion source, mold extent, fire and smoke zones, wind or hail impact. Findings tie to a cause-of-loss code, affected rooms and an estimate. For estimate line items and adjuster field notes, see property claim estimates and adjuster field reports.
- Pre-purchase home inspections follow a standards-of-practice structure by system (roof, exterior, structure, electrical, plumbing, HVAC, interior) and tend to have richer narrative but less consistent severity scales.
A claims-heavy dataset over-represents damage; an underwriting dataset over-represents clean homes with a handful of hazards. Know which base rate you are buying.
What a usable photo-finding record looks like
A usable record links one finding to one or more photos by ID, carries structured codes alongside the free text, and keeps capture context without exposing the household. Most source systems store photos as attachments on a report section, so finding-level linkage often has to be reconstructed by the supplier, which is the first thing to verify in a sample.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"inspection_id": "INSP-7f3a",
"inspection_type": "underwriting_renewal",
"property": {"region": "US-South", "year_built_band": "1970-1979", "dwelling_type": "single_family"},
"finding": {
"finding_id": "F-012",
"system": "electrical",
"component": "main_panel",
"code": "ELEC-PANEL-DOUBLE-TAP",
"severity": "repair_required",
"action": "licensed_electrician",
"text": "Two conductors under single breaker terminal, lower left."
},
"photos": [
{"photo_id": "P-0441", "role": "context", "room": "garage", "redacted": false},
{"photo_id": "P-0442", "role": "detail", "room": "garage", "redacted": true,
"regions": [{"bbox": [812, 604, 140, 96], "label": "double_tap"}]}
],
"capture": {"device_class": "phone", "exif_gps_removed": true, "capture_month": "2025-04"},
"inspector": {"inspector_hash": "a91c", "certification_level": "senior"}
}
For annotation export, COCO JSON is the common interchange: its info, licenses, categories, images and annotations sections map cleanly to photo IDs, finding codes and bounding boxes, and tools such as CVAT import and export it [6]. For large deliveries, WebDataset tar shards (often around 1 GB each) keep each image beside its JSON sidecar and stream well at training time [7].
In-home privacy: what interior photos expose and how to redact it
Interior photos are a privacy problem distinct from report text, because a camera inside a home captures far more than the finding. Define redaction rules per object class before any sample is pulled, and record which method was applied.
What typically appears in frame:
- People: occupants, children, the inspector's reflection in mirrors and appliance doors.
- Documents: mail, bills, prescription bottles, calendars and whiteboards with names or schedules.
- Identity and location cues: house numbers, license plates in driveways, street signs, mailbox names, EXIF GPS coordinates. See what EXIF metadata to strip and keep.
- Sensitive possessions and status: family photos, religious items, firearms, medical equipment, safes and valuables.
Faces are the highest-risk class. Illinois BIPA regulates collection, retention and disclosure of biometric identifiers [3], so any pipeline that runs face geometry on interior photos, including for blurring, needs a legal review of whether it creates biometric data. Under California's CCPA, data only counts as deidentified if the holder meets specific conditions, including reasonable measures against re-identification [4]. Memorization is also real: diffusion models have been shown to regenerate individual training images, including photos of people [5], which argues for redaction before training rather than relying on output filters.
Practical rules buyers can ask for: blur or mask faces and full bodies, documents and screens, house numbers and plates; strip EXIF GPS and device serials; coarsen capture date to month; replace addresses with region bands; and drop images where redaction would remove the finding itself.
Label quality, coverage and bias in inspection findings
Inspection labels are noisy because inspectors differ in thresholds, templates change, and findings reflect what the inspector chose to photograph. Audit before you train.
| Check | What to measure | Failure mode it catches |
|---|---|---|
| Inter-inspector agreement | Have 2-3 reviewers relabel a stratified sample of photos against the code list | Severity drift between senior and junior inspectors |
| Template versioning | Map every finding code across template versions to one taxonomy | Same defect under three codes after a software migration |
| Photo-to-finding linkage | Share of findings with at least one detail photo | Report-level attachments mislabeled as finding evidence |
| Negative coverage | Share of "inspected, no defect" photos per component | Classifier that never learns what a compliant panel looks like |
| Geography and housing stock | Region, construction era, dwelling type mix | Model trained on Sun Belt slab homes failing on basements |
| Outcome contamination | Whether labels were edited after claim or cancellation | Hindsight bias leaking into hazard labels |
Underwriting outcomes such as non-renewal can encode past decisions rather than true risk; the guide to historical decision bias in operational labels covers that audit. Ask suppliers for a written data statement describing who inspected, under which template, in which period, so reviewers can judge generalization [8]. Rare hazards (active arcing, structural cracking) will be thin; the rare defect coverage guide covers pooling and synthetic augmentation.
Rights, consent and insurance-regulator expectations
Rights in inspection photos are layered: the homeowner owns the property, the inspector or firm usually owns the photos, and the client (buyer, agent or insurer) may hold contractual use rights. Confirm the chain before evaluating samples.
Questions for counsel and the supplier:
- Do the inspection agreement and privacy notice permit use of photos and findings for model development, or only for the report?
- Did the homeowner or occupant consent to interior photography, and on what terms?
- Are any photos of third-party artwork, logos or identifiable people that need releases? See model and property releases for AI training images.
- If an insurer will deploy the model, can the data's provenance satisfy the NAIC Model Bulletin, a guidance document (not a model law) that a growing number of states have adopted as of October 2026 and that expects documented governance over AI systems and third-party data [2]? The NAIC bulletin guide for insurers goes deeper.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Sample request template for inspection photo-finding data
A precise request gets better matches than a category name. Describe the data, not the companies you think hold it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example entry |
|---|---|
| Inspection type | Underwriting renewal and four-point inspections; exclude claims |
| Geography and stock | US single-family, built before 1990, mixed regions |
| Systems | Roof, electrical panel, plumbing supply, water heater, HVAC |
| Linkage required | Each finding linked to at least one detail photo by ID |
| Labels | Finding code, severity, required-repair flag, free-text finding |
| Negatives | At least one "no defect" photo per component inspected |
| Privacy | Faces, documents, house numbers blurred; EXIF GPS removed; dates coarsened to month |
| Format | JPEG plus COCO JSON, or WebDataset shards with JSON sidecars |
| Intended use | CV hazard classifier and VLM finding-draft fine-tuning |
| Volume and refresh | Initial set plus quarterly additions |
How SourceX approaches inspection photo sourcing
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. Buyers describe the data; SourceX looks for US businesses that hold it, and every release is approved by the supplying company.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, allowed uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe your inspection photo requirements to SourceX using the template above; related categories are listed under images and inspection photos, underwriting files and the insurance buyers page.
Request property inspection photos with findings
SourceX finds US companies that hold the inspection photos and findings you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Nothing is contracted until a supplier agrees. Start a buyer request at sourcex.si/buyers.
Frequently asked questions
Can listing photos substitute for inspection photos?
Rarely. Listing photos are staged, wide-angle and chosen to sell, so they omit panels, attics and defects; see real estate listing photos for property AI for what they are good for.
Should findings be region-level or image-level labels?
Start with image-level codes from the report, then add bounding boxes or masks on a stratified subset. Region labels on detail photos give the largest gain for small components like breakers and fittings.
How do construction punch-list photos differ?
Punch-list photos document new-build QA against a spec rather than wear in an occupied home, so privacy risk is lower and defect classes differ; see construction QA and punch-list photos.
Sources
- Journal of Applied Remote Sensing (SPIE), "Residential roof condition assessment system using deep learning" (2018). https://journals.spiedigitallibrary.org/journals/journal-of-applied-remote-sensing/volume-12/issue-01/016040/Residential-roof-condition-assessment-system-using-deep-learning/10.1117/1.JRS.12.016040.full
- McDermott Will & Emery, "State Regulators Address Insurers' Use of AI: 11 States Adopt NAIC Model Bulletin". https://www.mcdermottlaw.com/insights/state-regulators-address-insurers-use-of-ai-11-states-adopt-naic-model-bulletin/
- Illinois General Assembly, "740 ILCS 14/15 (Biometric Information Privacy Act)". http://www.ilga.gov/legislation/ilcs/fulltext.asp?DocName=074000140K15
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- USENIX Security 2023 (Carlini et al.), "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
- CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- Hugging Face, "WebDataset (Hub documentation)". https://huggingface.co/docs/hub/datasets-webdataset
- Transactions of the ACL (Bender and Friedman), "Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science" (2018). https://aclanthology.org/Q18-1041/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.