Skip to content

Industry-specific operational data

Property claim estimates and adjuster field reports: line items, scope notes and photos

Quick answer

Property claim estimate data for AI is most useful when the estimate is the unit of record: every line item (category code, quantity, unit, unit price, depreciation, overhead and profit) is linked to the adjuster's scope notes, the inspection photos, and what happened next, meaning supplements, reinspection changes and the final settled amount. Buy the full claim lineage from carriers, independent adjusting firms or contractors, confirm the rights to any embedded estimating-software price data, and strip people, plates, addresses and EXIF before training.

By SourceX Editorial · Updated

What a complete property estimate record contains

A training-grade estimate record is a versioned claim file, not a single PDF. Field adjusters and desk reviewers produce several linked artifacts per loss, and estimating, review and photo-to-line-item models each depend on a different subset of them. Ask suppliers to export each layer separately with stable claim and estimate keys rather than flattening everything into printouts.

  • Estimate header: claim key, policy form (for example HO-3 or a commercial property form), cause of loss, date of loss, deductible, coverage A/B/C split, price list region and month.
  • Line items: category and selector codes from the estimating system (Xactimate and Symbility are the common platforms), description, quantity, unit (SF, LF, SQ, EA), unit price, labor and material split, tax, age and life, depreciation percentage and whether it is recoverable, and overhead and profit flags.
  • Sketch and measurements: room or roof-facet dimensions, pitch, waste factor, and the export format (ESX, XML or a structured JSON) so quantities can be checked against geometry.
  • Field report: the adjuster's narrative, scope notes per room or elevation, test squares and hail hit counts, moisture readings, and the causation statement.
  • Photo log: each image with its caption, room or elevation tag and the line items it supports.
  • Revisions: contractor supplements, desk-review adjustments, reinspection reports, appraisal or re-estimate results, and the final payment.

Claim-handling regulation explains why these layers often exist: the NAIC model claims-settlement regulation expects claim files to hold documentation detailed enough to reconstruct pertinent events and dates [7]. That means estimate revisions and reinspection notes are often retained in claim systems, even when they are not exported by default.

Damage detection only becomes an estimating product when it is tied to cost, so detection labels without priced line items train half a model. Practitioner pipelines pair a damage detector with a cost estimator trained on labeled images [4], and roof-assessment research applies deep learning to segment and classify roof condition from high-resolution aerial imagery [1]. Aerial roof-damage work positions its models to prioritize adjuster attention rather than replace field inspection [2].

Outcome labels matter because the first estimate is often wrong. The AI Incident Database records a homeowner non-renewal reportedly based on AI-analyzed roof imagery that an independent inspection disputed [3]; it was an underwriting decision, not an estimate, but it shows the same risk of trusting an unverified first read. Supplements and reinspection deltas give you the error signal: which line items were missed, overstated or reclassified, and why.

For imagery-first work, see the companion guides on property inspection photos linked to findings and adjuster comments used as image captions. This page keeps the priced estimate at the center.

Model use cases and the fields each one needs

Each estimating-AI use case needs a different join across the claim file, so specify the use case before you request data. The table below maps common targets to minimum fields and the most frequent defect.

Illustrative example: invented to show structure; it does not describe an available dataset.

Use caseMinimum fieldsLabel or targetCommon defect in supplied data
Photo-to-line-item suggestionPhoto, caption, room or facet tag, line-item codesLine items supported by each photoPhotos not mapped to line items; captions blank
Scope-of-loss extractionField narrative, scope notes, final line itemsStructured scope from free textNotes typed after the estimate, copying it
Desk review and leakage detectionInitial estimate, reviewer changes, reasonsAccepted or adjusted line itemsReviewer edits overwrite the original version
Supplement predictionInitial estimate, contractor supplement, approval decisionApproved supplement items and amountsSupplements stored as scanned PDFs only
Depreciation and O&P reviewAge, life, condition, policy form, recoverable flagCorrect depreciation treatmentPolicy form missing, so treatment is unlabeled
Evaluation setFinal settled estimate plus reinspectionField-verified line itemsFew reinspected files; heavy CAT-event skew

Sourcing channels and what each holds

Carriers, independent adjusting firms, third-party administrators and restoration contractors each hold a different slice of the estimate lineage. Carriers and TPAs typically hold the full revision history and payment outcome; independent adjusters hold field reports and photos for many carriers but may need carrier consent to release them; contractors hold supplements and the as-built scope from their side of the dispute.

Mix matters for generalization. A file set drawn from one hurricane or one hail season will overrepresent roof replacements and wind claims, while water, fire and theft losses carry very different line-item profiles. Ask for counts by peril, state, property type, catastrophe versus non-catastrophe, and year, and expect regional price list and code differences to shift unit prices.

If you are also buying related claim streams, compare with freight claims files, construction cost estimate and bid data for repair-cost modeling, and loss run extraction data for underwriting. The broader industry-specific operational data hub explains how SourceX approaches these categories.

Rights questions specific to estimate files

Estimate files mix the supplier's own records with third-party licensed content, so rights review has to cover both. Line-item unit prices usually come from an estimating vendor's regional price list, and the license under which the carrier or adjuster uses that software may limit redistribution of price data or code tables. Ask suppliers to confirm what may be included, and be ready to accept quantities and codes with prices removed or normalized if the price list cannot travel.

Photos raise separate questions. Inspection images often capture homeowners, contractors, neighbors, license plates, house numbers and interior personal items; stock-licensing practice treats recognizable people, including by surroundings, as needing a release [8]. For training, the practical response is redaction of faces, plates and address numbers plus a written basis for the use.

Insurer buyers have their own governance duties. The NAIC AI model bulletin expects insurers to govern AI systems and the third-party data behind them [6]; see the guide to NAIC expectations on external data. Keep supplier documentation in a format your model risk team can file.

Redaction and quality checks before training

Photos and narratives in claim files carry identifiers in places text-only pipelines miss, so build redaction and QA into acceptance. An audit of a large ML image dataset found non-empty Exif tags relating to timestamps, geolocation and individuals, and noted that its download tooling extracted that metadata [5]. Strip EXIF and XMP (including GPS and device serials) from every image and check that sketch files do not embed the loss address.

Narratives leak through quasi-identifiers: a rare loss type, a small town and a date can identify a claimant even after names are removed. The guide to indirect identifiers in business text covers the patterns. Then check for template duplication: adjusters reuse boilerplate scope paragraphs, and near-duplicate text inflates apparent volume and encourages memorization [9].

Buyer acceptance checklist

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Every line item resolves to an estimate version and a claim key; every photo resolves to at least one line item or an explicit "context" tag.
  • Initial, supplemented, reviewed and final estimates are all present and ordered by timestamp.
  • Price list region and month are recorded, or prices are removed with a documented reason.
  • Faces, plates and house numbers are redacted; EXIF, XMP and GPS fields are empty on a sampled re-scan.
  • Boilerplate scope paragraphs are deduplicated or flagged.
  • Peril, state, year and CAT-event mix are reported, along with known gaps.
  • A dataset card describes sources, collection, preparation and intended use [10].

How SourceX handles property claim estimate requests

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases; it holds no stock, and a request does not guarantee a match. You describe the data you need, such as estimates with supplements and reinspection outcomes, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and delivery runs through private, access-controlled workflows only after an executed agreement. SourceX does not train models and does not source scraped web content or generic photos. You can describe your estimate data requirement to SourceX or read more on insurance buyers, claims administration buyers and insurance claims datasets.

Sourcing property claim estimate data

SourceX sources operational records such as claim estimates, field reports and inspection photos from US companies on request, with each release approved by the supplier and delivered under a license after rights review. Share the fields, perils and revision history you need, and SourceX will look for matching suppliers. Start a buyer request.

Frequently asked questions

Can I get estimates without the estimating software's price data?

Often yes. Quantities, units, category codes and depreciation logic can carry most of the training value, and you can reprice line items against your own licensed price list. Agree with the supplier in writing which columns are dropped or normalized.

How many reinspected files do I need for an evaluation set?

There is no fixed number; size the set to the decisions you will make. Stratify by peril and region so a single catastrophe event does not dominate, and hold out by claim, not by line item, to avoid leakage between versions of the same estimate.

Are adjuster narratives usable if they were written after the estimate?

They are usable for summarization but weak for scope extraction, because the text often restates the estimate. Ask for note timestamps relative to estimate versions so you can filter.

Sources

  1. Journal of Applied Remote Sensing (SPIE), "Residential roof condition assessment system using deep learning". https://journals.spiedigitallibrary.org/journals/journal-of-applied-remote-sensing/volume-12/issue-01/016040/Residential-roof-condition-assessment-system-using-deep-learning/10.1117/1.JRS.12.016040.full
  2. SCITEPRESS, "Superpixel-wise Assessment of Building Damage from Aerial Images" (Lucks et al., 2019). https://www.scitepress.org/Papers/2019/72538/72538.pdf
  3. AI Incident Database, "AI Incident Database: entity Homeowners". https://incidentdatabase.ai/fr/entities/homeowners/
  4. Sogeti Labs, "Vehicle Damage Assessment whitepaper" (2019). https://labs.sogeti.com/wp-content/uploads/sites/2/2019/05/Whitepaper_VehicleDamageAssessment.pdf
  5. arXiv, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
  6. National Association of Insurance Commissioners, "NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers" (2024). https://content.naic.org/sites/default/files/inline-files/AI Model Bulletin - April 2024.pdf
  7. National Association of Insurance Commissioners, "Unfair Property/Casualty Claims Settlement Practices Model Regulation (Model 902)". https://content.naic.org:443/sites/default/files/model-law-902.pdf
  8. Adobe Stock Contributor Help, "Model release overview". https://helpx.adobe.com/ca/stock/contributor/content-policies-guidelines/model-property-releases/model-release-overview.html
  9. arXiv / ACL 2022, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  10. arXiv / FAccT 2022, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data