Document AI data
Loss Run Report Data for Commercial Underwriting Extraction
Quick answer
Loss run extraction models need real carrier and TPA loss run reports in many layouts, each paired with normalized claim-level labels: policy period, claim number, date of loss, status, cause of loss, and paid, reserved and incurred amounts as of a stated valuation date. The best training sets count distinct issuers, keep report totals so extracted values can be reconciled, and de-identify claimant names and injury narratives before delivery. Public document benchmarks do not cover this genre, so buyers usually license operational reports.
By SourceX Editorial · Updated
What a loss run extraction dataset has to contain
A usable dataset pairs each source report with a claim-level table that a reviewer has reconciled to the report's own totals. Vendor tools converge on the same target fields: policy number, named insured, carrier, line of business, claim number, date of loss, cause, status, paid, reserves and incurred [1]. Training data should label those fields plus the policy-level context that underwriters actually use.
- Document-level fields: issuing carrier or TPA, named insured, run date, valuation date, policy number and policy period, line of business (GL, auto, workers' compensation, property, umbrella).
- Claim-level fields: claim number, date of loss, date reported, claimant type, cause of loss or accident description, status (open, closed, reopened, subrogation), and paid, reserved and incurred amounts split by indemnity, medical and expense where the carrier reports them.
- Summary fields: per-policy-period claim counts and totals, plus "no losses reported" statements, which are labels in their own right.
Keep the source file as delivered (native PDF, scanned image, XLSX or CSV export) with page-level bounding boxes for every labeled value. Without coordinates you can train a text-to-schema model but cannot evaluate layout-aware extraction or route low-confidence fields to human review, which production pipelines rely on [1].
Why carrier loss run formats break template extractors
Format diversity across issuers, not field count, is the core difficulty. Loss runs arrive as carrier-generated PDFs, broker spreadsheets, TPA portal exports and scans with handwritten annotations [3]. One carrier prints one row per claim; another nests claimants under an occurrence; a third groups by policy year with subtotals that look like claims.
Common failure modes to test against:
- Multi-row claims: a description that wraps onto a second line gets read as a new claim with null amounts.
- Subtotal leakage: policy-year or coverage subtotals extracted as claims, inflating incurred.
- Column drift: "Outstanding" on one form and "Reserve" or "Case Reserve" on another; "Total Incurred" sometimes includes expense, sometimes not.
- Page breaks: tables that continue across pages with repeated or missing headers.
- Sign and format quirks: recoveries shown as negatives or in parentheses, dates in MM/DD/YY versus YYYY-MM-DD.
Table structure recognition research gives you evaluation tools here: PubTables-1M introduced the GriTS metric for scoring predicted table structure and cell content against ground truth [4]. For dedicated table-structure data, see table structure recognition data from real financial and operational documents. Track issuer count as a first-class dataset statistic, and hold out entire issuers for evaluation so you measure generalization to unseen layouts rather than memorized templates.
A normalized claim-level label schema
Normalization labels should map every issuer's vocabulary into one schema while preserving the raw string. That lets you train extraction and normalization separately and audit both. The record below shows the structure buyers can ask suppliers to deliver.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"document_id": "lr-000412",
"issuer_type": "carrier",
"issuer_layout_id": "layout-17",
"line_of_business": "workers_compensation",
"valuation_date": "2026-06-30",
"policy_period": {"start": "2024-07-01", "end": "2025-07-01"},
"claims": [
{
"claim_number": "WC-[REDACTED-04]",
"date_of_loss": "2024-11-14",
"status_raw": "CLSD",
"status": "closed",
"cause_raw": "Strain - lifting",
"cause_normalized": "strain_or_sprain",
"paid": {"indemnity": 4200.00, "medical": 3150.00, "expense": 610.00},
"reserve": {"indemnity": 0.00, "medical": 0.00, "expense": 0.00},
"incurred_total": 7960.00,
"bbox_page": 2,
"claimant": "[CLAIMANT-0091]"
}
],
"report_totals": {"claim_count": 3, "incurred_total": 21480.00},
"reconciled": true
}
Agree the schema before labeling starts. A shared document annotation schema for layout, reading order, tables and fields helps if the same team also labels ACORD applications; loss runs need their own claim-level child table, which is why they sit apart from ACORD form data for submission intake.
Consistency checks that turn report totals into label QA
Loss runs carry their own checksums, so a dataset can be validated without a second annotator. The basic rule is that paid plus outstanding equals incurred for each claim [2]. Stack further checks on top:
| Check | Rule | What a failure usually means |
|---|---|---|
| Claim arithmetic | paid + reserve = incurred, per component | Misread digit, wrong column, expense excluded |
| Report reconciliation | sum of claim incurred = printed policy-period total | Missed row, subtotal extracted as claim |
| Claim count | labeled claims = printed count | Wrapped row split, duplicate across page break |
| Date logic | date of loss within policy period; valuation date after loss | Swapped date fields, two-digit year error |
| Status logic | closed claims carry zero reserve | Status or reserve mislabeled |
Ask suppliers which checks were run and what share of documents passed, and keep failed documents flagged rather than silently dropped. Valuation date matters for evaluation too: two loss runs for the same policy at different valuations should disagree on reserves, and a model that ignores valuation date will look wrong on both. For a broader view of label reliability, see verifying ground truth in operational records.
De-identifying claimant names and injury narratives
Claimant names, claim numbers, adjuster names and free-text injury descriptions are the main privacy exposure in loss runs. Workers' compensation and bodily injury runs can describe diagnoses and body parts in plain language. Whether or not HIPAA applies to a given report, the HHS Safe Harbor list of 18 identifiers and its Expert Determination alternative are a practical reference for what to remove or replace [6].
Ask how replacement was done. Consistent pseudonyms (the same claimant token across reports) preserve multi-claim patterns; random redaction breaks them. Confirm that redaction was applied to the image layer as well as the text layer, since a scanned PDF with a text overlay can leak the original name.
Rights, provenance and documentation questions to put to suppliers
Loss runs are carrier or TPA records issued to the insured, often through a broker, and they describe third-party claimants, so who may license them should be established before any sample moves. Dataset licensing is often missing or wrong on public hosting sites [8]. As of October 2026, the NAIC model bulletin, adopted by many states, expects insurers using AI to document their systems and oversee third-party data and models [5].
Illustrative example: invented to show structure; it does not describe an available dataset.
Loss run data request checklist
- Lines of business and the share of each.
- Number of distinct issuing carriers and TPAs, and documents per issuer.
- File types: native PDF, scanned, XLSX, CSV.
- Valuation date range and whether multiple valuations per policy exist.
- Label schema, annotation guide and reconciliation pass rate.
- De-identification method for names, claim numbers and narratives.
- Who holds the right to license the reports, and what uses the license permits.
- A dataset card covering sources, collection, annotation and intended use [7].
For full submission packages rather than loss runs alone, see underwriting files for AI training and the insurance buyer overview.
How SourceX sources loss run data
SourceX sources operational datasets from US companies and manages the licensing process, including ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match; you describe the reports you need and SourceX looks for US businesses that hold them, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows only after an executed agreement. You can describe your loss run data requirements on the buyers page. Related task pages are collected under document AI datasets by task and the AI data hub.
Request loss run extraction data
Describe the lines of business, issuer variety, label schema and valuation dates your extraction model needs. SourceX follows a Find, Assess, Agree, Transact and Manage process, and nothing is contracted until a supplier agrees. Start a loss run data request.
Sources
- Sensible, "Extract Loss Runs". https://www.sensible.so/extract-old/loss-runs
- Sensible, "Automating data extraction from loss runs". https://sensible.so/blog/automating-data-extraction-from-loss-runs
- FinTech Global, "IntellectAI breaks new InsurTech ground with AI-powered loss run extraction" (2024). https://fintech.global/2024/03/04/intellectai-breaks-new-ground-in-insurtech-with-ai-powered-loss-run-extraction/
- arXiv (Smock, Pesala, Abraham), "PubTables-1M: Towards comprehensive table extraction from unstructured documents" (2021). https://arxiv.org/pdf/2110.00061
- National Association of Insurance Commissioners, "NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers" (2024). https://content.naic.org/sites/default/files/inline-files/AI Model Bulletin - April 2024.pdf
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.