Skip to content

Industry-specific operational data

Subrogation recovery files for subrogation-identification and demand AI

Quick answer

Subrogation data for AI is the closed-loop record of a paid claim's recovery attempt: the adjuster notes and loss facts that signaled a responsible third party, the referral decision, the demand package, the adverse carrier's response, any inter-company arbitration decision, and the dollars actually recovered. Models that find missed subrogation, draft demands or predict recovery need all of those linked by claim, with jurisdiction and dates, because a referral flag without an outcome teaches nothing about which opportunities pay.

By SourceX Editorial · Updated

This guide is for claims analytics leads at carriers, TPAs and recovery vendors who are scoping external data. It sits inside our industry-specific operational data hub and complements the first-party handling focus of insurance claims datasets.

What a usable subrogation recovery file contains

A usable file links five record types under one claim key: the source claim, the referral, the demand, the response or arbitration, and the recovery ledger. Most carriers hold these in different places, such as the claim system (Guidewire ClaimCenter, Duck Creek Claims or a homegrown platform), a subrogation unit's case-management tool, an outside counsel or vendor portal, and the general ledger for salvage and subrogation receipts. The NAIC's model claims-practices regulation, adopted in varying forms by states, expects claim files to hold notes and work papers detailed enough to reconstruct events and dates [1], which is why the raw material usually exists even when nobody has joined it.

The join is where value is created and where most datasets fail. A recovery check posted to a ledger without the originating claim number, or a demand PDF stored without the payment breakdown it relied on, cannot produce a training example.

Record layerTypical source systemFields that matter for models
Source claimClaim systemLoss date, loss state, line (auto PD, auto collision, homeowners, commercial property, workers' comp), cause of loss, paid indemnity and expense, deductible
Adjuster notesClaim notes, diary entriesFree text with third-party mentions, police report references, product or contractor involvement
ReferralSubrogation queue or case toolReferred flag, referral date, referral source (rule, model, adjuster), screen-out reason
DemandDemand letters and packagesDemand amount, components (paid loss, deductible, rental, loss of use), supporting documents, send date
Response and arbitrationCorrespondence, arbitration filingsLiability accepted or disputed, offered percentage, arbitration decision, assigned liability split
RecoveryLedger, salvage and subrogation receiptsAmount recovered, date received, net of fees, deductible reimbursed to the insured, close reason

Labels that train opportunity detection and recovery prediction

The most useful label set separates "was this claim recoverable" from "was it pursued" and "what came back." Collapsing these into one flag makes a model learn your past referral habits rather than recovery potential.

Ask suppliers for these label fields, defined in a data dictionary:

  • Referred (yes/no) and referral date, plus whether a rule, a model or an adjuster triggered it.
  • Recovery amount versus paid amount, so you can compute recovery ratio by line and cause of loss.
  • Time to recovery, measured from payment or referral to first receipt.
  • Closed-without-recovery reason, such as no liable party, uncollectible tortfeasor, policy limits exhausted, statute of limitations expired, waiver of subrogation, or cost exceeds expected recovery.
  • Liability percentage where an arbitration panel or negotiated settlement assigned comparative fault.

Missed-subrogation detection has a specific trap: claims that were never referred have no observed recovery outcome. This is the selective labels problem, in which the outcomes you can see depend on earlier human decisions, so naive evaluation overstates a model's accuracy [5]. Mitigations include requesting closed-file audit samples where a reviewer re-screened unreferred claims, and treating "not referred" as unlabeled rather than negative. Our guide to outcome-labeled evaluation data covers holdout design for decision-dependent labels.

Where missed-subrogation signals sit in claim notes

Most missed-subrogation signals sit in unstructured adjuster notes, not in coded fields. Cause-of-loss codes are often too coarse to separate a water loss from a failed supply line with a product-defect angle, or a collision from a rear-end impact with clear third-party fault.

Common textual signals include rear-end or "struck while parked" descriptions, a named other driver or other carrier, police report or citation references, contractor or plumber involvement, appliance or component failure, landlord or HOA responsibility, and utility or municipal causes. A detection model therefore needs note text aligned to the claim timeline, with note timestamps, so it learns from what was known at the decision point rather than from notes written after referral. Leakage is the usual failure mode: a note reading "sent to subro" or "AF filing complete" gives away the label and must be masked or cut off at a fixed day after first notice of loss.

For adjacent decision data, see claims adjudication decisions with coverage reasoning; for injury demands specifically, bodily injury claim valuation data covers medical specials and settlements.

Demand packages and arbitration decisions as training data

Demand packages train drafting and completeness models, and arbitration decisions train liability-assessment models, but each carries constraints a buyer must check. A demand package usually includes the demand letter, proof of payment, repair estimate or invoice, photos, police report and the insured's statement, so it is document-heavy and identity-heavy.

Inter-company arbitration programs used by many US carriers produce structured decisions with liability percentages and a short rationale, which makes them attractive labels. Those decisions are created under program membership rules that may restrict disclosure, so ask the supplying carrier to confirm, in writing, whether its program agreement permits licensing decisions for model training and in what form. If not, a supplier may still be able to provide its own derived fields (filed, decided, awarded percentage) without the opposing party's submissions.

Demand letters are useful for generation models only when paired with the response. A demand that was paid in full, one that was countered at 50 percent and one that was denied teach different things; a corpus of demands alone teaches tone, not effectiveness.

Rights, privacy and third-party data in recovery files

Recovery files are dense with information about people and companies who never dealt with the supplying carrier, so rights and privacy review is the gating step. The insured's data is typically nonpublic personal information under the Gramm-Leach-Bliley Act, whose privacy rules state insurance regulators apply to insurers [8]; the FTC's guide explains the core concepts [3], and the federal rule on redisclosure and reuse shows the limits that follow data a recipient obtains from a financial institution [4]. At-fault drivers, claimants, other carriers' claim numbers, adjuster names and repair-shop details are third-party data that should be removed or replaced before delivery.

Practical redaction targets include names, addresses, phone numbers, emails, policy and claim numbers of both carriers, VINs, license plates, driver's license numbers, bank and check details on recovery receipts, and signatures. Photos and police reports need image and document-level handling, not only text scrubbing. Our de-identification evidence package checklist lists what to request as proof, and the re-identification prohibition clause guide explains terms you will likely be asked to accept.

Governance expectations also apply on the buyer side. As of October 2026, many states have adopted the NAIC Model Bulletin on insurers' use of AI systems, which expects documented governance over AI systems and third-party data [2]. If you build generative tools offered in California, AB 2013 has required posting documentation about training data since January 1, 2026 [7]. Dataset documentation in the style of Data Cards, covering sources, collection, annotation and intended use [6], makes both easier.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Jurisdiction, time and line-of-business fields you cannot skip

Recovery outcomes depend on state law, so every record needs the loss state and the governing jurisdiction. Statutes of limitations, comparative-fault regimes (pure, modified at 50 or 51 percent, or contributory negligence) and anti-subrogation or made-whole rules differ by state and line, and a model that ignores them will rank time-barred or legally weak claims as strong opportunities.

Request the payment date, referral date, demand date and limitation deadline the unit tracked, so you can compute time remaining at decision. Line matters too: workers' comp liens, homeowners product-liability recoveries and auto property damage follow different workflows and should be stratified in training and evaluation. Freight losses have their own regime; see freight claims data for AI.

Request template for subrogation recovery data

A precise request describes the data and the labels, not the companies that might hold it. Use the template below as a starting point, whether you send it to your own data partners or describe the dataset to SourceX.

Illustrative example: invented to show structure; it does not describe an available dataset.

request: subrogation recovery files
use_case: missed-subrogation detection from claim notes; demand drafting; recovery prediction
lines: [auto_physical_damage, homeowners_property, commercial_property]
time_window: losses 2019-2024, all closed or recovery-final
unit_of_record: claim_id (tokenized) with linked referral, demand, response, recovery rows
required_fields:
  claim: [loss_date, loss_state, line, cause_of_loss_code, paid_indemnity, deductible]
  notes: [note_timestamp, note_text_redacted, author_role]
  referral: [referred_flag, referral_date, referral_trigger, screen_out_reason]
  demand: [demand_date, demand_amount, components, documents_redacted]
  outcome: [liability_pct, arbitration_flag, recovered_amount, recovery_date, close_reason]
audit_sample: re-screened unreferred closed claims, with reviewer finding
privacy: remove or replace parties, both carriers' claim numbers, VINs, plates, bank details
documentation: data dictionary, join logic, redaction method and QA sample results
restrictions_to_disclose: arbitration program rules, outside counsel or vendor files excluded

How to evaluate a sample before you license

Evaluate a sample on join integrity, label completeness and leakage before scaling a request. Check that every recovery row resolves to a claim, that recovered amounts never exceed paid plus deductible without explanation, and that close reasons are populated for non-recoveries.

Then read 50 to 100 notes against their labels. Confirm that the referral signal is present before the referral date, that "subro" vocabulary is masked, and that redaction did not remove the very facts (vehicle position, product type) that justify liability. Finally, compare recovery ratios by state and line with your own book; a sample from a single region or a single vendor's caseload can skew demand-response behavior.

Sourcing subrogation recovery data for AI

SourceX sources operational datasets from US companies on request, and every release is approved by the supplying company under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery and the method is recorded, though no method is perfect, and a request does not guarantee a match. Describe the recovery data you need at SourceX for buyers, or see how we support insurance buyers and claims administration teams.

Frequently asked questions

Is subrogation data different from general claims data?

Yes. General claims data centers on first-party handling from first notice of loss to closure, while subrogation data adds the recovery workflow: referral, demand, response, arbitration and receipts. Ask for both linked if you need detection and outcome prediction in one model.

Can arbitration decisions be licensed for AI training?

It depends on the supplying carrier's program agreement and the decision content. Ask the supplier to confirm whether disclosure is permitted and whether the opposing carrier's submissions must be excluded.

How much history is enough?

Recovery often closes months after payment, so the window should end early enough that most referred claims have reached a final outcome. Otherwise open recoveries look like failures.

Sources

  1. NAIC Unfair Property/Casualty Claims Settlement Practices Model Regulation (Model 902). https://content.naic.org:443/sites/default/files/model-law-902.pdf
  2. State Regulators Address Insurers' Use of AI: 11 States Adopt NAIC Model Bulletin. https://natlawreview.com/article/state-regulators-address-insurers-use-ai-11-states-adopt-naic-model-bulletin
  3. How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act. https://www.ftc.gov/business-guidance/resources/how-comply-privacy-consumer-financial-information-rule-gramm-leach-bliley-act
  4. 12 CFR 1016.11 Limits on redisclosure and reuse of information. https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  5. The Selective Labels Problem: Evaluating Algorithmic Predictions in the Presence of Unobservables. https://www.cs.cornell.edu/home/kleinber/kdd17-selective.pdf
  6. Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI. https://arxiv.org/pdf/2204.01075
  7. AB-2013 Generative artificial intelligence: training data transparency. https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  8. Privacy of Consumer Financial and Health Information Regulation (Model 672). https://content.naic.org/sites/default/files/model-law-672.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data