Skip to content

Industry-specific operational data

Bodily injury claim valuation data: demands, medical specials and settlements

Quick answer

Bodily injury claim data for AI is a set of closed casualty claim files that link each injury and treatment timeline to billed and paid medical specials, wage loss, liens, the claimant's demand, every offer and counteroffer, and the final settlement or verdict. Valuation and negotiation models need all of those linked at the claim level, plus venue and attorney flags. Buy closed claims, insist on paid-versus-billed detail, and de-identify the embedded medical records before delivery.

By SourceX Editorial · Updated

This page sits under the industry-specific operational data hub and focuses on amount: what a claim is worth and how it settles. Coverage decisions and document extraction are separate problems, covered by pages such as insurance claims AI evaluation. For the broader claims category, see the insurance claims datasets overview.

What a usable bodily injury claim record contains

A usable record joins clinical facts, economic damages and the negotiation path under one claim key, so a model can learn how injury severity and venue translate into dollars. Files that hold only a final payment amount train a reserving curve, not a valuation or negotiation tool. The components below come from adjuster notes, demand packages, medical bills, lien correspondence and payment ledgers in the claim system.

  • Injury description: diagnosis codes (ICD-10-CM where present), body part, injury type (soft tissue, fracture, TBI, surgical), and the adjuster's narrative.
  • Treatment timeline: dates of first treatment and last treatment, provider types (ER, chiropractic, PT, orthopedics, pain management), gaps in treatment, injections, and surgery flags.
  • Medical specials: billed amounts and paid or allowed amounts per provider, since the gap between them is central to many state collateral-source and "paid versus incurred" rules.
  • Economic losses: wage loss claimed and verified, future medicals, and property damage on the same occurrence.
  • Liens and reporting: health plan, hospital, Medicaid and Medicare liens, with amounts asserted and amounts compromised.
  • Negotiation history: time-limited demand amount and deadline, each offer and counteroffer with dates, mediation events, and policy limits.
  • Outcome: settlement, judgment, verdict, dismissal or abandonment, with the paid amount and date.
  • Context flags: jurisdiction and venue, attorney representation and the date it began, suit filed, liability split, and coverage line (auto BI, general liability, premises).

Regulators already expect much of this to exist. The NAIC's Unfair Property/Casualty Claims Settlement Practices Model Regulation includes a file and record documentation section that expects claim files detailed enough to reconstruct pertinent events and dates [1], which is why mature carriers can usually produce a dated activity log alongside the payment ledger.

Medicare Secondary Payer fields and liens

Medicare reporting gives casualty files some of their most structured fields, so ask for them explicitly. Under Section 111 of the Medicare, Medicaid, and SCHIP Extension Act (MMSEA), liability insurers (including self-insured entities), no-fault insurers and workers' compensation plans must report settlements, judgments, awards or other payments to Medicare beneficiaries; confirm current reporting rules in CMS's NGHP user guide [8]. Responsible reporting entities therefore keep beneficiary query results, injury diagnosis codes, the total payment obligation to the claimant and its date, and ongoing responsibility for medicals indicators.

For a valuation model, these fields are useful as validated diagnosis codes and as a clean settlement amount and date. They also carry a trap: Medicare beneficiary identifiers and query responses are direct identifiers and must be removed, while conditional payment and final demand letters often hold itemized medical claims that need the same handling as any medical record. Treat lien amounts as features and lien correspondence as sensitive documents.

Labels: which outcome fields are trustworthy

Settlement amount is the obvious label, but it is often contaminated, so define the label before you buy. A paid total may combine BI with property damage, med-pay or PIP, may be split across multiple claimants under one occurrence, or may be driven by a policy-limits tender rather than injury value. Our page on verifying outcome labels in operational records covers the general checks.

Specific failure modes to test in a sample:

  • Policy-limits censoring: settlements at the per-person limit understate value; flag them and treat them as right-censored.
  • Multi-claimant occurrences: a single per-occurrence limit split across claimants makes per-claimant amounts depend on other people's injuries.
  • Reopened claims: a "closed" status later reversed by a supplemental payment or lien dispute.
  • Verdict versus collected: a judgment above limits may never be collected; record both.
  • Reserve leakage: reserve fields set by the same adjuster who later negotiated can leak the outcome into features; drop or time-gate them.

For negotiation copilots, the label is the offer sequence, not the final number. Treat each offer and response as a dated step, which is the same trajectory reconstruction described in rebuilding agent trajectories from case histories.

Representation, venue and fairness testing

Attorney involvement and venue are among the strongest drivers of outcomes, so a dataset without them will produce a model that silently averages across very different populations. Ask for the date representation began, not just a yes/no flag, because a model can learn that late representation follows low offers. Venue should be at least state and county of suit or of the loss.

Valuation tools trained on historical settlements can reproduce past disparities, and regulators are watching insurers' AI use. The NAIC model bulletin on insurers' use of AI, adopted by many states as of NAIC's 2026 adoption map, expects a written AI program with governance, risk-based controls and oversight of third-party data and models [5]. Use the MEASURE function in NIST AI RMF 1.0 to document how you test for harmful bias [6]: compare predicted versus actual settlements across represented and unrepresented claimants, venues, and proxies such as ZIP code, and keep the results with the model card.

De-identifying medical content in casualty files

Most P&C insurers are not HIPAA covered entities, but bodily injury files are full of records that originated at covered providers, so buyers should expect HIPAA-grade de-identification. HHS describes two methods: Safe Harbor, which removes 18 identifier types and requires no actual knowledge that the remainder could identify someone, and Expert Determination, where a qualified expert finds the risk of re-identification very small [2]. The standard itself sits in 45 CFR 164.514(a)-(b), and 164.514(e) defines limited data sets under a data use agreement [3].

Safe Harbor's date rule is the core tension for valuation data. It removes all date elements except the year, which destroys treatment gaps, demand deadlines and time-to-settlement as calendar dates. Expert Determination is the usual route for keeping shifted dates or day offsets from the date of loss, and finer geography such as county, at the cost of an expert report. If the supplier relies on state law rather than HIPAA, California's CCPA defines deidentified data with conditions including public commitments and contractual limits on recipients [4].

Free text is where de-identification fails. Demand letters, IME reports and adjuster notes contain names of family members, employers, treating physicians and vehicle plates. Tools such as Microsoft Presidio help detect PII, but the official documentation warns it cannot guarantee finding all sensitive information [7], so require a documented method plus a human-reviewed sample.

Closed versus open files and privilege

Prefer closed claims with a final payment and no pending lien dispute, because open litigated files carry legal and labeling risk. Open suits can be subject to protective orders over medical records, and defense counsel correspondence, coverage opinions and claim-evaluation memos may be attorney-client privileged or work product. Ask the supplier to exclude counsel communications by document type, not by keyword, and to confirm no protective order covers the produced files.

Request template for a bodily injury valuation dataset

A precise request lets a supplier check its systems quickly and tells you early whether the fields you need actually exist.

Illustrative example: invented to show structure; it does not describe an available dataset.

Field groupSpecification to sendWhy it matters
ScopeClosed auto BI and premises GL claims, US, closed in a stated multi-year windowDefines population and drift
Claim keyStable pseudonymous claim and claimant IDs; occurrence ID for multi-claimant lossesJoins documents to outcomes
DatesAll dates as day offsets from date of lossKeeps intervals after de-identification
InjuryICD-10-CM codes, body part, surgery and injection flagsSeverity features
SpecialsBilled and paid per provider typePaid-versus-billed modeling
LiensType, asserted, compromisedNet-to-claimant and settlement friction
NegotiationDemand amount and deadline; each offer with offset day and partyNegotiation sequence labels
ContextState, county, represented flag and offset day, suit filed, policy limitOutcome drivers and censoring
OutcomeDisposition type, total paid, split by coverageClean label
DocumentsRedacted demand package, bills, ledger; counsel memos excludedExtraction and summarization training
PrivacyDe-identification method named, sample review resultsDiligence record

A matching illustrative record, flattened:

{
  "claim_id": "c_7f3a",
  "occurrence_id": "o_19b2",
  "line": "auto_bi",
  "state": "XX",
  "represented": true,
  "represented_offset_days": 21,
  "icd10": ["S13.4XXA"],
  "treatment": {"first_offset_days": 1, "last_offset_days": 142, "surgery": false},
  "specials": {"billed": 18400, "paid": 7350},
  "liens": [{"type": "health_plan", "asserted": 6100, "compromised": 4200}],
  "demand": {"amount": 85000, "offset_days": 190, "time_limited": true},
  "offers": [{"offset_days": 214, "amount": 12000}, {"offset_days": 260, "amount": 21500}],
  "disposition": "settled",
  "paid_bi": 27500,
  "policy_limit_per_person": 50000
}

How SourceX handles bodily injury claim requests

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases; claim histories are not held in stock, and a request does not guarantee a match. You describe the data, SourceX looks for US businesses that hold it, and the supplying company approves every release. Each dataset is reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, as discussed in the AI data license negotiation checklist.

Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Health records require HIPAA de-identification by Safe Harbor or Expert Determination. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Insurance buyers can also see how SourceX supports insurance AI teams and claims administration buyers, or start a buyer request.

Request bodily injury claim valuation data

SourceX sources casualty claim files from US companies on request, with rights review, recorded de-identification and a license defining records, uses, term and delivery. Nothing is contracted until a supplier agrees. Describe the claim lines, fields and volumes you need at sourcex.si/buyers.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Frequently asked questions

Can I train on settlement amounts from open claims?

Avoid it. Open amounts are reserves, not outcomes, and open litigated files can carry protective orders and privilege. Use closed claims and treat reopen events as a data quality check.

Is billed or paid medical specials the better feature?

Request both. Paid or allowed amounts reflect what was actually owed after contractual write-downs, while billed amounts are what demand letters cite, and the gap between them varies by state evidentiary rules.

How many years of closed claims should I request?

Long enough to cover litigated tails but recent enough to reflect current venue behavior and medical pricing. Split train and test by close date, not randomly, so you can measure drift.

Sources

  1. National Association of Insurance Commissioners, "Unfair Property/Casualty Claims Settlement Practices Model Regulation (Model 902)". https://content.naic.org:443/sites/default/files/model-law-902.pdf
  2. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  3. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  4. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  5. National Association of Insurance Commissioners, "NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers: adoption map" (2026). https://content.naic.org/sites/default/files/legal-adoption-map-ai-model-bulletin.pdf
  6. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
  7. Microsoft (Presidio Project), "Presidio - Data Protection API Documentation". https://microsoft.github.io/presidio
  8. Centers for Medicare & Medicaid Services, "MMSEA Section 111 MSPSurp User Guide Chapter IV: NGHP Reporting Requirement" (2024). https://www.cms.gov/files/document/mmsea-111-january-2024-nghp-user-guide-v74-chapter-iv.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data