Skip to content

Agent, workflow and domain-reasoning data

Claims adjudication decisions with coverage reasoning

Quick answer

Claims adjudication training data for AI agents is the decision layer of real claim files: each coverage determination (pay, partial pay, deny or reserve rights), the policy or plan provisions it cites, the adjuster's or reviewer's written reasoning, the evidence on file when the decision was made, and what happened next on appeal, reopening or in litigation. It becomes trainable only when those parts share one claim key and carry timestamps that separate what the adjuster knew from what surfaced later.

By SourceX Editorial · Updated

Where coverage decisions and their reasons are written down

Adjudication reasoning rarely sits in one table: it is split across the claims system of record, decision letters, and separate appeal and litigation files, and claims-handling rules are the main reason it exists in writing at all.

Property and casualty. The claims administration system holds claims, exposures, coverage verification, reserves, payments and adjuster notes, while the coverage position itself usually lives in letters (acknowledgment, reservation of rights, denial) in document management. The insurance claims workflow datasets page covers the full file from first notice of loss to closure and the systems it typically comes from; this page covers only the decision layer. The NAIC's Unfair Property/Casualty Claims Settlement Practices Model Regulation (Model 902) sets file documentation standards so that an insurer's activity on each claim can be reconstructed from the file, and it expects a denial based on a specific policy provision, condition or exclusion to reference it in writing [1]. States adopt the model with variations, so the governing state belongs on every record.

Health, disability and other employee benefits. For ERISA-covered plans, 29 CFR 2560.503-1 requires an adverse benefit determination notice to state the specific reasons and refer to the specific plan provisions relied on; on review, claimants can request the documents and, for group health and disability claims, the internal rules or guidelines relied on [2]. For disability claims, the Department of Labor's 2016 final rule adds a fuller discussion of why the claim was denied and the standards applied [3]. Coded medical claims as a product are covered on the medical coding and claims dataset page; here the unit is the determination and its stated basis.

Third-party administrators often adjudicate for carriers or self-funded employers, so the files may belong to their clients; see claims administration data for AI. Upstream, intake data is a separate need, described in FNOL intake conversations for claims-intake agents.

Fields that turn a claim decision into agent training data

A usable record pairs every decision with the provision wording it relied on, the evidence available at that moment, and the person or rule that made it; an extract of the claims table alone yields status codes without coverage reasoning.

Field groupWhat to requireCommon defect
Decision eventDecision ID, claim and exposure (or claim line) keys, decision type (accept, partial, deny, reservation of rights, closed without payment), timestamp, automated or manualOnly the current claim status survives; decision history overwritten
Decision makerRole, authority level, supervisor or medical reviewer sign-offRoles stripped along with names
Policy or plan contextForm and edition, coverage part, endorsements in force on the loss date, limits, deductibles; plan document version for benefit claimsProvision ID delivered without the wording
Cited provisionsEvery provision named in the letter or notice, mapped to the wording in forceFree text such as "per policy terms"
Reason codesInternal denial or adjustment codes plus the codebook and its version historyCodes whose meaning changed over the years
Reasoning textAdjuster or reviewer note, letter rationale, guideline or clinical criteria version appliedTemplated letters where the rationale is boilerplate
Evidence at decisionEach document with its received date: estimates, recorded statements, medical records, police or independent medical examination reportsDocuments received after the decision mixed in
Jurisdiction and timingGoverning state, report, proof-of-loss, acknowledgment and decision datesNo state, so timing rules cannot be checked
Downstream outcomeAppeal or reconsideration (date, new evidence, result), reopening, supplemental payment, regulator complaint, suit and dispositionLitigation kept in a legal system that was never joined

The guideline version matters more than it looks: a health plan denial that applied a medical-necessity criterion is only reproducible if the criterion text and version travel with it, which is the same material a claimant can request on review [2].

Illustrative decision record with a later appeal

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "decision_id": "dec_0007",
  "claim_key": "clm_8c41",
  "line_of_business": "homeowners",
  "governing_state": "OH",
  "policy": {"form_id": "HOME-STD", "edition": "2019-04", "coverage_part": "dwelling"},
  "decision": {"type": "deny", "made_at": "2024-03-18T15:02:00Z", "mode": "manual",
               "maker_role": "desk_adjuster", "authority_level": 2, "supervisor_review": true},
  "cited_provisions": [{"ref": "Section I Exclusions 2.c.(5)", "wording_version": "HOME-STD 2019-04",
                        "text_sha256": "9b1e..."}],
  "reason_code": {"code": "EXCL-SEEP", "codebook_version": "2023.2"},
  "reasoning_text": "[ADJUSTER] Moisture mapping shows staining over months behind vanity; [INSURED] reports odor since autumn. Continuous seepage excluded.",
  "evidence_at_decision": [{"doc": "inspection_photos", "received": "2024-03-11"},
                           {"doc": "mitigation_invoice", "received": "2024-03-14"}],
  "outcomes": [{"type": "appeal", "filed": "2024-05-02", "new_evidence": ["plumber_report"],
                "result": "partial_reversal", "resolved": "2024-06-10"}],
  "litigation": null
}

The plumber report arrived after the denial. An agent trained to reproduce the March decision must not see it; an evaluation of whether the March decision held up must.

Labels: the decision, whether it held up, and hindsight

Treat the original determination and its later fate as separate labels: the first teaches an agent how adjusters applied coverage with the evidence they had, and the second tells you only partly whether they were right.

SignalWhat it measuresBias to plan for
Original determinationHow the insurer or plan applied its coverage termsReproduces historical practice, errors included
QA audit of a sampleHandling and decision qualityDepends on sampling; openIMIS routes a share of unflagged claims to manual QA to estimate model validity [4]
Appeal or reconsideration resultError signal on contested decisionsCovers only decisions someone contested; unappealed denials carry no correctness label
Reopening or supplemental paymentMissed coverage or underpaymentAlso reflects new damage or late evidence
Suit, bad-faith allegation, regulator complaintSevere disputesRare, lagging and often settled confidentially

A vendor whitepaper describes adjudication models trained on historical claims, with feedback looped back in [5], and a published patent application describes reviewer input on flagged claims becoming training data that adjusts rules and routing [6]. That loop inherits whatever the historical process got wrong, so audit it with the methods in historical decision bias in operational labels and verifying outcome labels in operational records. Reviewer reversals of automated decisions are a related source, covered in human override and correction logs.

Legacy claims also lack fields that newer pipelines add: the openIMIS specification calls for a migration script to backfill its AI fields into claims already adjudicated [4]. Ask how any backfilled field was produced. If the records will become a test set, adjudication outcomes as evaluation labels covers split and scoring design.

Rights, privacy and AI-law questions for claims decisions

Claims decisions mix policyholder, claimant and often health information with the insurer's own coverage positions, so rights review has to establish who owns the file, how personal and health details were removed, and what downstream AI laws will ask about the training data.

  • Ownership. Carriers own their files; administrators and independent adjusters usually hold files for clients, whose authorization the license needs.
  • Health and disability claims. Claims held by a health plan or its business associate are protected health information. Disability and workers' compensation files carry similar medical detail even where HIPAA may not apply to the insurer, so ask which rules the supplier applied. HIPAA de-identification runs through Safe Harbor (removing 18 identifier types and having no actual knowledge of identifiability) or Expert Determination (an expert finds the re-identification risk very small) [7]. Safe Harbor's list includes every element of dates directly related to the individual except the year [8], which removes month and day from service dates and can remove them from claim receipt and decision dates as well, the dates you need to measure timing; compare the trade-off in Safe Harbor vs Expert Determination for AI training.
  • Substance use disorder records. Claims carrying records from federally assisted treatment programs may fall under 42 CFR Part 2; the 2024 final rule's compliance date was 16 February 2026 [9].
  • Narrative notes. Adjuster notes and recorded-statement summaries name claimants, witnesses, injuries and vehicles. Ask for the de-identification method used on free text and scanned letters, not only structured fields.
  • Colorado. Insurance and health-care services are consequential-decision areas under SB26-189. As of October 2026, from 1 January 2027 developers of covered automated decision-making technology must give deployers documentation that includes categories of training data, and deployers must explain the technology's role after an adverse outcome and allow human review and reconsideration [10]. Insurers and affiliated entities subject to Colorado's insurer AI disclosure statute (C.R.S. 10-3-1104.9) are deemed in compliance with this part in the practice of insurance, so these duties may not reach most claims AI directly [10]. The Attorney General released interim draft rules on 6 October 2026, with comments due 26 October 2026 [11]. See Colorado SB 26-189 training data documentation and, for insurer AI governance of third-party data, the NAIC bulletin and state rules on third-party data.

SourceX sources operational datasets from US companies, and every dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place. For health records, SourceX requires HIPAA de-identification (Safe Harbor or Expert Determination) before anything is considered for a license, and personal details such as names and account numbers are removed or replaced, with the method recorded for each dataset; no de-identification method is perfect. Teams can describe the claims decision records they need rather than approaching insurers or administrators themselves.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request checklist for claims adjudication records

State the lines of business, states, decision mix and outcome window up front, because those four choices decide whether the records can teach coverage reasoning or only reproduce payment patterns.

  • Lines and claim types (for example homeowners water damage, auto physical damage, workers' compensation medical, group disability), and whether automated decisions are in scope.
  • Governing states, decision years, and the policy form or plan document editions in force across that window.
  • Decision mix: target shares of denials, partial payments and reservations of rights, with paid-in-full decisions as contrast, sampled by stratum rather than at random.
  • Linked artifacts per decision: letter or notice, notes, cited provision wording, reason codebook, guideline or criteria version.
  • A received date on every evidence document, so the decision-time view can be rebuilt.
  • An outcome follow-up window long enough for appeals, reopenings and suits to resolve, with open matters flagged.
  • De-identification method for structured fields, narrative notes and scanned documents, plus a post-processing sample check.
  • Permitted uses: training, fine-tuning, evaluation, and whether derived benchmarks may be kept; see license terms for agent data.
  • A record of training data categories and collection period for your own downstream disclosures.

Red flags in a sample: every denial in a line cites identical boilerplate; codes arrive without a codebook; no claim has an appeal record, which usually means the appeals system was never extracted; or evidence documents lack received dates. For the general pattern of decision data, see decision records with rationale, the insurance counterpart in underwriting decision rationale for underwriting agents, and the AI agent training data hub.

Sourcing claims decisions with the reasoning attached?

Describe the lines of business, states, decision types, linked artifacts and outcome window you need, and the uses you need licensed. SourceX looks for US businesses that hold those records, checks the data and each supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees, and a request does not guarantee a matching dataset. Specify your claims decision dataset with SourceX.

Sources

  1. National Association of Insurance Commissioners, "Unfair Property/Casualty Claims Settlement Practices Model Regulation (Model 902)". https://content.naic.org:443/sites/default/files/model-law-902.pdf
  2. Legal Information Institute, Cornell Law School, "29 CFR 2560.503-1 - Claims procedure" (accessed 2026). https://www.law.cornell.edu/cfr/text/29/2560.503-1
  3. U.S. Department of Labor, Employee Benefits Security Administration, "Fact Sheet: Final Rule Strengthening Protections for Disability Benefit Claimants" (2016). https://www.dol.gov/sites/dolgov/files/EBSA/about-ebsa/our-activities/resource-center/fact-sheets/fact-sheet-final-rule-disability-benefits.pdf
  4. openIMIS Initiative, "Automated Claims Adjudication" (project wiki, accessed 2026). https://openimis.atlassian.net/wiki/x/BgDIN
  5. QBurst (vendor whitepaper), "Automating Claims Adjudication". https://www.qburst.com/downloads/automating-claims-adjudication.pdf
  6. Justia Patents, "Patents by Inventor Irene Victoria Tollinger" (accessed 2026). https://patents.justia.com/inventor/irene-victoria-tollinger
  7. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  8. Electronic Code of Federal Regulations, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (accessed 2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  9. U.S. Department of Health and Human Services, Federal Register, "Confidentiality of Substance Use Disorder (SUD) Patient Records (Final Rule)" (2024). https://www.govinfo.gov/content/pkg/FR-2024-02-16/html/2024-02544.htm
  10. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
  11. Colorado Attorney General, "Colorado Automated Decision-Making Technology & Chatbot Safety Rulemaking" (2026). https://coag.gov/ai/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data