Industry-specific operational data
Life underwriting medical evidence: APS summaries and impairment decisions for underwriting AI
Quick answer
APS summarization training data pairs attending physician statement pages with what an underwriter did with them: the extracted impairments, the written case summary, and the final decision (rating class, table rating, flat extra, postpone or decline). The usable unit is a case, not a document. Expect scanned, partly handwritten records released under underwriting-scoped authorizations, so buyers must plan for HIPAA de-identification [1], state insurance privacy rules and page-level labeling before any of it reaches a training run.
By SourceX Editorial · Updated
What a usable APS training case contains
A usable case links four layers by a stable case key: the raw evidence, the underwriter's reading of it, the decision, and later outcomes where they exist. Most carriers store these in different systems, so a dataset that ships only one layer is far less valuable than its page count suggests.
- Evidence layer. APS page images (TIFF or PDF), usually ordered by the provider rather than chronologically, plus any lab panels, prescription history hits, paramedical exam forms and Part B medical questionnaire answers attached to the case.
- Reading layer. The underwriter or nurse reviewer's case summary, impairment list, notes on what drove the debits, and any requirement orders (additional APS, EKG, specialist report) triggered by the evidence.
- Decision layer. Applied-for versus approved class, table rating (for example Table 2 or Table 4), flat extras with amount and duration, exclusion riders, postpone periods, decline reasons and the manual references cited.
- Outcome layer. Placement, reconsiderations, rescission reviews and, for reinsurers and audit teams, post-issue audit findings and early claims.
The reading layer is the scarce part. Page images without the underwriter's summary train OCR and retrieval; with it, they train summarization and impairment extraction against how underwriters actually weigh evidence. For the broader file structure, see the guide to underwriting decision rationale data for underwriting agents and the underwriting files dataset page.
Why APS reuse for AI depends on authorizations and de-identification
APS records are released by providers to the insurer under the applicant's signed authorization, and that authorization is typically scoped to underwriting the application, so reusing the pages to train a model usually needs separate authority or de-identification. Treat this as the working hypothesis to test with counsel for every source, not a settled rule. The insurer may be outside HIPAA as a covered entity for life business, but the evidence is still health information governed by the authorization terms, state insurance privacy law and the insurer's own notices.
De-identification is the common path. HHS recognizes two methods under the Privacy Rule: Safe Harbor, which removes 18 listed identifiers of the individual and of relatives, employers or household members with no actual knowledge that the rest could identify the person, and Expert Determination, where a qualified expert documents that re-identification risk is very small [1][2]. The guidance sets no numeric threshold for "very small" [1]. A limited data set under a data use agreement is a narrower, HIPAA-specific route that still keeps dates and some geography [2].
Three record types need extra care:
- Substance use disorder treatment records. Records from federally assisted SUD programs fall under 42 CFR Part 2; the 2024 final rule aligned much of Part 2 with HIPAA, with a compliance date of February 16, 2026 [4]. Flag these pages during intake rather than after labeling.
- Consumer health data under state law. Washington's My Health My Data Act defines consumer health data broadly, including inferred health status [5]; check whether any component of the pipeline sits outside insurer potential exemptions.
- Genetic test results, HIV status and mental health notes. Many states restrict these categories in insurance specifically; require the supplier to identify and either exclude or separately authorize them as their use for AI may be limited by state insurance privacy law.
The guide to de-identifying underwriting files covers the mechanics in more depth.
Why APS de-identification is harder than it looks
APS pages defeat simple redaction because identifiers hide in places structured PII scanners do not look. Letterhead carries practice names and addresses, fax headers carry phone numbers and timestamps, handwritten margin notes carry names of relatives, and lab printouts repeat medical record numbers in footers.
Common failure modes buyers should test for in a sample:
- OCR-gated redaction. If redaction runs on OCR text, every misread name survives on the image. Ask whether images were redacted, re-rendered, or dropped in favor of text.
- Date leakage. Safe Harbor requires removing all date elements except year that relate to the individual [1]; shifted dates must be shifted consistently per case or interval features (time since diagnosis, time since last A1C) break.
- Rare-impairment re-identification. A small-town applicant with a rare cardiomyopathy and an unusual face amount can be unique even with names removed. NIST's survey documents real re-identification of data that was believed de-identified [6].
- Free-text in the reading layer. Underwriter summaries often restate the applicant's employer, occupation and avocations; scrub them with the same rigor as the APS.
How to label APS pages for summarization and impairment extraction
Good APS labels exist at three levels: page, impairment and case. Page-level labels (page type, provider, date of service, legibility) let you train page classifiers and measure OCR quality before you trust anything downstream. Impairment-level labels tie each finding to its supporting page and span, which is what makes summary evaluation possible.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"case_id": "c-7f3a91",
"product": "term_20",
"face_band": "500k_1m",
"issue_age_band": "45_49",
"pages": [
{"page_id": "p012", "page_type": "progress_note", "date_shifted": "2019-03-14",
"handwritten": true, "ocr_confidence_mean": 0.71, "legible": "partial"},
{"page_id": "p031", "page_type": "lab_result", "date_shifted": "2021-06-02",
"handwritten": false, "ocr_confidence_mean": 0.96, "legible": "yes"}
],
"impairments": [
{"name": "type_2_diabetes", "icd10": "E11.9", "onset_year_offset": -6,
"control": "fair", "evidence": [{"page_id": "p031", "field": "HbA1c", "value": "7.4"}]},
{"name": "hypertension", "icd10": "I10", "control": "controlled",
"evidence": [{"page_id": "p012", "span": "BP 128/82 on lisinopril"}]}
],
"underwriter_summary": "45M T2DM dx ~6y, oral meds only, A1C 7.4, no end-organ findings; HTN controlled.",
"decision": {"applied_class": "preferred", "final_class": "standard",
"table_rating": null, "flat_extra": null,
"decline_reason": null, "manual_refs": ["diabetes_type2_section"]},
"decision_maker_role": "senior_underwriter",
"rules_engine_recommendation": "refer",
"outcome": {"placed": true, "audit_flag": false}
}
Fields worth insisting on: the evidence pointers (without them you cannot score hallucinated impairments), the rules-engine recommendation alongside the human decision (so you can separate triage models from decision models), and the underwriting manual version in force on the decision date. Manual changes, such as a reinsurer revising its build or A1C tables, silently shift labels across years.
Matching the dataset to the model you are building
Each underwriting AI application needs a different slice of the same case, and buying the wrong slice is the most common waste. Use this table to scope a request.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Application | Minimum layers | Critical labels | Typical evaluation |
|---|---|---|---|
| APS page classification and OCR | Evidence | Page type, legibility, handwritten flag | Page-type F1; character error rate on handwritten pages |
| APS summarization | Evidence + reading | Underwriter summary with evidence pointers | Impairment recall, unsupported-claim rate, reviewer edit distance |
| Impairment extraction and coding | Evidence + reading | Impairment, ICD-10, severity, control, onset | Per-impairment precision and recall on rare conditions |
| Rating and class recommendation | Reading + decision | Final class, table, flat extra, decline reason | Agreement with senior underwriters; disagreement review |
| Accelerated underwriting triage | Decision + outcome | Rules-engine route, full-UW result, audit finding | Misclassification rate against fully underwritten audit |
| Model monitoring and back-testing | Decision + outcome | Manual version, decision date, outcome | Drift by manual version and cohort |
For summarization specifically, the reference summaries must be written by underwriters for underwriting, not generated after the fact. See summarization evaluation sets with expert references and summarization fine-tuning data from business records for how to structure pairs. For decision labels, read verifying outcome labels in operational records before treating final class as ground truth: reconsiderations and exception approvals make some decisions noisy.
How insurance AI regulation shapes what you must document
Regulators expect insurers to explain the data behind underwriting AI, so training-data documentation is part of the deliverable, not an afterthought. New York DFS Insurance Circular Letter No. 7, issued July 11, 2024, sets expectations for insurers using AI systems and external consumer data in underwriting and pricing, including governance, documentation and testing for unfair discrimination [3]. Other states are adopting the NAIC model bulletin approach as of October 2026.
What that means for a dataset purchase:
- Provenance per record. Source system, extraction date, authorization basis or de-identification method, and which manual version governed the decision.
- Protected-class analysis readiness. If you will test for disparate outcomes, decide up front whether you need inferred or proxy demographic fields, and get counsel's view before requesting them, because the same fields raise re-identification risk.
- Dataset documentation. A data card covering upstream sources, annotation methods, intended use and known gaps [7] gives your model risk team something to review.
Treat EU requirements separately: life underwriting AI used in the EU may be high-risk under the AI Act, which brings Article 10 data governance duties starting 2 December 2027 for Annex III systems; see the EU AI Act Article 10 guide.
Request checklist for APS and underwriting decision data
Use this checklist to brief any supplier, internal or external, before samples change hands.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Scope. Product lines (term, UL, IUL, final expense), face-amount bands, issue-age bands, decision years, fully versus accelerated underwritten.
- Layers. Which of evidence, reading, decision and outcome are present, and the join key between them.
- Rights. The authorization language used at application, any later consent, and whether de-identification is Safe Harbor or Expert Determination with the expert's report available for review [1].
- Exclusions. How Part 2 records, genetic results, HIV and mental health notes were detected and handled [4].
- Redaction method. Image-level or text-level, the tooling, and the QA sample size and error findings.
- Label provenance. Who wrote summaries (underwriter, nurse, vendor), manual versions in force, and rules-engine outputs.
- Format. Page images and resolution, OCR output format (hOCR, ALTO or JSON with bounding boxes), and the case schema.
- License terms. Permitted uses (training, evaluation, fine-tuning), term, derivative-model rights and deletion obligations, all in writing.
How SourceX approaches life underwriting evidence requests
SourceX sources operational datasets from US companies on request and manages the commercial process, including the licensing agreement. Nothing is held in stock: a request for APS cases is a brief, and it does not guarantee a match. SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect; health records require HIPAA de-identification through Safe Harbor or Expert Determination. Delivery runs through private, access-controlled workflows only after an executed agreement. If you are scoping a request, start from the SourceX buyer intake or the insurance buyers page; related cases are covered in the industry data hub and the guide to healthcare LLM fine-tuning data under HIPAA.
Source APS summarization training data with SourceX
SourceX sources operational datasets, including documents and finance and legal workflow records, from US companies on request, with every dataset rights-reviewed and licensed with defined records, uses, term and delivery. A request does not guarantee a match, and nothing is contracted until a supplier agrees. Describe the cases, layers and labels you need at sourcex.si/buyers.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Frequently asked questions
Can we train on APS pages we already received for underwriting?
Not automatically. The authorization that let the provider release the pages was usually scoped to evaluating the application, so internal model training generally needs counsel's sign-off on the authorization language, state insurance privacy law and your privacy notices, or a de-identified copy [1].
Is Safe Harbor enough for APS images?
Safe Harbor is a legal standard, not a tool, and it requires no actual knowledge that remaining data could identify the person [1][2]. On scanned APS pages, achieving it reliably usually means image-level redaction plus a QA sample; many teams use Expert Determination when they need to keep dates or geography.
How is APS data different from clinical EHR data?
APS packets are what a provider chose to send in response to a request: selected notes, labs and letters, often faxed and scanned. They lack the structured codes of an EHR extract and carry insurer-side labels instead, which is why they suit underwriting summarization but not clinical prediction. For clinical record sourcing, see the medical records dataset page.
Sources
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- New York State Department of Financial Services, "Insurance Circular Letter No. 7 (2024): Use of Artificial Intelligence Systems and External Consumer Data and Information Sources in Insurance Underwriting and Pricing" (2024). https://www.dfs.ny.gov/industry-guidance/circular-letters/cl2024-07
- U.S. Department of Health and Human Services (SAMHSA and OCR), Federal Register, "Confidentiality of Substance Use Disorder (SUD) Patient Records, Final Rule (89 FR, No. 33)" (2024). https://www.govinfo.gov/content/pkg/FR-2024-02-16/html/2024-02544.htm
- Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act" (2023). https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
- National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
- Pushkarna, Zaldivar, Kjartansson (Google Research), FAccT 2022, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.