Skip to content

Industry-specific operational data

Claim denial appeal letters with outcomes: training data for appeal-drafting AI

Quick answer

A useful denial appeal letter dataset for AI links each letter to three things: the denial that triggered it (CARC, RARC and group code from the 835 remittance), the appeal level and regime it was filed under, and the payer's final decision with the amount paid. Without that outcome label, letters teach a model how appeals sound, not which arguments overturn denials. Buyers should specify the chain, the coverage mix and the PHI handling before pricing any volume.

By SourceX Editorial · Updated

What counts as one training record

One record is a chain, not a letter: denial notice, then one or more appeals, then the payer decision. The denial side comes from the 835 electronic remittance, where Claim Adjustment Reason Codes (CARCs) and Remittance Advice Remark Codes (RARCs) explain the adjustment; both lists are maintained nationally and revised several times a year [1]. The appeal side is the letter text plus whatever was attached: chart notes, an operative report, a letter of medical necessity, the payer's own medical policy excerpt. The outcome side is the payer's written decision and, ideally, a follow-up 835 showing the paid amount.

Ask suppliers whether each chain is joined on a stable claim key (patient control number, payer claim control number) or reconstructed by matching dates and amounts. Reconstructed joins are where mislabeled outcomes creep in. Our guide on verifying outcome labels in operational records covers the checks in more depth.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "chain_id": "APL-000417",
  "denial": {
    "group_code": "CO",
    "carc": "50",
    "rarc": [],
    "denial_category": "medical_necessity",
    "denial_date": "2025-03-04"
  },
  "appeals": [
    {
      "level": "level_1",
      "regime": "commercial_erisa",
      "letter_text_ref": "letters/APL-000417-L1.txt",
      "attachments": ["chart_note", "payer_policy_excerpt"],
      "cited_policy_id": "payer_policy_redacted_17",
      "decision": "upheld"
    },
    {
      "level": "level_2",
      "regime": "commercial_erisa",
      "letter_text_ref": "letters/APL-000417-L2.txt",
      "attachments": ["chart_note", "peer_reviewed_article", "physician_attestation"],
      "decision": "overturned",
      "paid_amount_ratio": 1.0
    }
  ],
  "payer_type": "commercial",
  "specialty": "orthopedic_surgery",
  "template_id": "tmpl_ortho_mn_v3",
  "deid_method": "expert_determination"
}

The pair of appeals in that chain is the most valuable shape: the same denial, one argument that failed and one that worked.

Why outcome labels change what the model learns

Outcome labels turn a style corpus into a supervision signal. Upheld and overturned letters for the same CARC and denial category form natural preference pairs, and methods such as Direct Preference Optimization train directly on preferred-versus-rejected pairs without a separately trained reward model [7]. For supervised fine-tuning, research on LIMA showed that about 1,000 carefully curated examples can be enough for effective alignment [8], which argues for selecting strong overturned letters over training on every letter. If you plan a reward model, see our page on training data for reward models.

Treat outcome labels with care. An overturn can follow a corrected claim rather than the letter's argument, and a "no response" is not an upheld appeal. Ask how the supplier distinguishes overturned on appeal, overturned after resubmission, partially paid and withdrawn.

Appeal levels and regimes to record per letter

Every letter should carry its appeal level and legal regime, because the rules and the reviewers differ. ERISA-governed group health plans follow the Department of Labor claims procedure in 29 CFR 2560.503-1, which requires adverse benefit determination notices to state the specific reason and reference the plan provisions relied on [2]. Medicare Advantage reconsiderations, state Medicaid managed care appeals and state-regulated fully insured plans follow different processes and external review paths.

Record the regime as a field rather than inferring it from the payer name, and verify current deadlines and levels with counsel before you build any rule-based logic on them. A model that drafts a Medicare Advantage reconsideration in the format of a commercial level 2 appeal is an avoidable failure.

Coverage fields to put in the request

Specify coverage along four axes, then ask for counts per cell before agreeing a price. Coverage gaps, not total volume, usually decide whether a model generalizes.

FieldValues to requestWhy it matters
Payer typeCommercial, Medicare Advantage, Medicaid managed careDifferent reviewers, policies and appeal paths
Denial categoryMedical necessity, coding, authorization, timely filing, eligibilityArgument structure differs sharply by category
CARC / RARCCode list with per-code countsTies letters to the remittance signal [1]
Appeal levelLevel 1, level 2, external reviewLater levels cite more evidence and policy
Specialty and settingSpecialty, inpatient vs outpatient, place of serviceClinical vocabulary and cited guidelines vary
OutcomeOverturned, partially paid, upheld, withdrawn, pendingThe supervision signal
Template IDSupplier or vendor template identifierExposes boilerplate inflation

For timely filing and eligibility denials, letters are short and procedural; for medical necessity, they cite clinical criteria and payer medical policy. A drafting model needs both, but in proportions that match your target workload. The adjacent claim denial prediction training data page covers the structured 837/835 side if you also want to predict denials before they happen.

PHI in appeal letters and attachments

Appeal letters are protected health information. They quote clinical history, dates of service, member IDs and provider details, and attachments are often scanned chart pages. HIPAA offers two de-identification routes, Safe Harbor removal of 18 identifier types or Expert Determination [3]; a limited data set under a data use agreement is a separate option with different permitted uses [4].

Free text is the hard part. Identifiers appear inside narrative sentences, signatures, fax headers and letterheads, so de-identification needs NLP-based PHI detection with human review, not find-and-replace on known fields [5]. Ask for the method, the recall measured on a held-out sample, and how dates were shifted so that intervals between denial, appeal and decision survive. Our guide to de-identifying clinical free text for LLM training goes further.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Boilerplate, templates and duplicate letters

Templated boilerplate inflates volume without adding signal. Billing teams and appeal vendors reuse the same paragraphs with a swapped patient name and code, so a corpus can look large while holding few distinct arguments. Research on language-model corpora found that near-duplicates are common and that deduplication cuts verbatim memorization while keeping or improving model quality [6], which matters doubly when the duplicated text once held PHI.

Before pricing, ask for:

  • Near-duplicate rate (for example, MinHash similarity above a stated threshold) across letter bodies.
  • Counts per template ID and the share of letters with substantive edits beyond the template.
  • Whether attachments are included as text, images or references only.
  • How letters drafted by third-party appeal vendors are marked.

Building a held-out evaluation set

Hold out evaluation chains by payer and time period, not at random. Random splits leak templates and payer-specific phrasing into the test set, which makes drafting quality look better than it is. A useful eval set pairs each held-out denial with its real winning letter and asks graders, or a rubric, whether the model's draft cites the right policy criteria, addresses the stated denial reason and includes the evidence the payer asked for. Related evaluation design is covered in insurance claims AI evaluation.

Where these records come from and how SourceX helps

Appeal letters with outcomes sit inside provider billing offices, physician groups, revenue-cycle service companies and appeal vendors, usually spread across a practice management system, a document store and remittance files. SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match; you describe the data, and every release is approved by the supplying company.

Each dataset is rights-reviewed for ownership and consents, and health records require HIPAA de-identification by Safe Harbor or Expert Determination. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. For the wider record family, see healthcare revenue cycle datasets, healthcare administration AI training data and the industry data hub. You can describe the appeal chains you need to SourceX.

Request denial appeal letter data with outcomes

SourceX sources operational datasets such as documents and finance workflows from US companies, rights-reviews each dataset and delivers it under a license defining records, uses, term and delivery. Pricing and allowed uses are agreed per deal, and nothing is contracted until a supplier agrees. Start a buyer request at SourceX.

Frequently asked questions

Are appeal letters without outcomes still useful?

They help with format, tone and structure, but they cannot teach which arguments win. Use them for style pretraining at most, and keep outcome-labeled chains for SFT selection, preference pairs and evaluation.

Should attachments be included or only the letter text?

Include at least a typed list of attachments per letter. Medical necessity overturns often depend on the evidence attached, so a model trained on letter text alone can learn to cite exhibits it never sees.

How do code-set changes affect multi-year appeal data?

CARC and RARC lists are revised several times a year [1], so a code's meaning or availability can shift across your date range. Record the code-list version per remittance; see code-set revisions in multi-year datasets.

Sources

  1. Commonwealth of Massachusetts (Mass.gov), "835 Payment Advice and EOB/CARC & RARC Lists". https://www.mass.gov/info-details/835-payment-advice
  2. Legal Information Institute, Cornell Law School, "29 CFR 2560.503-1 - Claims procedure". https://www.law.cornell.edu/cfr/text/29/2560.503-1
  3. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  4. Electronic Code of Federal Regulations, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  5. Ertas, "HIPAA-Compliant AI Training Data Guide". https://www.ertas.ai/blog/hipaa-compliant-ai-training-data-guide
  6. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2022). https://arxiv.org/abs/2107.06499v1
  7. Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  8. Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data