Industry-specific operational data
Claim denial prediction training data: matched 837 claims and 835 remittance outcomes
Quick answer
Claim denial prediction training data is a set of submitted X12 837P or 837I claims, each joined to the 835 remittance lines that settled it, with the outcome expressed as claim status plus CARC group and reason codes and RARC remarks. To train a pre-submission risk score that holds up in production, buyers need reliable claim-to-remit linkage, a defined label taxonomy, many months across several payers, time-based splits, and features stripped of anything posted after adjudication.
By SourceX Editorial · Updated
This page covers the modeling questions a data scientist has to answer before licensing: how claims and remits link, what the label really means, how to split and validate, and what to request from a supplier. For the record types a revenue-cycle dataset can contain, see the owner page on healthcare revenue cycle datasets; for the wider cluster, see industry-specific operational data for AI.
How 837 claims link to 835 remittances
Claims link to remittances through control numbers, not through patient names or dates, and the join is where most denial datasets quietly break. In the X12 005010 transactions, the provider's patient control number sent in the 837 CLM01 is generally echoed back in the 835 CLP01, while the payer assigns its own claim control number in CLP07; verify the exact element usage against the implementation guides and each payer's companion guide. Service lines can be tied back through line item control numbers (REF*6R) where the submitter populates them.
Several real-world patterns complicate a one-to-one join:
- Reversals and corrections. A payer may reverse a prior payment (claim status code 22 in CLP02) and re-adjudicate, so one claim can carry three or more 835 claim-payment loops over time.
- Corrected and voided claims. A frequency code of 7 (replacement) or 8 (void) in CLM05-3 creates a new claim that supersedes the original; the original's denial should not be counted as a final outcome if the replacement paid.
- Split and partial payments. Payers may split a claim across remits or pay some lines and deny others, so a claim-level label hides line-level reality.
- Crossover and secondary payers. Coordination of benefits produces a second 835 from another payer, with OA adjustments that are not denials.
- Front-end rejections. Claims rejected at the clearinghouse or payer front end appear in 999 or 277CA acknowledgments and never reach an 835.
Ask the supplier how they built the linkage, what share of claims matched at least one remit, and how unmatched claims are flagged rather than dropped. A dataset that silently drops unmatched claims removes exactly the rejections a scrubber model should learn to catch.
Defining the denial label from CARC, RARC and group codes
The label should be built from adjustment codes, not from a single "denied" flag, because the same paid-zero outcome can mean very different things. The 835 CAS segment carries a group code (CO for contractual obligation, PR for patient responsibility, OA for other adjustments, PI for payer-initiated reductions) paired with a Claim Adjustment Reason Code, and RARCs add detail at the claim or line level [1][2]. The CARC and RARC lists are updated several times a year, and payers such as Medicare contractors implement those updates on a schedule, so codes are added, revised and retired over the life of any multi-year dataset [1].
The CAQH CORE operating rule for CARCs and RARCs constrains which code combinations should be used for defined business scenarios, which makes label mapping more consistent across compliant payers [2]. Payer behavior still varies: a state Medicaid program, for example, publishes its own 835 guidance and the code lists it uses [3]. Expect a long tail of payer-specific usage that needs its own mapping.
A practical taxonomy separates at least four outcomes:
- Rejected before adjudication (999/277CA rejection, no 835).
- Denied (claim or line paid zero with a CO or PI reason that signals a correctable or non-covered issue, such as missing information or authorization).
- Paid with contractual adjustment only (CO-45-style fee-schedule reductions are normal, not denials).
- Patient responsibility (PR deductible, coinsurance or copay amounts, which are not payer denials).
Keep the raw CAS triplets (group, reason, amount) and RARCs in the delivered data so the label can be rebuilt as your taxonomy changes. Label quality matters even for operational records: audits of public benchmarks found an average test-set label error rate of at least 3.3%, enough to change model rankings [8]. For a deeper treatment, see verifying outcome fields as ground truth.
Features available at submission time
A pre-submission model may only use fields that exist when the claim is sent, so the feature list should be derived from the 837 and pre-billing context, never from the remit. Typical features include payer identifier, claim type (837P versus 837I), place of service, type of bill, billing and rendering provider taxonomy, CPT/HCPCS codes and modifiers, ICD-10-CM diagnosis pointers, units, charge amounts, prior authorization reference numbers, and days from service to submission.
Some of the strongest signals are relational: code pairs that trigger edits, missing authorization references on services that payers commonly gate, and modifier combinations. Prior authorization context deserves its own data request; see prior authorization submission packets and payer decisions.
Leakage traps in remittance-derived data
Leakage in denial data usually comes from fields written after adjudication that ride along in the practice-management export. Remove or quarantine:
- Adjustment and payment postings, write-off reason codes and balance fields.
- Appeal-filed flags, appeal dates and resubmission counts.
- Claim status updates from 276/277 inquiries made after submission.
- Account-level fields that summarize later events, such as "days in AR" or collector queue assignments.
- Corrected-claim frequency codes on the record you are scoring, if the correction was triggered by the denial.
Ask the supplier for a field dictionary with a "known at" timestamp for every column. The general audit method is covered in target leakage in licensed tabular data. Appeal and follow-up records are valuable for other models, such as appeal-drafting training data with outcomes and AR follow-up histories for claim status agents, but they belong outside the denial prediction feature set.
Temporal splits and payer policy drift
Split by submission or service date, not at random, because payer edits, fee schedules and coverage policies change and a random split lets the model learn from the future. Research on temporal tabular data with regime changes evaluates models forward in time, because shuffled splits mix later periods into training and can overstate performance [4]. In revenue cycle data, drift also arrives on a schedule: annual CPT and ICD-10 code updates, quarterly edit updates, and CARC list revisions [1].
A defensible design trains on an earlier window, validates on the following months, and tests on the most recent period, with remit lag accounted for. Claims submitted near the end of the extract often have no 835 yet, so define a maturity cutoff (for example, only label claims with enough elapsed time for final adjudication) or treat them as censored.
Coverage tables matter as much as volume. If one large payer contributes most rows and has an unusually high or low denial rate, a pooled AUC can look strong while per-payer performance is weak. Ask for row counts and denial base rates by payer, claim type, specialty and month, and evaluate per payer.
De-identification that keeps the features you need
The de-identification method decides which time features survive, so choose it with the model in mind. Under HIPAA Safe Harbor, all elements of dates directly related to an individual except the year must be removed, along with the other listed identifiers [5][6]. That eliminates service dates, submission dates and remit dates at day or month granularity, which are exactly the inputs for timely-filing and lag features.
Expert Determination lets a qualified expert certify that re-identification risk is very small under specified conditions, which can permit coarsened dates, date offsets or relative intervals [5][6]. Industry commentary describes it as the usual route when granular data, such as month-level dates, has to be kept [7]. A limited data set under a data use agreement is a separate HIPAA pathway with its own conditions [6]. The trade-offs are covered in HIPAA Safe Harbor vs Expert Determination for AI training.
Interval features (days from service to submission, days to remit) can often be computed before de-identification and delivered as integers, which keeps signal while dropping calendar dates. Confirm with the expert whether derived intervals are within the certified scope.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Request template for matched claim and remittance data
Use a structured request so suppliers can say quickly whether they hold the data and in what shape.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Requirement | What to specify |
|---|---|
| Transactions | 837P and/or 837I claims; all 835 claim-payment loops per claim, including reversals; 999/277CA rejections if available |
| Linkage | Join keys used (patient control number, payer claim control number, line item control number); match rate; unmatched claims flagged, not dropped |
| Label fields | Claim status code, CAS group/reason/amount triplets per line, RARCs, paid and allowed amounts |
| Payer coverage | Payer mix table with row counts and denial base rates per payer and month |
| Time span | Enough consecutive months for train/validation/test by date, plus a maturity cutoff for unadjudicated claims |
| Feature timing | Data dictionary with a "known at" marker per column; post-adjudication fields isolated |
| De-identification | HIPAA method (Safe Harbor or Expert Determination), date handling, and how derived intervals were computed |
| Format | Parquet or CSV with one claim table, one line table and one adjustment table, keyed consistently |
| Code versions | CARC/RARC and CPT/ICD code-set versions in effect for each period |
An illustrative adjustment row might look like this, with invented values:
{"claim_key": "c_000184", "line_no": 2, "remit_seq": 1, "clp_status": "1",
"cas_group": "CO", "carc": "197", "adj_amount": 240.00,
"rarc": ["N54"], "days_service_to_submit": 9, "days_submit_to_remit": 27}
How SourceX approaches denial prediction requests
SourceX sources operational datasets, including finance and other workflow records, from US companies on request, and manages licensing and ongoing purchases; it does not hold inventory, and a request does not guarantee a match. You describe the data you need, such as the template above, and SourceX looks for US businesses that hold it. Every dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery, and every release is approved by the supplying company.
Health records require HIPAA de-identification by Safe Harbor or Expert Determination. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement. You can describe a matched claims and remittance request to SourceX, and related coding and claims records are described on licensing medical coding and claims for AI and healthcare administration AI training data.
Get matched claim and remittance outcomes for your denial model
SourceX sources operational datasets from US companies on request and manages the license, from finding a supplier through assessment, agreement, transaction and ongoing purchases. Each dataset is rights-reviewed, de-identified as HIPAA requires, and released only with supplier approval. Describe the claims data your model needs.
Frequently asked questions
Can I train a denial model on 835 data alone?
Not for pre-submission scoring. The 835 carries outcomes and adjustment codes but only partial claim content, so you cannot reconstruct the full feature set the model sees at submission. An 835-only set is useful for denial-reason analytics and for evaluating code mapping.
Should rejections and denials share one label?
Usually not. Front-end rejections reported in 999 or 277CA acknowledgments reflect format and eligibility edits, while payer denials reflect adjudication rules. Keeping them separate lets a scrubber model and a denial model be tuned and evaluated independently.
How do I handle retired CARCs?
Map each code to your taxonomy using the code list in effect for the remit date, and keep the raw code [1]. Store the mapping version with the model so retraining reproduces the same labels.
Sources
- Centers for Medicare & Medicaid Services, "Remittance Advice Remark Code and Claim Adjustment Reason Code Update (JA6229)". https://www.cms.gov/Medicare/Medicare-Contracting/ContractorLearningResources/downloads//JA6229.pdf
- CAQH CORE, "CAQH CORE Phase III CARCs and RARCs 835 Rule". https://www.caqh.org/sites/default/files/core/phase-iii/policy-rules/CARCsRARCs_835_Rule.pdf
- Commonwealth of Massachusetts (MassHealth), "835 Payment Advice and EOB/CARC & RARC Lists". https://www.mass.gov/info-details/835-payment-advice
- arXiv, "Online learning techniques for prediction of temporal tabular datasets with regime changes" (2023). https://arxiv.org/pdf/2301.00790
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- Censinet, "Safe Harbor vs Expert Determination for PHI". https://censinet.com/perspectives/safe-harbor-vs-expert-determination-phi
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.