Skip to content

Industry-specific operational data

Healthcare fraud, waste and abuse case data for payment integrity AI

Quick answer

The most useful healthcare fraud detection training data is claims linked to closed special investigations unit (SIU) and prepayment review outcomes: each flagged claim or provider labeled confirmed fraud, waste, abuse or cleared, with the analyst's notes, the evidence reviewed and the recovered amount. Public options mostly use exclusion lists or teaching sets as proxy labels, which capture only a narrow, skewed slice of schemes. Licensed case files from payers or payment integrity vendors give you real negatives, graded outcomes and the reasoning a triage or summarization model must learn.

By SourceX Editorial · Updated

Why exclusion-list labels fail payment integrity models

Proxy labels from the OIG exclusion list measure who was sanctioned, not which claims were improper. Medicare machine learning studies typically label providers as fraudulent if they appear on the List of Excluded Individuals and Entities (LEIE), and one such study reports positive rates of only 0.038% to 0.074%, with labeled providers skewed toward overt, already-prosecuted schemes [1]. Exclusion status also changes over time as parties are reinstated, so a label pulled last year may no longer hold.

That makes every unlisted provider an "unknown," not a confirmed negative. Researchers have responded by treating enforcement labels as positive-unlabeled data or by abandoning them for unsupervised peer comparison [2]. Many public projects instead train on a Kaggle provider-fraud teaching set, which is common for initial modeling but may not meet production investigation standards [4]. For a production model that routes claims to prepayment review or opens SIU cases, you need labels that came out of an investigation. The broader question of when an operational outcome field is a trustworthy label is covered in verifying outcome labels in operational records.

What a licensable FWA case file contains

A usable FWA record joins the claim lines, the trigger that flagged them, the investigation and the final disposition. Most SIU case management systems and payment integrity platforms already hold these pieces; the work is in linking them with stable pseudonymous keys.

  • Claim lines: 837P/837I-derived fields such as CPT/HCPCS codes, modifiers, ICD-10-CM diagnoses, units, place of service, billed and allowed amounts, and service and paid dates.
  • Trigger: the rule, model score, tip line referral or data-mining lead that opened the review, with its date. This is what lets you measure the lift of a new model over the existing one.
  • Review artifacts: medical record request dates, records received, reviewer specialty, and the coding or medical necessity findings.
  • SIU notes: the free-text narrative, interview summaries and investigator conclusions, which feed case summarization and triage copilots.
  • Disposition and money: outcome category, overpayment identified, amount recovered, and whether the result was education, a corrective action plan, payment suspension or prepayment edit.

Prepayment review decisions are a distinct and valuable slice: each record pairs a held claim with a pay, deny or adjust decision and a reason, before money moves. This is different from denial prediction on remittances, which the claim denial prediction training data page covers.

Defining fraud, waste and abuse as separate labels

Treat fraud, waste and abuse as three labels with different evidentiary meanings, plus an explicit "cleared." Program integrity teams generally reserve fraud for knowing misrepresentation to obtain payment, waste for overuse or misuse of resources without criminal intent, and abuse for practices that raise costs without proof of knowing misrepresentation. Federal program definitions follow the same lines, so check which definition the supplier's SIU policy actually applied.

These distinctions matter for modeling. A provider whose upcoding is resolved through education is an abuse or waste case, not a fraud case, and collapsing them teaches a model that billing pattern equals intent. Ask the data holder for its label taxonomy, who assigned each label, and whether "fraud" means substantiated by the SIU, referred, or adjudicated.

Illustrative record schema

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
case_idCASE-7F3A91 (pseudonymous)Join key across claims and notes
provider_keyPRV-0C82 (pseudonymized NPI)Protects unproven-allegation subjects
provider_specialtyPhysical therapyPeer-group comparison
trigger_typerule: units per day above peer 99th percentileMeasures baseline detector
trigger_date2024-03-11Time-based splits
claims_in_scope412 lines, 37 membersCase size
review_typepostpayment medical record reviewSeparates pre- and postpayment
findingtime-based codes billed beyond documented minutesFeature and rationale target
dispositionabuse: education plus overpayment demandGraded label
overpayment_identified_usd18,240Cost-weighted evaluation
recovered_usd15,900Realized value
case_closed_date2024-09-02Verification latency
siu_noteFree text, identifiers replacedSummarization and triage training

FWA case data carries two distinct risks: member PHI and allegations about providers who were never found liable. Member data falls under HIPAA, so the file should be de-identified through Safe Harbor or Expert Determination [5], or, where the purpose is research, public health or health care operations, handled as a limited data set under a data use agreement per 45 CFR 164.514(e) [6]. Free-text SIU notes are the weak point, because they contain member names, addresses and phone numbers that structured scrubbing misses. If an Expert Determination is offered, our guide to reviewing a HIPAA Expert Determination report lists what to check.

Pseudonymize providers too. A cleared provider in a labeled fraud dataset is an unproven accusation attached to a real NPI, which creates defamation and contract exposure for both sides. Ask suppliers to limit files to closed cases and to exclude open investigations and anything referred to law enforcement or a Medicaid Fraud Control Unit that is still pending.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Building evaluation around imbalance and drift

Plan a time-based evaluation from day one, because fraud schemes adapt and labels arrive late. Fraud detection research describes the combined problem of concept drift, extreme class imbalance and delayed supervision (analogous to credit card fraud), where confirmed labels only exist for the small fraction of cases that were investigated [3]. In healthcare that latency can stretch across months between trigger date and case closure.

  • Split by trigger date, not randomly, so a model never sees closed outcomes from after its prediction date.
  • Report precision at review capacity (for example, top 200 cases per week) and dollar-weighted recall, not ROC AUC alone.
  • Stratify by specialty and scheme type, since a DME billing ring and physical therapy unit inflation look nothing alike.
  • Keep investigated-and-cleared cases as hard negatives; they are the examples that reduce provider abrasion.
  • Record selection bias: only flagged claims were reviewed, so the dataset reflects the old detector's blind spots.

For transaction-level fraud outside healthcare, see fraud-labeled transaction data; for analyst narratives in banking fraud operations, see fraud investigation case notes and analyst decisions.

Buyer checklist before licensing FWA case data

Confirm label provenance, scope and permitted use before any pricing discussion.

  1. Which systems produced the labels, and what does each disposition value mean in the supplier's policy?
  2. Are cleared cases included, and in what proportion to substantiated ones?
  3. Is the date range long enough to span at least one policy or coding change (for example, an annual CPT update)?
  4. Which lines of business (Medicare Advantage, Medicaid managed care, commercial) are covered, and do contracts with the plan sponsor permit secondary use?
  5. How were member and provider identifiers removed or replaced in both structured fields and free text, and was a sample checked?
  6. Does the license specify records, allowed uses, term and delivery?

How SourceX approaches FWA case data requests

SourceX sources operational datasets from US companies on request; it does not hold FWA data in stock, and a request does not guarantee a match. For a request like this, SourceX looks for US businesses that hold the described case and claims data, assesses the data and licensing permissions, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and health records require HIPAA de-identification by Safe Harbor or Expert Determination. Buyers can describe the data they need on the SourceX buyers page. Related context sits in claims administration buyers, licensing medical coding and claims data, healthcare administration buyers and the industry-specific operational data hub.

Request healthcare FWA case data

Describe the dispositions, claim types, lines of business and date range your payment integrity model needs, not the companies you think hold them. SourceX looks for US businesses with matching data and handles the licensing process, with personal details removed or replaced before delivery and nothing contracted until a supplier agrees. Start a buyer request at SourceX.

Sources

  1. medRxiv, "Utilization Analysis and Fraud Detection in Medicare via Machine Learning" (2024). https://www.medrxiv.org/content/10.1101/2024.12.30.24319784.full.pdf
  2. arXiv, "Unsupervised Machine Learning for Explainable Health Care Fraud Detection" (2022). https://arxiv.org/pdf/2211.02927
  3. IEEE Transactions on Neural Networks and Learning Systems (Dal Pozzolo et al.), "Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy" (2017). https://boracchi.faculty.polimi.it/docs/2017_FraudsTNNLS.pdf
  4. NYC Data Science Academy, "Healthcare Fraud Detection (student project)". https://nycdatascience.com/blog/student-works/machine-learning/healthcare-fraud-detection-2
  5. U.S. Department of Health and Human Services, "Guidance Regarding Methods for De-identification of Protected Health Information". https://www.hhs.gov/guidance/document/de-identification-guidance
  6. eCFR, Office of the Federal Register, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data