Industry-specific operational data
Risk adjustment chart review data for HCC coding AI
Quick answer
Useful HCC coding training data is not a list of diagnosis codes. It is a set of de-identified charts in which certified coders marked each supported HCC, linked it to an evidence span that meets MEAT-style support criteria, recorded deletes as well as adds, tagged the CMS-HCC model version, and, where available, attached a later audit or second-level review outcome. Without the evidence spans and the deletes, a model learns to capture codes, not to defend them.
By SourceX Editorial · Updated
What a risk adjustment chart review record should contain
A training-grade record joins one encounter's clinical documentation to the coder's decision about each candidate condition, with the text that justifies it. The unit is usually a chart (all encounters for one member in a service year) broken into encounter documents such as progress notes, H&Ps, discharge summaries and problem lists. Each candidate ICD-10-CM code then carries a decision: confirmed, added, deleted, or queried.
The highest-value label is the evidence span: character offsets into the note that show the condition was Monitored, Evaluated, Assessed or Treated on that date of service. A code supported only by a problem-list entry or a historical mention is exactly the error audits look for, so the span lets your model learn the difference.
Buyers building adjacent payer models should compare scope with utilization management review records and clinical registry abstraction records, which share the chart-plus-span pattern but answer different questions.
Why model version tags matter for CMS-HCC V24 and V28 labels
Every HCC label should state which CMS-HCC model it maps to, because the same ICD-10-CM code can map to a different HCC, or to none, across versions. CMS phased in the 2024 CMS-HCC model (commonly called V28) over several payment years, blending it with the 2020 model (V24); verify the year-by-year blend and the current model weights against the CMS Advance Notice and Rate Announcement for each payment year before relying on any label set.
For training data, that means a 2023 chart review labeled under V24 is not wrong, but it is stale for a V28 model unless the supplier keeps the raw ICD-10-CM decision alongside the HCC roll-up. Ask for the code-level decision, the HCC under each model version, and the payment year. If a supplier can only provide rolled-up HCCs, you cannot re-map history when CMS recalibrates again.
How RADV audit exposure changes what counts as ground truth
Ground truth for risk adjustment AI should be the decision that survives validation, not the decision that was submitted. CMS's 2023 Risk Adjustment Data Validation (RADV) final rule addressed contract-level audit methodology, including extrapolation of audit findings. That rule has since been challenged in court, so confirm its current status with counsel as of October 2026 rather than assuming a settled methodology.
Either way, unsupported codes carry financial risk, which makes three label sources valuable: second-level coder QA results, internal mock-RADV findings, and actual RADV medical record review outcomes. Treat each as a separate field with its own provenance; an internal QA "agree" is weaker evidence than an external validation result. The verification method is covered in verifying outcome labels in operational records.
Adds, deletes and two-way review: avoiding a biased label set
A dataset built only from add-only chart reviews teaches a model to find more codes and never to remove unsupported ones. One-sided reviews that only add codes have drawn enforcement attention in Medicare Advantage, so ask each supplier whether their reviews were two-way (adds and deletes submitted) and what share of decisions were deletes. Verify that history with your own counsel; this page does not characterize any particular case.
Ask also how the review queue was selected. Charts pulled by a suspecting algorithm over-represent high-value HCCs such as diabetes with complications, CHF and major depression, which skews both prevalence and the evidence patterns your model sees. A dataset bias audit should compare the label distribution with a random chart sample from the same population.
HIPAA de-identification route for full medical charts
Charts are full medical records, so de-identification is the gating step, and Expert Determination is the more likely route when you need dates of service and free text preserved. HHS OCR describes two methods under 45 CFR 164.514: Safe Harbor, which removes 18 categories of identifiers, and Expert Determination, in which a qualified expert concludes that re-identification risk is very small [1][2]. Safe Harbor strips dates more specific than the year, which breaks date-of-service logic that risk adjustment depends on.
A limited data set under a data use agreement is a different instrument with its own conditions, defined in 164.514(e) [2]. Whichever route applies, ask for the expert's report scope, the free-text scrubbing method (NER plus surrogate substitution is common), and the residual-risk statement. Free-text clinical notes leak identifiers through family names, facility names and rare-condition combinations, so no method is perfect.
Buyer checklist for an HCC chart review dataset
Use this request template to describe the data you need before talking to suppliers.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | What to specify | Why it matters |
|---|---|---|
| chart_id, encounter_id | Stable pseudonymous keys | Joins encounters without identity |
| doc_type, dos | Note type and shifted or retained date of service | MEAT is date-of-service specific |
| icd10cm_code | Code-level decision, not just HCC | Lets you re-map across model versions |
| hcc_v24, hcc_v28, payment_year | HCC under each model | Avoids stale labels |
| decision | confirmed / added / deleted / queried | Two-way review signal |
| evidence_spans | Character offsets plus MEAT element | Core training signal |
| reviewer_credential, qa_level | CRC/CPC, first or second level | Label reliability |
| audit_outcome, audit_source | Internal QA, mock RADV, RADV | Ground-truth strength |
| deid_method | Safe Harbor or Expert Determination, report date | HIPAA basis |
Before acceptance, double-code a random 5 to 10 percent sample and compute inter-annotator agreement per HCC family, following annotation quality audit practice. Document the result in a Data Card that records sources, annotation method and intended use [3].
Rights, licensing and AI use of payer chart data
Chart review data usually involves at least three parties: the provider that authored the note, the health plan that holds it for risk adjustment, and the coding vendor that produced the labels. Licensing must establish which party can authorize AI training use, whether business associate agreements permit it, and whether de-identified data is released under the plan's own terms. In-house reviewers should follow the counsel review guide for AI data licenses.
Related payer and coding datasets are covered on the medical coding and claims and medical records pages, and buyer context for the sector sits under healthcare administration buyers. For other operational datasets in this cluster, start at the industry-specific operational data hub or the AI data guide.
How SourceX handles chart review data requests
SourceX sources operational datasets from US companies on request; it does not hold them in stock, and a request does not guarantee a match. You describe the data, not the business, and SourceX looks for US organizations that hold it; every release is approved by the supplying company. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, and personal details are removed or replaced before delivery, with the method recorded and a sample checked. You can describe the HCC review data you need to SourceX.
The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, through private, access-controlled workflows. SourceX does not train models and does not publish prices.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Request risk adjustment chart review data
If you are training or validating an HCC coding model and need charts with evidence spans, adds and deletes, and audit outcomes, describe the dataset and intended use. SourceX will assess whether a US supplier holds it and can license it. Start a buyer request.
Sources
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- Pushkarna, Zaldivar, Kjartansson (Google Research), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.