Evaluation and benchmarking datasets
Medical coding AI evaluation sets: audited codes as ground truth
Quick answer
A medical coding AI evaluation dataset should pair de-identified encounter documentation with the final codes that survived a coding audit, not the first-pass codes a coder or vendor submitted. Score against ICD-10-CM, ICD-10-PCS and CPT/HCPCS codes pinned to the code-set version in effect on the date of service, slice results by specialty, setting and payer, and keep coder-auditor disagreements as a labeled signal. Public code-lookup benchmarks cannot tell you how an autonomous coder performs on your claim mix.
By SourceX Editorial · Updated
Why public coding benchmarks miss production accuracy
Public benchmarks measure code recall from descriptions or from hospital discharge notes, while production coding is a documentation-to-claim task with payer rules attached. A Mount Sinai study in NEJM AI found that GPT-4, the best model tested, reproduced the exact code from its description in under half of cases (45.9% ICD-9-CM, 33.9% ICD-10-CM, 49.8% CPT) [1]. A follow-up preprint reported large gains when a retriever and reranker were added, but on 100 single-term conditions, and its author flags the need for realistic cases [2].
Neither setup resembles a coder reading an operative report, applying the Official Guidelines for Coding and Reporting, sequencing a principal diagnosis and adding modifiers 25 or 59. MIMIC-derived ICD coding sets are useful for research, but they reflect one academic medical center's inpatient and ICU population, mostly older code sets and no payer adjudication. For an agent that will touch claims, you need held-out encounters from the settings you sell into; see private evaluation sets vs public benchmarks for the general case.
Post-audit codes as the gold label
The gold label should be the code set after a credentialed auditor (for example, a CPC or CCS reviewer) has reviewed the encounter, because first-pass codes carry the same errors you are trying to measure. Production coding data mixes coder output, CDI query responses, edits from claim scrubbers and, sometimes, payer-driven corrections after a denial. If you score against the submitted claim, a model that matches a coder's undercoding looks accurate.
Label noise matters even in curated benchmarks: one audit of 10 widely used test sets estimated an average label error rate of at least 3.3% [5]. Expert-validated golden sets are now a common vendor practice [6], but ask who validated, against which guideline year, and how disagreements were resolved. The same logic applies to verifying outcome fields as reliable labels.
Keep three code layers per encounter where the supplier has them:
- Original: codes as first assigned by the coder or the incumbent engine.
- Audited final: codes after audit, with a change reason (added, deleted, resequenced, modifier change, specificity change).
- Adjudicated: remittance outcome (CARC/RARC codes from the 835) where available, which separates coding errors from payer policy denials.
The original-to-audited delta is your coder-auditor disagreement dataset: it shows where humans themselves fail and where an autonomous coder must beat a realistic baseline.
Code-set versions and date-of-service pinning
Every gold code must be valid for the date of service, because ICD-10-CM changes each October and sometimes in April. As of October 2026, CMS lists the FY 2027 ICD-10-CM files for encounters from October 1, 2026 through September 30, 2027, and ICD-10-PCS is published separately [3]. NCHS lists the October 1, 2026 release as FY27, replacing the April 1, 2026 FY26 update [4].
Store a code_set_version field per encounter and score against that version, or a model that emits a newly created code for a 2025 discharge is marked wrong for the wrong reason. CPT is maintained by the AMA under license, so confirm your own license covers using CPT descriptors in prompts and scoring tools before building the harness.
Slices that expose failure modes
An evaluation set is only as informative as its slices, so allocate encounters by specialty, setting, payer and documentation type rather than sampling at random. Random samples over-represent routine E/M office visits and hide failures in interventional radiology, multi-procedure surgery or inpatient DRG assignment. Use stratified allocation for rare and high-risk cases to oversample:
- Inpatient vs outpatient vs professional fee (ICD-10-PCS and MS-DRG apply only to inpatient facility claims).
- Specialty: cardiology, orthopedics, oncology, emergency medicine, behavioral health.
- Payer: Medicare fee-for-service, Medicare Advantage, Medicaid, commercial, because NCCI edits and local coverage determinations differ.
- Documentation: scanned faxes and outside records, templated EHR notes, dictated operative reports, addenda.
- Risk adjustment: encounters where an HCC-relevant diagnosis is supported or unsupported by MEAT evidence.
Scoring: exact match, hierarchy and claim impact
Report exact code match as the headline metric, then add hierarchical partial credit and dollar-weighted impact so you can see whether errors are harmless or costly. Exact-match precision and recall at the code level show over- and undercoding separately. Hierarchical credit (same category, wrong fourth to seventh character) shows whether the model reads specificity such as laterality or encounter type.
Claim-level metrics matter most to buyers of autonomous coding: principal diagnosis accuracy, MS-DRG match, E/M level match, modifier accuracy and the share of encounters a model would route to a human. Freeze the test split and keep it out of any fine-tuning pipeline; contaminated test data can make a benchmark obsolete quickly [9]. See contamination-resistant evaluation design.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
encounter_id | enc_7f3a91 (pseudonymous) | Joins notes, codes and remittance without identifiers |
setting | outpatient_hospital | Selects code sets and edits |
specialty | orthopedic_surgery | Slice key |
payer_class | medicare_advantage | Slice key; NCCI and policy differences |
date_of_service_shift | +14 days, month retained | Keeps code-set version resolvable while shifting the day |
code_set_version | ICD-10-CM FY2026 | Scoring version |
documents | operative report, H&P, discharge note (text) | Model input |
codes_original | M17.11; 27447 | Baseline for disagreement |
codes_audited | M17.11; 27447-RT | Gold label |
audit_change_reason | modifier_added | Error taxonomy |
auditor_credential | CPC | Label provenance |
remit_outcome | paid / CARC code | Separates coding from payer policy |
De-identification routes for encounter notes
Encounter notes are protected health information, so an evaluation set must be de-identified under HIPAA or shared as a limited data set under a data use agreement. HIPAA offers two de-identification methods: Safe Harbor, which removes 18 identifier types including all date elements except year, and Expert Determination, where a qualified expert finds the re-identification risk very small [7][8]. A limited data set may keep dates and some geography but requires a data use agreement [8]; see HIPAA limited data sets and DUAs.
Safe Harbor date removal can break code-set pinning and length-of-stay logic, which is why buyers often consider Expert Determination with controlled date shifting for coding evaluation. Free-text notes also leak identifiers that structured scrubbing misses, such as names in dictation headers or MRNs in scanned faxes. Ask for the expert's report and review it with this Expert Determination checklist.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Buyer request checklist for coding audit data
Describe the evaluation set in terms a supplier's coding compliance team can check against their audit logs. A precise request gets a faster yes or no on feasibility than "medical billing data."
- Settings, specialties and payer classes, with target counts per slice.
- Gold standard: audited final codes, auditor credential, audit sampling method and guideline year.
- Code layers wanted: original, audited, adjudicated (835 remittance).
- Code-set coverage: ICD-10-CM, ICD-10-PCS, CPT/HCPCS, modifiers, MS-DRG, HCC.
- Note types and formats: HL7 v2, C-CDA, FHIR DocumentReference, PDF scans.
- De-identification route, date handling and residual-risk testing.
- Permitted use: evaluation only, held out from training, with access controls.
How SourceX sources coding evaluation sets
SourceX sources operational datasets from US companies on request and manages the licensing process; it holds no inventory, and a request does not guarantee a match. Buyers describe the data, such as audited encounters for a given specialty and setting, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, and each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. You can describe the coding evaluation set you need.
For training on coding and claims data rather than evaluation, see licensing medical coding and claims data, the broader healthcare administration AI data page, and the evaluation datasets hub.
Request a medical coding evaluation dataset
If your team needs held-out encounters with audited codes to test a coding or billing agent, describe the settings, specialties, code sets and de-identification route you need. SourceX looks for US companies that hold that data, assesses rights and permissions, and nothing is contracted until a supplier agrees. Start a buyer request at sourcex.si/buyers.
Sources
- Soroush et al., "Large Language Models Are Poor Medical Coders — Benchmarking of Medical Code Querying" (2024). https://ai.nejm.org/doi/full/10.1056/AIdbp2300040
- arXiv, "Large language models are good medical coders, if provided with tools" (2024). https://arxiv.org/pdf/2407.12849
- Centers for Medicare & Medicaid Services, "ICD-10 Codes" (2026). https://www.cms.gov/medicare/coding-billing/ICD-10-codes
- CDC National Center for Health Statistics, "Comprehensive Listing of ICD-10-CM Files" (2026). https://www.cdc.gov/nchs/icd/comprehensive-listing-of-icd-10-cm-files.htm
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Sigma AI, "Expert-validated golden datasets for AI evaluation". https://sigma.ai/?p=26358
- HHS Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- eCFR, Office of the Federal Register, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- arXiv (White et al.), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.