Skip to content

Industry-specific operational data

Clinical documentation integrity (CDI) query data for documentation-gap AI

Quick answer

A useful CDI query dataset links four things per encounter: the clinical indicators that triggered the query (labs, vitals, medications, with timestamps), the query text and format, the provider's response, and the coded result, including any MS-DRG, CC/MCC or severity change. Without the indicators you cannot train gap detection; without the response and final codes you cannot measure yield; and without query-format metadata a drafting model can learn leading, noncompliant phrasing.

By SourceX Editorial · Updated

This guide is for NLP and product teams building gap detection, query drafting or query-quality evaluation. It covers record structure, compliance labels, de-identification trade-offs and the questions to ask before licensing. For the wider category view, start with the industry-specific operational data buyer's guide.

What a CDI query record needs to contain

A trainable CDI record is a joined view across the EHR, the CDI worklist and the coding abstract, not a free-text query export. CDI platforms typically store the query, its status and the reviewer, while the indicators live in the EHR and the final codes live in the encoder or abstracting system. Ask suppliers how those three are joined (encounter ID, query ID, account number) and whether the join survives de-identification.

The core field groups are:

  • Trigger context: working diagnoses, lab values with collection times, vital signs, medication orders (for example, IV diuretics or antibiotics), and the note excerpts the CDI specialist cited.
  • Query metadata: query type (concurrent, retrospective, post-bill), format (open-ended, multiple choice, yes/no), template ID, query author role, recipient role, issued and answered timestamps.
  • Response: agree, disagree, clinically undetermined, other, or no response, plus where the provider documented the answer (progress note, addendum, discharge summary).
  • Outcome: codes before and after, MS-DRG before and after, CC/MCC captured, and severity of illness (SOI) and risk of mortality (ROM) where an APR-DRG grouper is used.

Compliant query practice as a labeling requirement

Compliance rules belong in the data as labels, because a model trained on raw query history learns whatever style the organization actually used. The joint ACDIS/AHIMA "Guidelines for Achieving a Compliant Query Practice" (2026 update) is the recommended industry standard for provider queries. It covers query templates, clinical indicators, who and how to query, and query technology, and supersedes the 2019 brief [3].

As of October 2026, confirm with suppliers which edition their templates were built against. The practical point for model builders is that queries should present clinical indicators and reasonable options, including an "other" or "clinically undetermined" choice for multiple-choice formats, without indicating financial impact or steering toward a diagnosis. Yes/no formats have narrower permitted uses under the brief, so label them separately.

Request these labels where the supplier's compliance or audit team produced them:

  • query_format and template_id, so you can filter out ad hoc free text.
  • compliance_review_result (pass, fail, not reviewed) and the failure reason, such as leading language or missing indicators.
  • escalation_flag for queries routed to a physician advisor.

DRG, CC/MCC and severity outcomes

Outcome fields only mean something when they carry the grouper version used. CMS publishes MS-DRG grouper software and Medicare Code Editor files by version each fiscal year. The definitions manual sets assignment from principal and secondary diagnoses, procedures, sex and discharge status, and lists CC/MCC severity exclusions and MCCs that count only if the patient was discharged alive. A secondary diagnosis that is excluded as a CC for a given principal diagnosis produces no DRG shift, and a model that ignores this will overestimate query value.

Ask for the grouper and version recorded with each before-and-after pair, plus the discharge date, so you can regroup consistently. Payer mix matters too: a query that changes an MS-DRG for a Medicare stay may change an APR-DRG SOI level for a Medicaid stay instead. If the supplier also has claim and remittance data, see the medical coding and claims owner page and the prior authorization packets guide for adjacent records.

Avoiding upcoding bias in query histories

Query histories optimized for revenue can teach a model to find financially valuable diagnoses rather than undocumented clinical truth. Signs to look for: queries concentrated on high-weight conditions (sepsis, acute respiratory failure, malnutrition, encephalopathy) with weak supporting indicators, agree rates near 100 percent, or few disagree and undetermined responses.

Counterweights a buyer can request:

  • Clinical validation denials and their outcomes, linking post-bill payer audits to the original query.
  • Internal audit samples with second-reviewer determinations.
  • Queries that resulted in no code change, including those that lowered severity.

Payment integrity teams treat the same signals as risk, which is covered in the healthcare FWA case data guide.

De-identification choices that change model quality

The HIPAA de-identification method determines whether your indicator timelines survive. HIPAA allows Expert Determination or Safe Harbor [1][2]. Safe Harbor removes 18 identifier categories, including all date elements other than year that relate to the individual [1], which collapses "lactate 4.1 at 02:10, query at 09:30, response at 16:45" into undated values. Expert Determination can retain shifted or relative dates when an expert documents low re-identification risk, at higher cost and documentation burden [1].

For gap detection, relative time offsets from admission are usually enough. Ask for the expert's report scope, the date-shifting method, and how free-text notes were scrubbed of names, MRNs and provider identifiers. The Safe Harbor vs Expert Determination guide and the de-identification evidence package checklist go deeper. Full source records are covered on the medical records owner page.

Illustrative CDI query record

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "query_id": "Q-000431",
  "encounter_offset_hours": { "admit": 0, "query_issued": 31.5, "response": 38.2 },
  "query_type": "concurrent",
  "query_format": "multiple_choice",
  "template_id": "RESP-FAIL-v3",
  "clinical_indicators": [
    { "type": "lab", "name": "pO2", "value": 55, "unit": "mmHg", "offset_hours": 29.0 },
    { "type": "vital", "name": "SpO2", "value": 86, "unit": "%", "offset_hours": 29.1 },
    { "type": "order", "name": "BiPAP initiated", "offset_hours": 29.4 }
  ],
  "options": ["Acute hypoxic respiratory failure", "Hypoxemia without respiratory failure", "Other", "Clinically undetermined"],
  "response": "agree_option_1",
  "documented_in": "progress_note",
  "codes_before": ["J18.9"],
  "codes_after": ["J18.9", "J96.01"],
  "grouper": { "name": "MS-DRG", "version": "44.0", "drg_before": "TBD", "drg_after": "TBD" },
  "cc_mcc_captured": "MCC",
  "compliance_review_result": "pass",
  "post_bill_audit": null
}

Buyer checklist for a CDI query dataset

A short pre-license checklist catches the gaps that matter most for this data type:

  • Query, indicator and final-code records join on a stable encounter key after de-identification.
  • Query type and format are labeled; free-text queries are separable from template-based ones.
  • Grouper name and version are stored with before-and-after DRG values.
  • Non-response, disagree and undetermined outcomes are included, not filtered out.
  • Compliance review and denial outcomes are available for at least a sample.
  • De-identification method is documented, with date handling stated.
  • The supplier has authority to license provider-authored notes and query text; see client data held by service providers if the source is an outsourced CDI vendor.
  • A held-out evaluation set was double-reviewed; benchmark test sets commonly carry label errors of a few percent, enough to distort model comparisons [4].

Use the data provider due diligence questionnaire for the broader rights and security review.

How SourceX approaches CDI query data requests

SourceX sources operational datasets from US companies on request rather than from stock, so a request for CDI query data is a description of the records you need, not a catalog pick, and a request may not produce a match. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, and every release is approved by the supplying company. Healthcare teams can also review buyers by industry: healthcare or describe the data on the buyer page.

Request CDI query and response data

SourceX sources operational datasets, including documents and workflow records, from US businesses that hold them, and every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Describe the CDI queries, indicators and outcome fields you need at sourcex.si/buyers.

Frequently asked questions

Can ambient scribe audio substitute for CDI query data?

No. Ambient audio captures the encounter conversation, while CDI query data captures the gap between what was documented and what can be coded, along with the provider's response. They are complementary; see doctor-patient conversation audio for that data type.

Should retrospective and post-bill queries be mixed with concurrent ones?

Keep them labeled and usually split. Post-bill queries carry different rebilling and compliance considerations, and their indicator timing differs, so mixing them can blur what a concurrent gap-detection model should flag.

Is synthetic query text enough for drafting SFT?

Synthetic queries can teach format, but they lack real provider response behavior and DRG outcomes. Real query-response pairs are needed to model yield and to evaluate whether drafted queries stay compliant. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Sources

  1. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  2. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  3. ACDIS/AHIMA, "Guidelines for Achieving a Compliant Query Practice—2026 Update" (2026). https://acdis.org/resources/acdisahima-guidelines-achieving-compliant-query-practice%E2%80%942026-update
  4. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data