Privacy, de-identification and sensitive data
HIPAA-compliant AI training data: what the label must mean before you license it
Quick answer
"HIPAA-compliant" is not a certification, and no regulator issues it. For licensed training data, the only defensible meaning is that the records meet the HIPAA de-identification standard in 45 CFR 164.514(a) through one of two documented methods, Safe Harbor or Expert Determination, so they are no longer protected health information (PHI) [1][2]. Before you license, ask which method was applied, by whom, to which fields, and what evidence proves it.
By SourceX Editorial · Updated
What "HIPAA-compliant" should mean on a dataset listing
A HIPAA-compliant training dataset is one that was de-identified under 164.514(b), with written records showing how. The regulation treats health information as not individually identifiable when there is no reasonable basis to believe it can identify a person, and it gives two ways to meet that standard [1]. Properly de-identified data falls outside HIPAA's use and disclosure restrictions, which is why it is the usual route for commercial AI development; a limited data set does not, because it is still PHI [3]. A vendor that says "HIPAA-compliant" but cannot name the method is usually describing its hosting and business associate agreements, not the data.
Three labels get confused in health data deals:
- De-identified (Safe Harbor or Expert Determination): not PHI, so HIPAA no longer restricts the licensee's use [1][3].
- Limited data set: PHI with 16 categories of direct identifiers removed; dates, city, state and ZIP code may remain, and use is limited to research, public health or health care operations under a data use agreement [1]. The trade-offs are covered in limited data sets and DUAs for AI development.
- "HIPAA-secure" or "HIPAA-ready": a claim about infrastructure and contracts, silent on whether the records identify anyone.
The legal status of each label outside HIPAA is compared in de-identified, anonymized, pseudonymized and aggregated, and the wider cluster lives in the buyer's guide to de-identified data for AI training.
Safe Harbor or Expert Determination: which features survive
The method decides which training signal survives: Safe Harbor is a fixed removal rule, while Expert Determination is a documented risk analysis. Safe Harbor requires removing 18 identifier types of the patient and of relatives, employers and household members, including every date element except year that relates directly to the individual, and the covered entity must have no actual knowledge that what remains could identify someone [1][2]. ZIP codes survive only as the first three digits, and only where that area holds more than 20,000 people; ages over 89 collapse into a single 90-or-older category [1][2].
Expert Determination lets a qualified expert apply statistical and scientific methods, conclude that the risk of identification by an anticipated recipient is very small, and document the methods and results [1]. That route can keep date offsets, service intervals or finer geography that Safe Harbor strips, which is why some practitioners describe Safe Harbor as a poor fit for high-fidelity ML [5]. The catch is that the determination is bound to the data, the recipient and the conditions the expert assumed. The buyer decision is set out in Safe Harbor vs Expert Determination for AI training, and the feature losses in dates, ages and ZIP codes under Safe Harbor.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Training use | What Safe Harbor removes | Method usually needed |
|---|---|---|
| Denial prediction on X12 837 claims and 835 remittances with CARC/RARC codes | Service, submission and payment dates cut to year, so days-to-pay and timely-filing features disappear | Expert Determination with date offsets |
| No-show and scheduling optimization | Appointment dates, lead times and patient ZIP below three digits | Expert Determination |
| Coding assistant (ICD-10-CM, CPT, HCPCS) | Names, record numbers and dates in notes; the codes themselves remain | Safe Harbor often workable |
| Contact-center agent or summarization model | Spoken names, member IDs, dates of birth and callback numbers in transcripts | Safe Harbor often workable if redaction is verified |
| Regional utilization or network adequacy | Geography below the three-digit ZIP area | Expert Determination |
Where identifiers hide in health operational data
In administrative health data, identifiers sit in structured segments and free text that generic PII scanners miss. Safe Harbor names medical record numbers, health plan beneficiary numbers (member IDs) and account numbers explicitly, and its catch-all category, "any other unique identifying number, characteristic, or code," reaches claim control numbers and prior authorization numbers [1]. Identifiers inside free text, such as a portal message or a claim note, must also be removed under Safe Harbor [2]. Check each of these locations in a sample:
- X12 837 claims: subscriber and patient NM1 names, N3/N4 addresses, DMG dates of birth and REF identifiers.
- X12 835 remittances: CLP patient control numbers, payer claim control numbers and patient NM1 segments.
- HL7 v2 and FHIR exports: PID-3 identifier lists, PID-5 names, PID-7 birth dates, PID-11 addresses, and FHIR Patient.identifier values.
- Scheduling and reminder logs: appointment timestamps, callback numbers and SMS reminder bodies.
- Call and chat transcripts: callers spelling names, reading member IDs and confirming birth dates for verification; see redacting spoken PII from call recordings.
- Scanned documents and metadata: insurance card images, faxed referrals, DICOM headers and export filenames that embed medical record numbers.
Safe Harbor targets identifiers of the patient and of relatives, employers or household members [1], so provider NPIs and staff names are not automatically in scope; check whether supplier agreements or state law require removing them anyway. Free-text de-identification also faces newer attacks, covered in LLM-assisted re-identification of de-identified text.
Evidence to request before you sign
Ask for evidence of method, scope and quality assurance rather than a compliance badge. For Expert Determination, request the report itself, the expert's qualifications, the risk threshold, the assumptions about the anticipated recipient and environment, and any expiration date; OCR notes the rule does not mandate an expiration date but that experts may set one because available data changes over time [2]. A walkthrough of the report is in reviewing a HIPAA Expert Determination report, and the full document list in the de-identification evidence package checklist.
Join keys deserve their own question. Under 164.514(c), a re-identification code may stay with de-identified data only if it is not derived from information about the individual and the mechanism is not disclosed [1]. A SHA-256 hash of a medical record number or SSN is derived from the individual, so ask for random tokens instead.
Illustrative example: invented to show structure; it does not describe an available dataset.
deidentification_statement:
dataset: "Outpatient billing and scheduling extract"
method: expert_determination # or safe_harbor
citation: "45 CFR 164.514(b)(1)"
expert: { name_on_file: true, qualifications_attached: true }
anticipated_recipient: "single licensee, access-controlled training environment"
risk_threshold: "documented in report section 4"
determination_date: 2026-08-14
expiration_or_review: 2027-08-14
field_map:
patient_name: removed
date_of_birth: generalized_to_year
service_date: shifted_per_patient_offset
zip5: truncated_to_zip3_with_population_check
member_id: replaced_with_random_token
free_text_notes: ner_redaction_plus_manual_review
qa_sample: { records_reviewed: 2000, residual_identifiers_found: 3, remediated: true }
license_terms: [no_reidentification, no_linkage_without_approval, notify_on_suspected_reidentification]
Ask also for coverage facts: date range, sites, specialties, payer mix and demographic distribution. Research on de-identified clinical datasets finds they are often outdated and demographically limited [4]. A 2019 extract predates several annual ICD-10-CM code set updates, so a coding model trained on it learns retired codes.
Where the HIPAA label stops protecting you
HIPAA de-identification answers one question under one statute, and other regimes can still attach to the same records. Substance use disorder records fall under 42 CFR Part 2, whose February 2024 final rule aligned it with HIPAA, including its de-identification standard, with compliance required by 16 February 2026, so as of October 2026 those rules apply in full [6]; see 42 CFR Part 2 records in AI training data. Buyers established in the EU should note that GDPR anonymity is judged against all means reasonably likely to be used [7], a test a US expert report does not settle.
Combination is the quieter risk. An Expert Determination premised on a single recipient can fail once you join the data to another licensed source with overlapping geography and dates, as explained in linkage and mosaic risk. Put the no-linkage and no-re-identification terms in the license, and plan for state privacy rules in deidentified data under US state privacy laws.
Matching the record type to the model
Administrative records often carry less clinical detail than charts but are easier to de-identify and closer to what payer and provider operations agents need. Clinical versus administrative trade-offs for fine-tuning are covered in healthcare LLM fine-tuning data. For record-type specifics, see licensing medical coding and claims data and licensing medical records for AI training; the regulatory background is in HIPAA and AI training data.
How SourceX handles health records
SourceX sources operational datasets from US companies on request, including support histories, documents and finance workflows, and manages the licensing process; categories are not inventory, and a request does not guarantee a match. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, and personal details such as names, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Healthcare administration buyers can start at buyers in healthcare administration or describe the data on the SourceX buyer page.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Source HIPAA de-identified health data for training
Describe the health operational data you need and the de-identification method your counsel will accept, and SourceX looks for US businesses that hold it and assesses data and licensing permissions. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Start a buyer request for de-identified health data.
Sources
- Electronic Code of Federal Regulations (eCFR), Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- HealthExec, "Healthcare AI and HIPAA compliance: 5 key legal questions + answers". https://healthexec.com/topics/artificial-intelligence/healthcare-ai-and-hipaa-compliance-5-key-legal-questions-answers
- Stanford University Open Journal Systems (GRACE), "Article on de-identified clinical datasets (GRACE, article 3837)". https://ojs.stanford.edu/ojs/index.php/grace/article/download/3837/1799/11712
- Tonic.ai, "HIPAA AI compliance guide". https://www.tonic.ai/guides/hipaa-ai-compliance
- Troutman Pepper, "Final Rule Aligns 42 CFR Part 2 With HIPAA/HITECH" (2024). https://www.troutman.com/insights/final-rule-aligns-42-cfr-part-2-with-hipaahitech.html
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation), Recital 26" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.