Skip to content

Privacy, de-identification and sensitive data

HIPAA Safe Harbor vs Expert Determination for AI training data: which should a buyer require?

Quick answer

The Safe Harbor vs Expert Determination choice for AI training data turns on which details your model needs. Require Safe Harbor when year-only dates, state or three-digit ZIP geography and scrubbed notes are enough, and you want a checklist a reviewer can audit. Require Expert Determination when the model needs month- or day-level timing, finer geography, cross-source linkage or date-shifted clinical text; it keeps that detail only for the recipient, environment and controls the expert assessed. Either method takes the data outside HIPAA's protected health information rules [1].

By SourceX Editorial · Updated

Safe Harbor is a removal rule and Expert Determination is a risk judgment; both implement the same HIPAA de-identification standard in 45 CFR 164.514 [2]. That standard treats health information as not individually identifiable when it does not identify a person and there is no reasonable basis to believe it can be used to identify one [2]. OCR's guidance describes the two methods, and data de-identified by either one is no longer PHI and falls outside the Privacy Rule [1]. Short definitions sit in the glossary entries for Safe Harbor de-identification and Expert Determination.

Safe Harbor, 164.514(b)(2). The supplier removes 18 listed identifiers of the patient and of the patient's relatives, employers and household members, and must have no actual knowledge that what remains could identify the person [1][2]. Three items on the list do most of the damage to training data:

  • Geography. Every subdivision smaller than a state goes, except the first three digits of a ZIP code when the combined three-digit area holds more than 20,000 people under current Census data; smaller areas are recoded to 000 [2].
  • Dates. Every date element except the year goes for dates directly related to the patient, such as birth, admission, discharge and death, and ages over 89, along with any date element (including the year) that reveals such an age, collapse into a single 90-or-older category [2].
  • Catch-all codes. "Any other unique identifying number, characteristic, or code" goes too, except a re-identification code under 164.514(c): one not derived from or related to information about the individual, not used or disclosed for any other purpose, and whose re-identification mechanism the supplier keeps to itself [2].

Expert Determination, 164.514(b)(1). A person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods determines that the risk is very small that the information could be used, alone or with other reasonably available information, by an anticipated recipient to identify someone. The expert documents the methods and results [2]. OCR sets no numerical threshold for "very small" and requires no expiration date, though it notes that changes in technology, social conditions and available data can make re-examination appropriate [1].

Two things are not on this menu. A limited data set under 164.514(e) keeps dates and town, city, state and ZIP code, but it remains PHI and needs a data use agreement [2]; see HIPAA limited data sets and DUAs for AI development. NIST SP 800-188 is federal-agency guidance on de-identification techniques and governance that an expert may draw on, not a third HIPAA method [3].

Safe HarborExpert Determination
Legal test18 identifiers removed; no actual knowledge of identifiabilityExpert finds a "very small" risk for an anticipated recipient
Who stands behind itThe supplierA named expert, with written methods and results
DatesYear only; ages over 89 become 90+Whatever the analysis supports, such as month or shifted dates
GeographyState, or three-digit ZIP above 20,000 peopleWhatever the analysis supports, such as three-digit ZIP everywhere
ScopeNot tied to a recipientThe recipient, environment and controls analyzed
Shelf lifeHolds while the removals hold and no actual knowledge arisesNo required expiry; re-examination as conditions change
What the buyer auditsField-by-field removal evidenceThe report, its assumptions and its conditions

What each method does to claims, encounter and documentation fields

For the encounter, claims, revenue-cycle and documentation records that healthcare AI teams license, Safe Harbor's losses concentrate in dates, sub-state geography and identifiers buried in prose, while clinical and billing codes mostly survive. The table maps common fields from X12 837 claims, 835 remittances, HL7 ADT feeds and FHIR resources to each method. The right-hand column shows what an expert may accept, not what every determination allows.

Field (typical source)Under Safe HarborWhat an Expert Determination may allow
Service, admission and discharge dates (837, ADT, FHIR Encounter.period)Year only [2]Month, or per-patient shifted dates that keep intervals
Remittance and adjudication dates (835), appeal datesYear only [2]Same options; days-to-payment can survive
Birth date and age (FHIR Patient.birthDate)Age in years; over 89 becomes 90+ [2]Age bands or birth year, depending on the population
Patient ZIP and countyCounty removed; ZIP cut to the first three digits if the area exceeds 20,000 people, else 000 [2]Three-digit ZIP everywhere, or county with small-cell suppression
MRN, member ID, claim control and account numbersRemoved; a random patient key survives only if counsel confirms it meets 164.514(c) [2]Consistent tokens, including tokens for approved cross-source linkage
ICD-10-CM, CPT/HCPCS, CARC/RARC denial codes, chargesKept; not on the listKept, with review of rare combinations
Treating provider names and NPIsNot on the list; OCR's guidance limits name removal to patients and their relatives, employers and household members [1]May be generalized if a rare specialty in a small area points to patients
Note text, appeal letters, CDI queriesEvery listed identifier removed wherever it appears in the proseSurrogates and generalizations the expert accepts in context

Codes, age, sex and three-digit ZIP are quasi-identifiers: harmless alone, identifying in combination.

Rocher and colleagues estimated, with a generative model, that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes; that is a model estimate, not a count of people actually re-identified [4]. Sweeney's k-anonymity model formalizes one defense, each record indistinguishable from at least k-1 others, and also shows attacks that succeed on k-anonymous releases unless accompanying policies hold [5]. Safe Harbor runs neither test; an expert's analysis is where that measurement belongs. For the field-level cost of year-only dates and truncated ZIPs, see what temporal and geographic models lose under Safe Harbor.

Matching the method to what the model must learn

Write down the features the model cannot do without, then require the least detailed method that delivers them. Safe Harbor is preferable whenever it is sufficient, because it attaches no recipient conditions to the data.

Model needExample taskSafe Harbor resultRequire
Intervals between eventsDays from submission to denial, length of stay, prior-auth turnaroundDates are year-only; ask counsel whether a derived interval field is acceptableExpert Determination with date shifting or interval fields
SeasonalityMonth-end billing cycles, respiratory-season volumeLostExpert Determination with month-level dates
Sub-state geographyPayer mix by region, site-of-care modelsThree-digit ZIP where the area exceeds 20,000 peopleExpert Determination with generalized geography
Patient history within one datasetDiagnosis history, encounters per patient per yearDepends on a random patient key that counsel confirms meets 164.514(c)Safe Harbor if the key holds up; otherwise Expert Determination
Joining two suppliers' recordsClaims linked to EHR notesA hash of an MRN or SSN is derived from patient information, so it does not meet 164.514(c) [2]Expert Determination that covers the linkage
Coding and denial classificationICD-10 or CPT assignment, CARC predictionCodes and amounts surviveSafe Harbor is often enough
Documentation and summarizationCDI query drafting, note summarizationFeasible once notes are scrubbed; dates in text become yearsEither; Expert Determination if date-shifted notes matter
Evaluation set with rare conditionsHeld-out cases for rare diagnosesRare cases raise the actual-knowledge questionExpert Determination with small-cell rules

Vendor guides describe the same trade-off. One says Safe Harbor may not suit use cases that need high fidelity, such as data used to train ML models [9]. Another says Expert Determination can keep month-level dates and regional geography but demands more effort, expertise and ongoing governance [10]. Treat both as evidence of market practice, not authority. If the choice is still between clinical and administrative records, see healthcare LLM fine-tuning data.

Why a Safe Harbor label on clinical text needs evidence

On structured claims, Safe Harbor is close to mechanical; on notes, letters and scanned attachments it is only as good as the detection pipeline, so the label is really a claim about recall. NISTIR 8053 separates structured data from the harder problem of unstructured data such as medical text [6]. One vendor guide says finding all 18 categories in clinical text takes a tuned NLP pipeline, not find-and-replace [7], and the open-source Presidio project cautions that its ML-based detection cannot promise to find all sensitive information [8].

Identifiers hide in predictable places: fax numbers on referral cover sheets, "seen 3/14" in progress notes, a daughter's name in a discharge plan, an employer in an occupational history, implant serial numbers, and portal URLs pasted into messages. Each maps to a Safe Harbor category, because the list covers fax numbers, dates, names of relatives, employers, device identifiers and URLs [2]. Recorded speech adds a harder question, since the list includes biometric identifiers such as voice prints [2]; see doctor-patient conversation audio for ambient documentation. Ask for per-category recall on a hand-annotated sample, as covered in de-identifying clinical free text for LLM training.

Training does not change the data's HIPAA status, but it changes the exposure. Carlini and colleagues extracted verbatim training sequences, including contact details, from GPT-2 by querying it [11], and later work recovered thousands of training examples from aligned production models [12]. A residual identifier in a note can become a model output; see training-data extraction and memorization risk.

What an Expert Determination buys back, and what it binds you to

Expert Determination buys back detail by tying the risk judgment to an anticipated recipient and its controls, so it helps only if its assumptions match how your team will store, combine, train on and release the data. Read the scope before the conclusion:

  1. Recipient and environment. The regulation measures risk against an anticipated recipient [2]. Check that the report covers your organization, cloud account or enclave, and any annotation vendors who will touch the records.
  2. Uses, including model release. Check whether the expert considered model training and third-party access to the trained model or its outputs. Given the extraction results above [11][12], a determination that is silent on model release leaves a gap.
  3. Linkage. The regulation weighs other reasonably available information [2], so a determination rests on assumptions about what else the recipient holds. Licensing a second dataset into the same environment can break those assumptions; see linkage and mosaic risk when combining datasets.
  4. Conditions. List the controls the conclusion depends on, such as access limits, no re-identification, no onward disclosure and retention, because they become license obligations. NIST SP 800-188 describes enclave sharing and governance measures such as a Disclosure Review Board and re-identification studies [3].
  5. Time and refresh. OCR requires no expiry but anticipates re-examination [1]. Ask whether monthly refreshes are covered cumulatively and what triggers re-certification.

OCR sets no number for "very small" [1], so ask which risk measure and threshold the expert used and what outside data the attack model assumed. OCR's guidance also does not tie expert status to a single degree or certification, so ask for the expert's relevant experience [1]. The full review process is in how to review a HIPAA Expert Determination report.

Obligations that survive de-identification

Meeting 164.514(b) ends the Privacy Rule's hold on that data, but California law, non-US privacy law, the rules for substance use disorder records and the license itself can still matter to a buyer.

  • California. As of October 2026, the CCPA has separate provisions for deidentified patient information in Civil Code sections 1798.146 to 1798.148. They prohibit re-identification except for listed purposes and require contracts for its sale or license to state that it includes deidentified patient information, prohibit re-identification by the buyer and bar onward disclosure to third parties not bound by the same or stricter terms [13]. The CCPA also requires a business that sells or discloses such data to say in its privacy policy which HIPAA method it used [13]. Confirm the current text before relying on any clause.
  • Substance use disorder records. Records from federally assisted programs fall under 42 CFR Part 2 [14], so ask how the supplier handled them before de-identification. The 2024 final rule aligned parts of Part 2 with HIPAA, and its compliance date was 16 February 2026 [14]. See 42 CFR Part 2 records in AI training data.
  • EU and UK data. HIPAA de-identification is not GDPR anonymization. Rocher and colleagues concluded that heavily sampled anonymized datasets are unlikely to meet GDPR's standard [4]; see de-identified, anonymized and pseudonymized compared.
  • Identifiable intake. If your team would receive PHI and de-identify it, the deal is a different structure; see business associate agreement or data license and SourceX's note on what a HIPAA-covered supplier may license.

Writing the method into a health-data request

State the method, the retained detail and the downstream uses up front, so the supplier designs de-identification around the model instead of retrofitting it after a sample fails.

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset: professional claims with remittances (837P + 835) and linked denial appeal letters
hipaa_method_required: expert_determination   # Safe Harbor cannot support the date shift below
retained_detail:
  service_and_remit_dates: per-patient shift, intervals preserved
  patient_age: years, top-coded at 90+
  geography: three-digit ZIP, small areas suppressed
  patient_key: supplier-assigned random key, not derived from identifiers
  free_text: names, contact details and ID numbers replaced with typed surrogates
intended_uses:
  - supervised fine-tuning of a denial-appeal drafting model
  - held-out evaluation set
  - trained model offered to customers; outputs visible to third parties
recipient_environment: buyer cloud account, named-user access, no other health data joined
refresh: monthly; determination must cover cumulative releases
evidence_requested:
  - signed determination with methods, risk measure, threshold and conditions
  - per-category recall on a hand-annotated free-text sample
  - re-certification triggers

Evidence to request under each method:

  • Safe Harbor: a map of all 18 categories to the fields and text locations where each was removed; the three-digit ZIP population table and the Census data used; handling of ages over 89, including ages implied by dates in notes; free-text recall by category; how any record key was generated; and how rare cases were reviewed for actual knowledge.
  • Expert Determination: the signed report, plus written answers on the five scope points and the risk metric covered above.

The full document list is in the de-identification evidence package checklist, and the wider cluster is mapped in de-identified data for AI training. When SourceX sources health records, it requires HIPAA de-identification by Safe Harbor or Expert Determination before anything is considered for a license, records the method used for each dataset and prepares diligence materials for the buyer's review, under its legal framework. No de-identification method is perfect, so name the method and retained fields you need when you describe your health-data requirements to SourceX.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Sourcing de-identified health records for model training?

Describe the records you need, the HIPAA method you require, and the dates, geography and text detail your model depends on. SourceX looks for US companies that hold such records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Share your data and privacy requirements.

Sources

  1. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  2. Electronic Code of Federal Regulations, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (current text). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  3. NIST and U.S. Census Bureau, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  4. Rocher, Hendrickx and de Montjoye, "Estimating the success of re-identifications in incomplete datasets using generative models," Nature Communications (2019). https://www.nature.com/articles/s41467-019-10933-3
  5. Latanya Sweeney, "k-Anonymity: A Model for Protecting Privacy" (2002). https://dataprivacylab.org/people/sweeney/kanonymity.html
  6. NIST (Simson L. Garfinkel), "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
  7. Ertas (vendor blog), "HIPAA-Compliant AI Training Data: A Practical Guide for Healthcare Organizations." https://www.ertas.ai/blog/hipaa-compliant-ai-training-data-guide
  8. Microsoft presidio project, "Presidio - Data Protection API." https://pkg.go.dev/github.com/microsoft/presidio
  9. Tonic.ai (vendor guide), "HIPAA and AI compliance guide." https://www.tonic.ai/guides/hipaa-ai-compliance
  10. Nirmitee (vendor blog), "Healthcare data de-identification for AI: Safe Harbor vs Expert Determination." https://nirmitee.io/blog/healthcare-data-deidentification-ai-safe-harbor-expert-determination
  11. Carlini et al., "Extracting Training Data from Large Language Models," USENIX Security (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  12. Nasr et al., "Scalable Extraction of Training Data from Aligned, Production Language Models," ICLR (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
  13. California Legislature, "California Consumer Privacy Act of 2018 (Civil Code Title 1.81.5)" (codified text). https://leginfo.legislature.ca.gov/faces/codes_displayText.xhtml?division=3.&lawCode=CIV&part=4.&title=1.81.5
  14. U.S. Department of Health and Human Services, "Confidentiality of Substance Use Disorder (SUD) Patient Records (Final Rule)," Federal Register 89, No. 33 (2024). https://www.govinfo.gov/content/pkg/FR-2024-02-16/html/2024-02544.htm

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data