Privacy, de-identification and sensitive data
How to review a HIPAA Expert Determination report before licensing health data
Quick answer
Review an Expert Determination report by testing whether it covers your deal, not whether it exists. Under 45 CFR 164.514(b)(1), a qualified expert must find that re-identification risk is very small for an anticipated recipient and must document the methods and results [2]. So check five things: the named recipient and environment, the exact field list, the analysis behind the conclusion, the expiry and re-determination triggers, and whether model training was in scope.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
It assumes the supplier has already chosen Expert Determination. If you are still deciding which method to require, start with Safe Harbor vs Expert Determination for AI training, and for the term itself see the Expert Determination glossary entry.
What the regulation requires the report to show
The rule requires two things from the expert: a determination of very small risk and documentation of the methods and results that justify it [2]. The standard in 164.514(a) is that there is no reasonable basis to believe the information can identify an individual [2]. The expert must be a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable [2].
HHS OCR's 2012 guidance adds the working detail reviewers rely on. It says no particular degree or certification is required, describes risk principles such as replicability, data source availability and distinguishability, and lists techniques such as generalization, suppression and perturbation [1]. A report that names none of these principles and none of these techniques is a conclusion letter, not a determination.
Vendor guidance on AI use makes the same point in plainer terms: the expert has to analyze the dataset, the context it will be used in and who receives it, and record the methods and results [3]. An internal compliance sign-off does not substitute for an expert's determination [4].
Recipient and environment: the clause AI buyers should test first
A determination is only as good as its description of the anticipated recipient, because the rule measures risk "by an anticipated recipient" [2]. OCR's guidance ties the expert's analysis to the recipient's environment and the other data that recipient can reasonably reach [1]. If the report was written for the supplier's analytics vendor, a hospital research partner or "licensees" in general, it may not cover an AI lab with large public-web corpora and in-house linkage capability.
Read the recipient section for these specifics:
- Named recipient or recipient class. Your legal entity, or a class that clearly includes it. Affiliates, contractors, annotation vendors and evaluation partners should be listed if they will touch the data.
- Controls the expert relied on. Access restrictions, no-linkage and no-re-identification clauses, audit rights, encryption, retention limits. Each relied-on control is one you must actually implement and, usually, contractually promise.
- Hosting environment. Cloud provider, region and tenancy model. A move from one region or provider to another can change the risk assumptions if the report relied on a specific enclave.
- Reasonably available external data. What the expert assumed you hold. A lab training on broad web text, voter files or social media scrapes differs from a hospital analytics team.
If any relied-on control conflicts with how your training pipeline works, flag it before signing. The de-identification evidence package checklist lists the companion documents to request alongside the report.
Field list and distributions: does the report describe what will be delivered
The determination covers the dataset the expert analyzed, so the report's data dictionary must match your delivery manifest field by field [1]. Fields added after the analysis, new free-text columns, extra date precision or a new source system are outside the determination until the expert re-analyzes them. This gap is easy to miss in multi-drop health data deals.
Compare the report against a sample extract on these points:
- Quasi-identifiers and their transforms. Age bands, three-digit ZIP or larger geography, date shifting or year-only dates, rare diagnosis or procedure codes (ICD-10-CM, CPT, HCPCS), provider NPI handling and facility identifiers.
- Free text. Clinical notes, call transcripts and claim narratives need their own analysis: which PHI detector, recall measured on what annotated sample, and how surrogates were generated. See de-identifying clinical free text for LLM training.
- Record linkage keys. Tokenized patient IDs or hashes that let you join across files or deliveries. Under 164.514(c), a re-identification code must not be derived from information about the individual and must not be disclosed for other purposes [2]; a plain hash of a medical record number generally fails that test.
- Small cells. Minimum equivalence class size or k-threshold, and what happens to records that fall below it.
Method evidence: separating analysis from assertion
A credible report shows its work: the threat model, the risk metric, the measured result and the threshold it was compared with [1][3]. Look for a stated metric (for example, maximum or average re-identification probability across equivalence classes, or a journalist, prosecutor or marketer risk model), the threshold the expert treated as very small, and the measured value on the actual data.
Weak reports share recognizable failure modes. They state a threshold but no measured value, analyze a sample drawn from a different month than the delivery, ignore longitudinal linkage across encounters, or treat free text as covered because structured fields passed. Ask for the expert's credentials and a statement of independence from the supplier's commercial team, since the rule depends on the expert's qualification [2][4].
Expiry and re-determination triggers
The Privacy Rule does not itself set an expiration date, but OCR's guidance notes experts may limit a determination in time because data availability and re-identification techniques change [1]. Treat an undated or open-ended report as a question to raise, not a comfort. For recurring purchases, each new extract should either fall inside the original analysis or carry an updated determination.
Common re-determination triggers to write into your diligence and contract:
- A refresh that adds fields, extends the date range or adds a new source system.
- A new recipient: a contractor, a new affiliate, or a vendor running evaluation.
- A change in hosting provider, region or tenancy.
- Linkage of this dataset with another licensed dataset; see combining de-identified datasets and linkage risk.
- Evidence of new re-identification methods against similar data, including LLM-assisted re-identification of de-identified text.
Model training, weights and outputs: was AI use in scope
Properly de-identified data is not PHI, so HIPAA does not restrict using it for AI development [5]. The open question is whether the expert's risk analysis assumed your use. A determination written for aggregate analytics may not have considered that a fine-tuned model can memorize and regurgitate rare strings, or that released weights put that risk in third-party hands.
Check whether the report addresses training, memorization and extraction risk, whether weights stay internal or may be released, whether model outputs reach end users, and whether retrieval-augmented generation (RAG) will surface records verbatim. If none of this is in scope, ask the expert for an addendum that names the training use, or for the controls the expert would require (for example, internal-only weights, output filtering, deduplication of rare records). Then match those controls to your own PII scan before fine-tuning.
Reviewer checklist for an Expert Determination report
Use this table during the review call with the supplier and, where possible, the expert.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Area | Question to answer | Evidence to request | Red flag |
|---|---|---|---|
| Expert | Who made the determination and how are they qualified? | CV, independence statement, engagement letter scope | Signed only by supplier's privacy office [4] |
| Recipient | Is your entity or class named? Are contractors covered? | Recipient section, list of relied-on controls | "Licensees" with no environment description [2] |
| Environment | Which hosting, access and linkage assumptions apply? | Control schedule, data use terms | Assumes enclave access you will not use |
| Fields | Does the data dictionary equal the delivery manifest? | Dictionary, sample extract, row counts by file | Columns in sample not in report |
| Free text | How were notes or transcripts handled and tested? | Detector, recall on annotated sample, surrogate method | "Free text redacted" with no measurement |
| Method | Which metric, threshold and measured result? | Risk tables by equivalence class | Threshold stated, measured value absent [1] |
| Linkage keys | How are tokens generated and who holds the key? | Tokenization spec, key custody | Plain hash of MRN or SSN [2] |
| Time | Is there a date, expiry and refresh rule? | Effective date, trigger list | Undated, covers "future extracts" [1] |
| AI use | Was training, weight release and output exposure analyzed? | Use-scope section or addendum | Analytics-only use context |
A worked excerpt of what a usable recipient clause looks like:
Illustrative example: invented to show structure; it does not describe an available dataset.
Anticipated recipient: Licensee legal entity and named subprocessors
(annotation vendor; cloud host, single US region, dedicated tenancy).
Intended use: supervised fine-tuning and internal evaluation of language
models; model weights not distributed outside Licensee.
Relied-on controls: no linkage to external identified data; no
re-identification attempts; role-based access; retention per license.
Analysis covers extract dated 2026-08 (fields per Appendix A only).
Determination expires 24 months from signature or on any trigger in §6.
How this fits a licensing process with SourceX
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases. Every dataset is rights-reviewed for ownership and consents, delivered under a license that defines records, uses, term and delivery, and health records require HIPAA de-identification by Safe Harbor or Expert Determination. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and a report like this one belongs in that package for your own review.
Personal details are removed or replaced before delivery and the method is recorded and sample-checked, but no method is perfect, which is why the checks above still matter. If you are scoping healthcare data, the healthcare LLM fine-tuning data guide and the cost of de-identification help frame the request, and you can describe the dataset you need to SourceX. More context sits in the privacy and de-identification hub and the AI data hub.
Request de-identified health data with an Expert Determination you can review
Describe the health records, fields and training use you need; SourceX looks for US businesses that hold that data, and nothing is contracted until a supplier agrees. Every release is approved by the supplying company and delivered only through private, access-controlled workflows after an executed agreement. Start a buyer request at sourcex.si/buyers.
Sources
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations (eCFR), Office of the Federal Register, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- Nirmitee, "Healthcare data de-identification for AI: Safe Harbor vs Expert Determination". https://nirmitee.io/blog/healthcare-data-deidentification-ai-safe-harbor-expert-determination
- Tonic.ai, "HIPAA compliance for AI". https://www.tonic.ai/guides/hipaa-ai-compliance
- HealthExec, "Healthcare AI and HIPAA compliance: 5 key legal questions answered". https://healthexec.com/topics/artificial-intelligence/healthcare-ai-and-hipaa-compliance-5-key-legal-questions-answers
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.