Privacy, de-identification and sensitive data
De-identifying clinical free text for LLM training: PHI detection, surrogates and residual risk
Quick answer
Clinical notes can be licensed as HIPAA de-identified training data only after a measured PHI-detection pipeline, not find-and-replace, has removed or replaced identifiers across HIPAA's 18 Safe Harbor categories or passed an Expert Determination [1][2]. Buyers should require per-entity recall on a held-out labeled sample, a stated surrogate strategy, a residual-PHI rate from blind review and a check for contextual identifiers that NER misses. Leftover PHI does not stay in the file: a fine-tuned model can repeat it [4].
By SourceX Editorial · Updated
This page covers what a buyer should accept from a supplier's pipeline. For the preparation side, see the guide to de-identifying medical records; for the broader framework, start at the de-identified data for AI training hub.
Why free-text PHI needs an NLP pipeline, not regex
Free text defeats pattern matching because identifiers appear in irregular forms, inside sentences, with no field boundaries. A structured EHR export has a patient_name column you can drop; an encounter note says "pt's daughter Maria drove her in from Lodi" and hides a relative's name and a small place in prose. HIPAA Safe Harbor covers identifiers of the individual and of relatives, employers and household members, so a pipeline that only targets the patient's own name fails the standard [1].
Mature pipelines are hybrids: pattern matching for well-formed items such as phone numbers and MRNs, machine-learning named-entity recognition for names and places, and a replacement step. Open toolkits such as Microsoft Presidio follow the same recognizer-plus-anonymizer design, and the project itself warns that its trained models cannot guarantee finding all sensitive information [5].
LLM-based redaction is now a serious option. Research on LLM redaction reports that models can catch PII types that evade regex and traditional NER, while also introducing new failure modes that need their own evaluation [3]. For a comparison of approaches, see LLM-based PII redaction vs NER and regex.
What the PHI entity inventory must cover in clinical notes
The detection schema should be broader than the 18 Safe Harbor identifiers, because notes carry identifying detail that the list does not name. Shared-task corpora used to benchmark clinical de-identification have annotated a wider set than HIPAA lists, including professions, full dates, and the names of clinicians and facilities, because combinations of individually harmless details across a longitudinal record can still single out a patient. HHS guidance likewise notes that Safe Harbor fails if the covered entity has actual knowledge that remaining information could identify someone [1].
For a buyer, the practical schema has these layers:
- Direct identifiers: patient, relative and clinician names; MRN, account, health plan and device numbers; phone, fax, email, URL and IP addresses; SSN.
- Dates and ages: all date elements more specific than year, and ages over 89 under Safe Harbor [1].
- Locations: street address, city, ZIP code, facility names, unit names ("4 West"), and named employers.
- Contextual identifiers: rare diagnoses, unusual occupations ("retired state senator"), named events, distinctive family structures, and free-text references to media coverage.
Contextual identifiers are where most residual risk lives. NER models are trained on span-level labels for names and numbers; they rarely flag "the only pediatric lung transplant in the county this year." See indirect identifiers in business text for a fuller taxonomy, and the PHI glossary entry for definitions.
Surrogates vs placeholder tags: how the choice changes your fine-tuned model
Realistic surrogates hide missed PHI among fake PHI, while placeholder tags expose every miss and teach the model a tag vocabulary. The clinical NLP literature calls the surrogate approach "hiding in plain sight": because both manual and automated removal leave residual identifiers, replacing every detected identifier with a realistic fictitious one makes the leftovers hard to tell apart from the surrogates. With tags, a missed "Mrs. Alvarez" stands out in a note where every other name reads [NAME].
The trade-off matters for SFT. Placeholder tags such as [NAME] or <DATE> produce outputs full of tags, which is fine for coding assistance but awkward for summarization or documentation agents that must write natural notes. Surrogates keep fluency, but naive generation can introduce distribution artifacts: gender-mismatched names, impossible date intervals, or one surrogate pool reused so often the model memorizes it.
Consistency is a design decision, not a default. A consistent mapping (same real name always becomes the same surrogate within a patient) preserves coreference across a longitudinal record, but it also means one leaked real name sits next to a stable, repeated pattern that an adversary can exploit; random per-mention substitution breaks coreference instead. Ask the supplier which strategy they used, at what scope (note, patient or corpus), and how date shifting preserves intervals. Keyed mapping options are covered in pseudonymisation techniques for training data.
Acceptance evidence to request from a supplier's pipeline
The evidence that matters is measured recall per entity type on notes like yours, not a vendor's headline F1 on a public benchmark. Public shared-task corpora are useful for system comparison, but notes from a different specialty, EHR template or dictation style can shift performance sharply. Ask for metrics computed on a labeled sample drawn from the delivered corpus itself.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Evidence item | What good looks like | Red flag |
|---|---|---|
| Labeled evaluation sample | Random, stratified by note type (progress, discharge, call note, appeal letter); double-annotated with adjudication | Sample chosen by the pipeline vendor; single annotator |
| Per-entity recall | Reported separately for NAME, DATE, LOCATION, ID, CONTACT, AGE>89, PROFESSION | One blended F1 number |
| Precision | Reported, with examples of over-redaction of drug names or eponyms ("Hodgkin," "Parkinson") | Not reported, or recall reported alone |
| Residual-PHI audit | Blind human review of post-pipeline text; residual rate per 1,000 notes with confidence interval | "Spot-checked" with no count |
| Surrogate method | Named strategy, scope and date-shift rule | "Masked" with no detail |
| Contextual review | Documented pass for rare conditions, named events, small facilities | None |
| HIPAA basis | Safe Harbor checklist or Expert Determination report with assumptions [2] | "HIPAA compliant" with no method |
For sample sizes and thresholds, use auditing residual PII in a delivered dataset. For the complete document set, use the de-identification evidence package checklist.
Illustrative example: invented to show structure; it does not describe an available dataset.
A hypothetical acceptance note (invented numbers) might read: "On 500 randomly sampled delivered notes, two reviewers found 3 residual direct identifiers (2 relative first names, 1 facility unit phone extension); per-entity recall on the 200-note labeled set was lowest for PROFESSION and LOCATION." That level of specificity lets you decide whether to accept, request a rerun, or add your own scrubbing pass before training.
How residual PHI becomes model risk after fine-tuning
Residual identifiers that survive de-identification can be reproduced by the model you train, which turns a data-handling issue into a product issue. A 2025 study reports a small open model emitting email addresses, phone numbers and account details when prompted, illustrating how personal data in training text resurfaces at inference [4]. Clinical SFT sets are small and repetitive compared with pretraining corpora, which raises the chance that a rare note is memorized.
Three controls reduce exposure. Deduplicate notes and copy-forward templates before training, since repeated text is more likely to be memorized (near-duplicate detection with MinHash and LSH). Run canary or extraction probes against the fine-tuned checkpoint (training-data extraction and memorization risk). And retest the de-identified text with an LLM adversary, since de-identification that resists a human reader may not resist a model (LLM-assisted re-identification).
Safe Harbor, Expert Determination and the legal floor
HIPAA treats health information as de-identified when there is no reasonable basis to believe it can identify an individual, met either by Safe Harbor or by Expert Determination [2]. Safe Harbor also requires that the covered entity have no actual knowledge that remaining information could identify the person, which is hard to certify for narrative text full of context [1]. Expert Determination lets a qualified expert assess risk given the recipient and controls, which often fits training data that needs to keep clinical detail; see Safe Harbor vs Expert Determination for AI training data.
Neither method makes risk zero. NIST SP 800-188 cautions about the inherent limitations of traditional de-identification and stresses governance around releases [6]. State law can add obligations: California Civil Code 1798.148 restricts re-identifying deidentified patient information, so expect license terms that prohibit it (re-identification prohibition clauses) [7].
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Note types and the failure modes to test for each
Each clinical text type fails differently, so the evaluation sample should be stratified by source. The table below lists the patterns to probe.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Note type | Typical leak | Test to run |
|---|---|---|
| Encounter and progress notes | Relative names, social history ("works at the mill on Route 9") | Recall on NAME (non-patient) and LOCATION |
| Discharge summaries | Copy-forward headers with MRN, attending, unit | Header regex audit plus dedup |
| Payer call notes | Callback numbers, reference numbers, rep names | CONTACT and ID recall; see provider-to-payer calls |
| Appeal letters | Letterhead, signatures, claim numbers, dates of service | Layout-aware extraction before NER; see claim denial appeal letters |
| Dictated or ambient-scribe transcripts | Spelled-out names, numbers spoken as words ("five five five...") | Normalization before detection; see doctor-patient conversation audio |
How SourceX handles clinical text requests
SourceX sources operational datasets from US companies on request and manages licensing; it does not hold clinical notes in stock, and a request does not guarantee a match. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, every dataset is rights-reviewed for ownership and consents, and personal details are removed or replaced before delivery with the method recorded and a sample checked. No de-identification method is perfect, which is why the acceptance evidence above still belongs in your review. You can describe the notes you need on the buyer request page or review the medical records licensing overview.
Request de-identified clinical notes for LLM training
SourceX looks for US businesses that hold the clinical or administrative notes you describe, assesses the data and licensing permissions, and agrees allowed uses in a license before anything is delivered through a private, access-controlled workflow. Every release is approved by the supplying company. Describe your note types, volume and de-identification requirements at sourcex.si/buyers.
Frequently asked questions
Is Presidio or a similar open-source tool enough for clinical notes?
It is a starting framework, not an acceptance criterion. Presidio's own documentation says it cannot guarantee finding all sensitive information, [5]. Its general-purpose recognizers are not built around clinical identifiers such as MRNs or hospital unit names, so measure recall on your own labeled notes.
Should date shifting be per patient or per corpus?
Per patient, with a consistent offset, preserves intervals such as length of stay and days to readmission, which summarization and coding models often need. A corpus-wide offset is easier to reverse if one true date leaks.
Does de-identified text still need a data use agreement?
Data that meets the HIPAA standard is no longer PHI [2], but licenses for clinical text commonly still restrict re-identification, linkage and onward transfer. See combining de-identified datasets for linkage risk.
Sources
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
- arXiv, "Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models" (2025). https://arxiv.org/pdf/2507.04478
- Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- FindLaw (California Civil Code), "California Civil Code 1798.148". https://codes.findlaw.com/ca/civil-code/civ-sect-1798-148/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.