Skip to content

Privacy, de-identification and sensitive data

LLM-assisted re-identification: why de-identified text needs a stronger test now

Quick answer

Yes, capable LLMs can re-identify or profile people in de-identified text, usually not by recovering a redacted name but by inferring location, employer, role, age band or health status from residual context, then narrowing the candidate set. Recent work shows LLMs catching PII that pattern matching and NER miss [1], and the same reading ability lets a model exploit whatever residual cues remain. Before licensing or releasing a text corpus, add an LLM-assisted motivated intruder test to acceptance, alongside detector recall checks.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What changed: the intruder now reads like a domain expert

The capability shift is that inference, not memorization, has become cheap. Inferring who wrote a ticket, or where they work, from scattered context used to need a patient human analyst with domain knowledge; a general-purpose model can now process large volumes of records and reason over job titles, local references and timelines the way an investigator would. Buyers should treat this as the working assumption as of October 2026.

The same capability shows up on the defender side. A 2025 preprint reports that LLMs detect PII types that pattern matching and named-entity recognition miss [1], which suggests that residual cues your regex and NER stack left behind are likely legible to a model. Tools such as Microsoft Presidio state plainly that their ML-based detectors cannot guarantee finding all sensitive information [9]. Assume the attacker uses at least as strong a model as your best detector.

This is distinct from memorization and extraction risk, where a trained model regurgitates its training text. That threat is covered in training-data extraction and memorization risk. Here the concern is the corpus itself: a licensee, insider or downstream user pointing a model at released records.

Why de-identified text leaks: residual context, not residual names

De-identified free text leaks through quasi-identifiers that no identifier list enumerates. Support tickets, call transcripts, email threads and operational notes are full of residual identifiers. Carrell et al. found that both manual and automated de-identification leave residual PHI, which is why they proposed realistic surrogates to conceal what gets missed [8]. In business text the residue looks like this:

  • Rare events and dates: "the flood at our second warehouse last March", an outage ticket that references a specific incident number.
  • Role plus organization size: "I'm the only CFO-level person in a 12-person firm", "our sole Spanish-speaking agent".
  • Small places: a branch in a town with one employer, a ZIP3 that maps to a sparsely populated area.
  • Stylometry: signature phrasing, recurring typos and sign-offs that support author identification across documents even after names are surrogated.
  • Cross-record linkage: the same account surfacing in 40 tickets, each harmless alone, jointly unique.

Uniqueness compounds quickly. Rocher et al. showed that a generative model can predict whether a person is unique from a handful of attributes even in heavily sampled data, with AUC of 0.84 to 0.97 across 210 populations; the frequently quoted 99.98% figure is a model estimate, not a count of actual re-identifications [7]. NIST's survey of the field documents multiple cases where released de-identified data was re-identified [6], and recent legal scholarship notes repeated re-identification of de-identified educational datasets using auxiliary information [10]. For the cluster overview see the de-identified data for AI training buyer's guide; for a deeper catalog of these cues, see indirect identifiers in business text.

Attribute inference matters even when identity does not resolve

Attribute inference is a privacy harm in its own right, even if no name is recovered. A model that concludes "this caller is likely pregnant, in Ohio, employed by a regional hospital" has derived sensitive data about a real person, and that output can be joined to other sources later. For corpora destined for pre-training, SFT or evaluation, treat inferable health, financial, union, religious and location attributes as findings, not as acceptable residue.

The legal tests point the same way. GDPR Recital 26 judges anonymity by all the means reasonably likely to be used to identify someone [3], and the UK ICO frames identifiability as a risk assessment anchored by the motivated intruder test [2]. HIPAA's standard in 45 CFR 164.514(a) asks whether there is a reasonable basis to believe information can identify an individual [4]. When frontier models are a commodity, "reasonably likely means" plausibly includes them, so a test that ignores LLMs tests yesterday's intruder.

Detector recall is necessary but not sufficient

Measuring what your redaction pipeline misses answers a different question from whether the output is identifiable. A detector recall test asks: of the names, emails, phones and account numbers present, how many were caught? An intruder test asks: given everything left, can someone work out who this is or what is true about them? Run both. Recall testing methods are covered in PII redaction for LLM training data.

HIPAA Safe Harbor illustrates the gap. It requires removing 18 listed identifiers [5], yet a clean Safe Harbor pass on free text can still leave a narrative that a model resolves to one patient. HHS guidance itself notes that neither Safe Harbor nor Expert Determination drives risk to zero [5], and an Expert Determination that predates LLM-capable intruders deserves a fresh look. See Safe Harbor vs Expert Determination for AI training data and de-identifying clinical free text.

How to run an LLM-assisted motivated intruder test

An LLM-assisted intruder test samples records, prompts strong models to infer identity and attributes, and scores hits against ground truth the data holder keeps. It is a direct extension of the ICO's motivated intruder test [2], with a model standing in for the intruder: score anonymization by what an LLM adversary can still infer, then use those inferences to guide further edits. The design choices below make it repeatable enough to put in an acceptance gate.

Illustrative example: invented to show structure; it does not describe an available dataset.

StepWhat to doOutput
1. ScopeList target attributes: person identity, employer, city or region, age band, health condition, financial status, author identity across recordsAttribute list with sensitivity tier
2. SampleStratified sample of 300 to 1,000 records across source system, record type and length; oversample long free-text fieldsSample manifest with record IDs
3. Ground truthData holder labels true attribute values from the pre-redaction source, inside their environmentSealed answer key, never shipped
4. Attack promptsAt least two strong models, zero-shot and with reasoning; prompt per attribute ("Infer the city, employer and job role of the writer; give top-3 guesses and confidence")Raw model guesses per record
5. Linkage passGroup records by surrogate account or thread ID and re-run the attack on concatenated contextJoint-context hit rates
6. Authorship passAsk the model to cluster records by likely author and compare with true authorsAuthor-linkage precision
7. ScoreTop-1 and top-3 hit rates per attribute against a base-rate guess (e.g., most common city)Lift over baseline, per attribute
8. TriagePull every high-confidence correct hit; tag the cue that enabled it (rare event, role, place, style)Cue taxonomy feeding remediation

Three design rules keep the result honest. First, run the attack where the answer key lives, on the data holder's side, so testing never creates a new disclosure. Second, report lift over a base-rate guess rather than raw accuracy, since guessing the largest metro area is often right by default. Third, re-test after remediation with fresh prompts, because tuning prompts against a fixed sample overfits the test.

An illustrative acceptance record for one attribute might read: attribute=employer; model=A; n=500; top1_hit=0.04; baseline=0.03; lift=1.3; high_conf_correct=2; cues=[rare_event, role_title]; decision=pass_with_fixes. The format matters less than the habit of recording model, prompt version, sample and decision for every run.

Remediation that reduces inference without destroying utility

The most effective mitigations remove the cues the intruder test surfaced rather than adding more detectors. In rough order of impact:

  • Drop fields you do not need. If the use case is ticket routing, the free-text "customer background" field may be pure risk.
  • Generalize rare details. Replace specific incident dates with month or quarter, branch towns with region, exact headcounts with bands; techniques are covered in de-identifying tabular and transactional data.
  • Use realistic surrogates consistently. Surrogate names, companies and places so residual misses blend in [8], keeping mappings stable within a thread but not across the corpus.
  • Break cross-record linkage. Re-key account and thread IDs per release, and cap records per entity.
  • LLM-guided rewriting for high-risk records. Feed the intruder's correct inferences back as edit targets, then re-test; review outputs for meaning drift.
  • Consider synthetic text for the riskiest slices. See differentially private synthetic text.

Every generalization costs signal, so measure it. Track task metrics before and after remediation and check over-redaction, which can quietly ruin a fine-tuning or eval set. If you plan to join this corpus with other licensed data, re-run the test on the joined view; linkage and mosaic risk explains why.

What buyers should ask for before licensing de-identified text

Buyers should ask the data holder for evidence that the corpus was tested against an LLM-capable intruder, not only a redaction log. A practical request list:

  • De-identification method, tool versions and configuration for each field, with known detector gaps.
  • Recall results on a labeled sample, by identifier type.
  • Intruder test design: models used, prompt templates, sample size, attributes targeted, baseline and lift.
  • Cue taxonomy from triage and the remediations applied.
  • For health data, the Safe Harbor attestation or Expert Determination report, and whether it considered LLM-based inference.
  • Contract terms prohibiting re-identification attempts and onward linkage.

The full document set is in the de-identification evidence package checklist, and broader methods are in re-identification risk assessment for licensed datasets. For definitions, see the SourceX glossary entry on re-identification and the explainer is de-identified data truly anonymous?.

When SourceX sources operational text from US companies, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, while recognizing that no method is perfect. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Buyers can describe the text data they need at the SourceX buyer page.

Sourcing de-identified text with LLM-era testing in mind

SourceX sources operational datasets, including support and sales histories and other business text, on request from US companies; every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Describe the corpus and the testing evidence you need at https://sourcex.si/buyers.

Frequently asked questions

Can an LLM recover a name that was redacted?

Rarely directly; the realistic risk is that the model infers enough attributes (role, employer, place, timing) to shrink the candidate pool until a search or a second dataset resolves the person [7].

Does passing HIPAA Safe Harbor mean the text is safe from LLM inference?

No. Safe Harbor removes 18 listed identifiers [5], but free-text narrative can still carry rare events and context that a model links to one person, so an intruder test is a useful supplement.

Is the motivated intruder test a legal requirement?

The ICO recommends it as a starting point for UK identifiability assessments [2]; other regimes use different wording, such as GDPR's "means reasonably likely to be used" [3] and HIPAA's "reasonable basis" standard [4].

Sources

  1. arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
  2. Information Commissioner's Office (ICO), "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  3. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation), Recital 26" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  4. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (current). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  5. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  6. National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
  7. Rocher, Hendrickx, de Montjoye, Nature Communications, "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
  8. Carrell et al., Journal of the American Medical Informatics Association, "Hiding in plain sight: use of realistic surrogates to reduce exposure of protected health information in clinical text" (2013). https://academic.oup.com/jamia/article/20/2/342/897812
  9. Microsoft presidio project (indexed on pkg.go.dev), "Presidio - Data Protection API". https://data-privacy-stack.github.io/presidio
  10. Michigan Journal of Environmental & Administrative Law (mjeal-online.org), "Clarke, Spring 2026 article on de-identified educational data" (2026). https://www.mjeal-online.org/2026/04/12/clarke-spring-2026/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data