Skip to content

Privacy, de-identification and sensitive data

Re-identification risk assessment for licensed datasets: methods a buyer can request

Quick answer

A re-identification risk assessment estimates how likely it is that someone holding a de-identified dataset could single out, link or infer facts about a real person, given the data, its context and who receives it. Before licensing, ask the supplier for four things: a quasi-identifier inventory with uniqueness or k-analysis results, a linkage simulation against named auxiliary sources, a documented motivated-intruder test, and a report stating the threat model, risk thresholds, residual risks and the controls the result depends on.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why "de-identified" means nothing without a risk assessment

A de-identification label tells you a process was applied, not how much identifiability remains. Legal standards are framed as risk judgments: HIPAA asks whether there is "no reasonable basis to believe" the information can identify someone [2], GDPR Recital 26 tests anonymity against "all the means reasonably likely to be used" [4], and the CCPA definition of deidentified information attaches ongoing conditions to the business that holds it [9]. Each standard depends on facts a buyer cannot see from a data dictionary alone.

The failure mode is well documented. Rocher, Hendrickx and de Montjoye showed that a generative model can estimate the uniqueness of individuals even in heavily sampled datasets, so releasing "only a sample" is not a defense on its own [7]. Legal commentary on de-identified education records makes the same point: auxiliary information held outside the dataset can undo de-identification [10]. For the conceptual background, see our re-identification glossary entry and re-identification risk in business data.

Methods a supplier can run, and what each one proves

No single method is sufficient; a credible assessment combines a structural metric, a linkage test and a human or adversarial test. The table below maps each method to the question it answers and the artifact you should receive.

Illustrative example: invented to show structure; it does not describe an available dataset.

MethodQuestion it answersArtifact to requestCommon weakness
Quasi-identifier inventoryWhich fields, alone or combined, could link to outside data?Field list with QI classification and rationale (e.g., ZIP3, birth year, job title, account open date)Free-text fields and timestamps left off the list
Uniqueness and k-anonymity analysis [8]How many records are unique or in small equivalence classes on the QI set?Distribution of equivalence-class sizes, % records with k < threshold, max and average riskComputed on the sample, not estimated for the population [7]
l-diversity / t-closeness checksCan a sensitive value be inferred even when k is met?Sensitive-attribute diversity per classSkipped entirely for "non-sensitive" business data
Linkage simulationCan records be joined to a realistic auxiliary dataset?Named auxiliary sources, join keys tried, match rate and precisionUses an auxiliary set the attacker would never hold, or a weaker one
Motivated intruder test [3]Could a competent, resourced person with no special access identify someone?Test protocol, intruder profile, time budget, outcome per targetRun by the same team that de-identified the data
LLM-assisted intruder testCan a language model with web search infer identity from text context?Prompts, model and tools used, hit rate on seeded targetsTreated as complete; this is an emerging practice without a standard protocol
Formal privacy accountingIs there a mathematical bound on disclosure?Mechanism, epsilon/delta, composition across releases [6]Parameters too loose to mean anything

Structural metrics such as k-anonymity were built for tabular data [8]. For transcripts, tickets and documents, the quasi-identifiers hide in prose: a rare job title, a store location, a dated incident. That is why text-heavy data needs the adversarial tests, covered in more depth in LLM-assisted re-identification of de-identified text.

Context, recipient and environment change the score

The same file can be low-risk in one environment and high-risk in another, so the assessment must name who receives the data and under what controls. HHS OCR's guidance on Expert Determination tells the expert to consider the anticipated recipient and the context of disclosure, and lists principles such as replicability, data source availability and distinguishability [1]. The ICO frames identifiability the same way, as a risk judgment about the means reasonably likely to be used, including by a motivated intruder [3].

For an AI buyer, the recipient environment includes things the supplier may not anticipate:

  • Model training and outputs. Memorized strings can surface at inference time; see training-data extraction and memorization risk.
  • Joining with your other data. Your CRM exports, web data or other licensed corpora are auxiliary data. Combining sets is covered in linkage and mosaic risk.
  • Access breadth. A dataset in a locked research enclave and the same dataset copied to a shared bucket for an annotation vendor are different risk cases.

The EDPB's Opinion 28/2024 takes a case-by-case view of whether a trained model itself is anonymous [5], which means your downstream use can reopen the question the supplier's report closed. Tell the supplier your intended environment before the assessment is run, not after.

What a re-identification risk report should contain

A usable report is reproducible: another expert should be able to read it and reach the same conclusion. Ask for these sections.

Illustrative example: invented to show structure; it does not describe an available dataset.

Re-identification risk report: minimum contents checklist

  1. Scope and version. Dataset name, extract date, record and field counts, file hashes so the report binds to the exact delivery.
  2. De-identification methods applied. Suppression, generalization, date shifting, surrogate substitution, tokenization, with parameters and the tools used.
  3. Threat model. Attacker types considered (prosecutor, journalist, marketer), their assumed knowledge and the auxiliary sources named.
  4. Recipient and environment assumptions. Who gets the data, storage and access controls, whether joins are permitted.
  5. Quasi-identifier inventory with justification for any field excluded.
  6. Metrics and thresholds. Equivalence-class distribution, maximum and average risk, the threshold chosen and why.
  7. Adversarial test results. Linkage simulation and motivated-intruder protocol, effort spent, targets attempted, outcomes.
  8. Free-text findings. Detector recall on a hand-labeled sample, with the residual PII rate and examples of misses (redacted).
  9. Residual risks. What remains and why it is accepted.
  10. Required controls. The conditions the conclusion depends on, such as no linkage, no re-identification attempts, access restrictions, retention limits.
  11. Validity period and triggers. When the assessment must be rerun (new fields, new releases, new auxiliary data).
  12. Assessor identity and independence. Qualifications and relationship to the data holder.

Item 10 matters most to counsel, because those controls become contractual obligations. Compare them with the language you will be asked to sign in re-identification prohibition clauses. For the wider document set, use the de-identification evidence package checklist.

How to read the metrics without being misled

Risk numbers are only meaningful against a stated threshold, population and attacker, so read every metric with its assumptions attached. Four checks catch most weak reports.

  • Sample versus population. A record unique in a 2% sample may or may not be unique in the population. Ask whether uniqueness was estimated for the population, as in the Rocher model, or simply counted in the file [7].
  • Average versus maximum risk. A low average can hide a tail of unique records. Ask for the share of records above the threshold, not just the mean.
  • Weak auxiliary assumptions. If the linkage simulation used only public voter-style fields while your environment holds richer data, the result understates your risk.
  • Formal guarantees with no parameters. "Differential privacy applied" without epsilon, delta and composition accounting is not evidence; NIST SP 800-188 contrasts formal privacy methods such as differential privacy with the limits of traditional de-identification [6].

For health data, the assessment sits inside a defined legal route: Safe Harbor's identifier removal or an Expert Determination [1][2]. See how to review a HIPAA Expert Determination report for that specific document.

Requesting the assessment during diligence

Ask for the assessment at the same time you ask for the sample, and state your threat model in the request so the supplier does not test against an easier one. A short request often gets a better answer than a long questionnaire.

Illustrative example: invented to show structure; it does not describe an available dataset.

Request: re-identification risk assessment Dataset: de-identified support ticket histories, 2021-2025, US customers. Intended use: supervised fine-tuning; data held in a private training environment, no joins to customer or marketing data; annotation by an internal team only. Please provide: QI inventory; equivalence-class distribution on the agreed QI set; linkage simulation against named public and commercial sources; motivated-intruder results on 50 sampled tickets including free text; residual-risk statement; controls the conclusion depends on; assessor qualifications; validity period. Question: which fields or text patterns drove the highest-risk records, and what was done about them?

How SourceX approaches this: SourceX rights-reviews every dataset for ownership and consents, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked; no method is perfect. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Buyers can describe the data they need through the SourceX buyer intake.

For where this assessment fits among other pre-acquisition checks, see training data risk assessment before acquisition and the privacy cluster guide. Whether de-identified data is ever truly anonymous is discussed in is de-identified data truly anonymous?.

Get a de-identified dataset with diligence prepared

SourceX sources operational datasets from US companies on request, rights-reviews each one and delivers it under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery and the method is recorded. Describe the data you need at sourcex.si/buyers.

Frequently asked questions

What is a motivated intruder test?

It is a structured attempt by a competent person, using resources reasonably available to them but no special access or insider knowledge, to identify individuals in a dataset. The ICO recommends it as a starting point for assessing identifiability [3]. Ask who ran it, how long they spent and which targets they tried.

Is k-anonymity enough for text data?

No. k-anonymity measures indistinguishability on a defined set of structured quasi-identifiers [8]. Free text contains identifying context that does not map to fixed columns, so it needs detector recall measurement and adversarial testing alongside any structural metric.

Does the assessment stay valid after we train a model?

Not automatically. The supplier assessed the data in an assumed environment; training, joins and output exposure change that environment, and the EDPB treats model anonymity as a separate, case-by-case question [5]. Plan your own reassessment when your use changes.

Sources

  1. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  2. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  3. Information Commissioner's Office, "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  4. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation), Recital 26" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  5. CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  6. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  7. Nature Communications (Rocher, Hendrickx, de Montjoye), "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
  8. Data Privacy Lab (Latanya Sweeney), "k-Anonymity: A Model for Protecting Privacy" (2002). https://dataprivacylab.org/projects/kanonymity/
  9. California Legislative Information, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  10. Michigan Journal of Environmental & Administrative Law, "Clarke, Spring 2026 (re-identification of de-identified education data)" (2026). https://www.mjeal-online.org/2026/04/12/clarke-spring-2026/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data