Privacy, de-identification and sensitive data
The de-identification evidence package: documents to request before licensing sensitive data
Quick answer
A de-identification documentation checklist for a licensed dataset should yield seven documents about that exact release: a field inventory, a method statement with tool versions, measured QA results, an expert or risk-assessment report, a residual-risk statement, key-custody and recipient controls, and a signed attestation tied to the delivered files. A privacy policy or a description of the supplier's usual process proves none of these. Each document should answer the legal test the supplier says the data meets.
By SourceX Editorial · Updated
Why a privacy policy cannot show that a dataset is de-identified
De-identification standards are tests applied to one dataset in one context, so the evidence must describe that dataset. HIPAA asks whether there is no reasonable basis to believe the information can identify someone [1]. GDPR Recital 26 asks what means are reasonably likely to be used to identify a person [2], and the CCPA adds a public commitment and contracts that bind every recipient [3]. A third-party analysis of one health company's privacy policy notes that the adequacy of its de-identification cannot be judged from the policy alone [4].
Tools do not close the gap. The open-source Presidio SDK warns that, because it relies on trained models, there is no guarantee it finds all sensitive information [5], and the DICOM standard says its confidentiality profiles do not guarantee removal of all identifying information [6].
The claimed standard decides which evidence matters most.
| Standard claimed | Legal test, as of October 2026 | Evidence that answers it |
|---|---|---|
| HIPAA Safe Harbor, 45 CFR 164.514(b)(2) | 18 identifiers of the patient, relatives, employers and household members removed; no actual knowledge that the rest could identify the individual [1][7] | Field-by-field removal map, free-text QA, signed no-actual-knowledge statement |
| HIPAA Expert Determination, 164.514(b)(1) | A qualified expert finds a very small risk for the anticipated recipient and documents the methods and results [1]; OCR sets no numerical threshold [7] | The report, the expert's qualifications, and its conditions on you |
| CCPA "deidentified", Civil Code 1798.140(m) | Reasonable measures, a public commitment not to re-identify, contracts obligating recipients [3] | Method and QA evidence, the public commitment, the flow-down clause |
| GDPR anonymous information, Recital 26 | Not identifiable by means reasonably likely to be used; pseudonymized data (Article 4(5)) that can be re-attributed with additional information stays personal data [2] | An identifiability assessment covering linkage and data you already hold |
| UK GDPR | The ICO recommends the motivated intruder test as a starting point [8]; its March 2025 guidance is under review after the Data (Use and Access) Act 2025 [9] | A dated motivated-intruder assessment |
| FERPA, 34 CFR 99.31(b) | A reasonable determination that students are not identifiable, considering multiple releases and other reasonably available information [10] | The written determination and rules for any research record code |
Compare the labels in de-identified, anonymized, pseudonymized and aggregated data and the wider topic in the de-identified data buyer's guide.
The seven documents and what each must contain
A complete package holds seven documents, each with a named owner and all naming the same dataset ID and version.
| # | Document | Must contain | Owner | Weak version to reject |
|---|---|---|---|---|
| 1 | Field inventory | Dataset ID and version, counts, source systems; every column, free-text field, attachment and metadata field classed as direct identifier, quasi-identifier, sensitive or other; data subject locations | Supplier data owner | A "PII fields removed" list with no inventory of what remains |
| 2 | Method statement | Standard claimed; the transformation for each field (drop, generalize, date-shift, surrogate, keyed token, suppress); detector, version, entity types and thresholds for free text; image, audio and metadata handling | Supplier engineer, privacy lead | "Industry-standard anonymization" |
| 3 | QA report | Sample design and size, labelers and adjudication, recall per entity type, residual identifiers found and fixed, over-redaction rate | Reviewer independent of the pipeline | One accuracy figure scored against the tool's own output |
| 4 | Expert or risk-assessment report | Methods, results, recipient and environment assumptions, date, conditions | Named expert or disclosure review function | A report for another dataset version or recipient |
| 5 | Residual-risk statement | Identifying information retained and why, known linkage sources, re-review triggers | Supplier privacy lead | "No residual risk" |
| 6 | Key-custody record and recipient controls | Who holds any re-identification code, crosswalk, hashing key or date-offset table; obligations that flow to you | Supplier security lead, counsel | Keys held by an unnamed processor |
| 7 | Attestation and manifest | Signed statement that the delivered files, identified by hash, are the files assessed | Accountable officer | No file identifiers |
Consent and notice records (what to request) and rights, quality and security evidence (evidence to request from a data vendor) are reviewed alongside this package.
A method statement a reviewer can audit field by field
A method statement is auditable when a reviewer can pick any inventory field and see what happened to it, with which tool, version and settings. For free text, record the detector family (rules, named-entity recognition or an LLM), version and entity list; research on LLM-based redaction reports that LLMs can catch PII types that evade pattern matching and traditional NER [11]. For DICOM images, ask which Annex E profile and options were applied, since pixel data is covered only by specific options, and spot-check the de-identification method attributes recorded in each header [6].
Illustrative example: invented to show structure; it does not describe an available dataset.
Method statement excerpt for a support-ticket dataset claiming the CCPA standard:
| Field | Class | Transformation | Settings recorded |
|---|---|---|---|
| customer_name | Direct identifier | Surrogate name, consistent within a ticket thread | Surrogate list version, run date |
| customer_email | Direct identifier | Dropped | Inventory row reference |
| account_id | Direct identifier | Keyed hash (HMAC-SHA-256); key not delivered | Key holder and storage location |
| created_at | Quasi-identifier | Shifted by a per-customer offset of up to 30 days; intervals kept | Date the offset table was destroyed |
| ship_to_zip | Quasi-identifier | First three digits kept; low-population areas set to 000 | Population table and cutoff used |
| ticket_body | Free text | Rules plus NER; PERSON, EMAIL, PHONE, ADDRESS and CARD replaced with typed placeholders | Detector name, version, confidence threshold |
Two rows would change status under other standards. A keyed hash of an account number is derived from information about the person, so it would not qualify as a HIPAA re-identification code under 164.514(c) [1]. While the supplier holds the key, the token is pseudonymized data under GDPR Article 4(5), not anonymous information [2]; its status in your hands is a separate question (pseudonymized data from the recipient's side).
QA results: the numbers that show what the pipeline missed
QA evidence answers one question: on a labeled sample of this dataset, how many identifiers of each type did the pipeline miss? A credible report states:
- Sampling. A random sample stratified by source system, document type and period, with its size and dataset version.
- Reference labels. Made by people who did not build the pipeline, double-annotated and adjudicated.
- Recall by entity type. Names, emails, phone, account numbers, addresses and birth dates reported separately with confidence intervals, plus residual identifiers per 10,000 records.
- Over-redaction. How much non-identifying content was masked (measuring over-redaction).
- Fix-and-rerun record. Changes made after misses, and the re-measured result.
A vendor blog notes that Presidio publishes no official accuracy benchmark and expects users to measure on their own data [12], so a tool's published scores never replace a report on the supplier's records. NIST SP 800-188 recommends a de-identification standard with measurable performance levels [13]. Pipelines and metrics are compared in PII redaction for LLM training data.
The expert report and the residual-risk statement
A residual-risk statement is the supplier's written account of what identifying information remains, why it was kept, and when the assessment stops holding. It should rest on an expert or risk-assessment report showing how the risk was measured.
For HIPAA Expert Determination, check the expert's qualifications, the documented methods and results, and the anticipated recipient against 164.514(b)(1) [1]. OCR requires no expiration date but says changes in technology, social conditions and available data can make re-examination appropriate [7], so compare the report date with the dataset version (how to review an Expert Determination report).
Outside HIPAA, ask which metrics were used. K-anonymity makes each record indistinguishable from at least k-1 others on its quasi-identifiers, though Sweeney showed such releases can still be attacked unless accompanying policies are respected [14]. Rocher and colleagues' generative model estimated that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes, an estimate rather than a count [15]. NIST SP 800-188 recommends re-identification studies to gauge residual risk [13]; see the risk assessment methods a buyer can request.
Illustrative example: invented to show structure; it does not describe an available dataset.
residual_risk_statement:
dataset: support-tickets-r2 # same ID and version as the method statement
standard_claimed: CCPA 1798.140(m) deidentified # GDPR anonymity not claimed
assessment:
quasi_identifiers: [age_band, ship_to_zip3, product_line, plan_tier]
metric: k-anonymity, k >= 5 after suppressing or generalizing 412 records
attack_test: motivated-intruder exercise against public reviews; 200 targeted records, 0 matched
assessed_by: privacy engineering lead (named in the signed copy)
date: 2026-08-20
retained_on_purpose:
- product_line and plan_tier (needed for intent classification)
- shifted dates with intervals between events preserved
known_linkage_risks:
- recipient's own CRM or support records for the same customers
- public product reviews that quote ticket text
assumptions:
- recipient trains in an access-controlled environment and does not join other customer data
revalidate_when:
- a new dataset version or new fields
- a change in recipient, environment or permitted use
- 12 months after the assessment date
For EU personal data, the statement also feeds your model documentation. Per a law-firm summary of EDPB Opinion 28/2024, whether a model trained on personal data is anonymous must be assessed case by case [16], and a vendor summary notes the EDPB stresses documentation when a supervisory authority evaluates anonymity [17]. As of September 2026, the Digital Omnibus amendments narrowing the GDPR's personal-data definition were still proposals [18]; no statement should rely on them. See training-data extraction and memorization risk.
Key custody, flow-down terms and the signed attestation
The last three documents decide whether de-identification survives the handover.
Key custody. Ask who holds every re-identification code, crosswalk, hashing key, salt and date-offset table, and which were destroyed. HIPAA permits a re-identification code only if it is not derived from information about the individual, is not used or disclosed for other purposes, and its mechanism is not disclosed [1]. Under GDPR, data re-attributable with separately held information remains personal data [2].
Flow-down terms. The CCPA standard requires the supplier to contractually obligate recipients [3], so expect clauses that bar re-identification and linkage, limit onward transfer and require you to report identifiers you find. See re-identification prohibition clauses and the CCPA obligations a buyer inherits. In the UK, the ICO flags that attempting to reverse pseudonymization or ineffective anonymization can be a criminal offense [19].
Attestation. The attestation ties the documents to the bytes you receive: dataset version, manifest file hashes, and a duty to notify you if identifiers turn up. Write it into the license as a warranty (data warranties for AI training licenses).
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
Supplier attests that the files listed in Delivery Manifest [ID], identified by the SHA-256 values in Schedule A, were processed as described in the Method Statement dated [date]; that the QA Report dated [date] was performed on a sample drawn from those files; and that, to Supplier's knowledge after reasonable inquiry, the files contain no direct identifier listed in the Field Inventory except as disclosed in the Residual-Risk Statement. Supplier will notify Licensee within [number] days of learning that any identifier remains in the files.
When to request each document, and who reviews it
Request documents in step with the deal so that nothing you rely on arrives after signature.
| Deal stage | Request | Reviewer |
|---|---|---|
| Shortlist | Standard claimed, data subject locations, one-page method summary, QA metrics available | Procurement lead, privacy counsel |
| Sample under NDA or evaluation license | Field inventory, method statement, de-identified sample records | ML engineer, privacy engineer |
| Before signature | QA report, expert or risk report, residual-risk statement, key-custody record, draft attestation, flow-down clauses | Privacy counsel, security reviewer |
| At delivery | Signed attestation, manifest hashes, your own detector run on delivered files | Data engineering |
| After delivery | Re-review triggers, route for reporting found identifiers | Privacy lead, data owner |
Related: scanning a corpus for PII before fine-tuning, the playbook for found personal data and, for vendor-wide controls, a vendor privacy review.
For datasets sourced through SourceX, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded, a sample is checked, and diligence materials on source, rights, preparation and allowed use are prepared per dataset (how personal details are removed). For health records, SourceX requires HIPAA de-identification by Safe Harbor or Expert Determination. No method is perfect, so state the privacy standard your reviewers need and keep the checks above.
Questions that expose a weak package
Weak packages usually fail on version, scope or file identity, and four direct questions show which.
- Which dataset version does each document describe? A report on version 1 does not cover a version 2 with added fields.
- What happened to free text, attachments and file metadata? Ask separately about ticket bodies, document properties, DICOM private attributes and burned-in pixel text [6].
- Which recipient did the expert assume? An Expert Determination is made for an anticipated recipient [1]; joining other sources may take you outside it (linkage and mosaic risk).
- Will you sign the attestation against file hashes? A refusal means the package describes a process, not your files.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Licensing de-identified records from US businesses?
Describe the records you need, where the people in them live, and the de-identification standard and evidence your reviewers require. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery through private, access-controlled workflows. Share your data and privacy requirements.
Sources
- eCFR (Office of the Federal Register / HHS), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (current text). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)" (current codified text). https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Conduct Atlas, "Hims & Hers Privacy Policy: AI model training using de-identified data (provision analysis)". https://conductatlas.com/platform/hims-hers/hims-hers-privacy-policy/ai-model-training-using-de-identified-data/
- Microsoft presidio project (pkg.go.dev listing), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- NEMA / DICOM Standards Committee, "DICOM PS3.15 - Security and System Management Profiles, Annex E: Attribute Confidentiality Profiles" (current edition). https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Information Commissioner's Office, "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
- Information Commissioner's Office, "Anonymisation guidance: About this guidance" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/about-this-guidance/
- U.S. Government Publishing Office / U.S. Department of Education, "34 CFR 99.31 - Under what conditions is prior consent not required to disclose information?" (CFR 2018 edition). https://www.govinfo.gov/content/pkg/CFR-2018-title34-vol1/pdf/CFR-2018-title34-vol1-sec99-31.pdf
- arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
- Grepture (vendor blog), "Blog post on the limits of Presidio for PII redaction". https://grepture.com/blog/presidio-not-enough-pii-redaction
- National Institute of Standards and Technology and U.S. Census Bureau, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- Latanya Sweeney, "k-Anonymity: A Model for Protecting Privacy" (2002). https://dataprivacylab.org/people/sweeney/kanonymity.html
- Rocher, Hendrickx, de Montjoye, "Estimating the success of re-identifications in incomplete datasets using generative models" (Nature Communications, 2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
- CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (law firm summary). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
- Securiti (vendor summary), "Summary of EDPB Opinion 28/2024 concerning AI models and the processing of personal data". https://securiti.ai/summary-of-edpb-opinion-282024-concerning-ai-models-processing-of-personal-data
- Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/
- Information Commissioner's Office, "Pseudonymisation" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.