Skip to content

Privacy, de-identification and sensitive data

De-identified data for AI training: a buyer's guide to privacy-safe licensed datasets

Quick answer

De-identified data for AI training is personal data transformed so that it no longer identifies people under a specific legal test, and the test changes with the regime. HIPAA accepts Safe Harbor or Expert Determination, the GDPR still treats pseudonymised data as personal data, and the CCPA requires the holder to bind every recipient by contract. Before licensing, ask which standard was met, by what method, with what measured miss rate and what you must sign, then plan for risk left in the trained model.

By SourceX Editorial · Updated

A dataset described as "de-identified" can pass one regime's test and fail another's, because HIPAA, the GDPR and the CCPA ask different questions about the same records. The table states each test as of October 2026.

RegimeTest the data must meetStill regulated?What the recipient takes onRead next
HIPAA Safe Harbor, 45 CFR 164.514(b)(2)Remove 18 identifiers of the person and of relatives, employers and household members; no actual knowledge that the rest identifies anyone [1][2]No; not protected health information (PHI) [1]Nothing under HIPAA; license terms applySafe Harbor vs Expert Determination
HIPAA Expert Determination, 164.514(b)(1)An expert finds a "very small" risk that an anticipated recipient could identify anyone and documents the analysis [2]; no numeric threshold [1]No, while its assumptions holdConditions in the report, such as access limitsReviewing an Expert Determination report
HIPAA limited data set, 164.514(e)Direct identifiers removed; dates, town or city, state and ZIP code may remain [2]Yes; still PHI, only for research, public health or health care operations [2]A data use agreement, including no identifying or contacting individuals [2]Limited data sets and DUAs
GDPR anonymous information, Recital 26No one identifiable by means reasonably likely to be used [3]NoLittle, if the test genuinely holdsAnonymised or pseudonymised?
GDPR pseudonymised data, Article 4(5)Attributable to a person only with separately held additional information [3]Yes for the controller; for recipients, see EDPS v SRB [4]Lawful basis, Article 9 condition for health or biometric data, transfer safeguards [3]Pseudonymised data from the recipient's side
CCPA "deidentified", Cal. Civ. Code 1798.140(m)Not reasonably linkable to a consumer; the holder takes reasonable measures, publicly commits not to re-identify and contractually binds recipients [5]No, while all three conditions holdThe same commitments, by contract [5]CCPA obligations a buyer inherits

In the EU, the Court of Justice held in EDPS v SRB (C-413/23 P, 4 September 2025) that sufficiently strong pseudonymisation can leave data personal for the controller holding the key but not for a recipient with no means to reverse it or identify people otherwise; the assessment is case by case [4]. As of September 2026, Digital Omnibus proposals to narrow the definition of personal data and to confirm legitimate interest for AI development were still proposals, not law [6].

In the UK, the ICO's anonymisation guidance (published March 2025; under review) recommends the "motivated intruder" test [7][8] and warns that re-identifying people from pseudonymised or ineffectively anonymised data can be a criminal offence [9]. Compare the four labels in de-identified, anonymised, pseudonymised and aggregated; SourceX's glossary defines data anonymization.

Sector and state rules that follow particular records

Some records carry rules that a general de-identification claim does not settle, so identify the source system and the people in it first.

  • Education records. FERPA allows release without consent once all personally identifiable information is removed and the releaser reasonably determines, across multiple releases and other available information, that no student is identifiable [10] (FERPA guide).
  • Financial records. Nonpublic personal information received under a Regulation P exception may be used only for the purpose it was received, even by a recipient that is not a financial institution [11]; ask counsel whether de-identified transactions still qualify (GLBA guide).
  • Substance use disorder records. HHS's 2024 final rule aligned parts of 42 CFR Part 2 with HIPAA, with a compliance date of 16 February 2026 [12]; ask whether any records came from a Part 2 program (Part 2 records).
  • Biometrics. Illinois' BIPA treats voiceprints and face geometry scans as biometric identifiers, requires a written release before collection and bars profiting from them [13]. Removing names does not help when the voice or face is the identifier (biometric data guide).
  • Washington consumer health data. The My Health My Data Act counts inferences drawn from non-health data as consumer health data, and selling it needs a separate authorization that seller and purchaser must both keep for six years [14].

Why stripping identifiers does not make data anonymous

Removing names and account numbers leaves quasi-identifiers, such as dates, places, job titles and rare events, that outside data can link back to a person. Sweeney's k-anonymity requires each person to be indistinguishable from at least k-1 others on those attributes, and even then attacks can succeed when release policies are ignored [15]. Rocher et al. estimated with a generative model, not from actual re-identifications, that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes, so releasing only a sample does not make data anonymous [16].

NIST documents real re-identification cases and notes that medical text, photographs and genetic data are harder to de-identify than tables [17]. SP 800-188 (2023) cautions about the limits of traditional methods compared with formal ones such as differential privacy [18]. Watch for indirect identifiers in business text and linkage risk when combining licensed sets; SourceX's explainer is de-identified data truly anonymous? gives the short version.

Where identifiers hide in text, audio, images and tables

The method has to match the modality, because identifiers sit in different places in a support ticket, a call recording, a video frame and a claims table. Test each type's weak point in a sample.

Data typeWhere identifiers hideTypical methodKnown weak pointRead next
Free text: tickets, email, chat, clinical notesNames, contacts and account numbers in prose, signatures, quoted headersEntity detection, then masks or consistent surrogatesDetectors miss entities; Presidio says it cannot guarantee finding all sensitive data [19]PII redaction for LLM data; masking vs surrogates
Call and meeting audioSpoken names and card numbers; the voice itselfAudio masking aligned to transcripts; speaker anonymizationIn one benchmark, a naive attacker's equal error rate (higher is more private) reached about 52.12%; an attacker trained on anonymized speech cut it to about 10.7% [20]Spoken PII in call recordings; speaker anonymization
Images and videoFaces, plates, screens and papers in frame; EXIF GPS tagsBlur, pixelation or synthetic replacementMissed small or occluded faces; Gaussian blur can be partly reversed [21]Video anonymization; faces: consent or anonymization
Medical images (DICOM)Header attributes, private tags, burned-in pixel textPS3.15 Annex E confidentiality profilesProfiles do not guarantee removal of all identifying information [22]HIPAA-compliant training data
Scanned documentsPixels, the OCR text layer, file metadataRedact all three layersBoxes over the image while the OCR layer stays searchablePII in scanned documents
Tables and transactionsQuasi-identifier and free-text columns, stable keysGeneralization, suppression, keyed tokensRare value combinations stay unique [16]De-identifying tabular data
Screen recordings and agent trajectoriesOn-screen records, typed input, URLs, tool-call argumentsFrame- and event-level redactionText visible in frames but absent from event logsScreens and computer-use trajectories

Privacy risk that carries into the trained model

Training can memorize personal data that survived de-identification, so the model itself can become the disclosure. Carlini et al. extracted hundreds of verbatim training sequences from GPT-2, including contact details, using only queries [23], and Nasr et al. (ICLR 2025) recovered thousands of training examples from aligned production models [24]. Alignment is not a privacy control.

A law-firm summary of EDPB Opinion 28/2024 (17 December 2024) reports that a model trained on personal data cannot be presumed anonymous; it is anonymous only if extracting personal data from it, directly or through queries, is insignificant [25] (EDPB Opinion 28/2024 for data buyers). In the US, the FTC's 2021 Everalbum order required deletion of models and algorithms built from users' photos and videos [26].

Residual personal data therefore decides whether you can ship weights or open a public API. Differential privacy bounds what any single record can reveal, usually at some cost in accuracy [18]; see memorization risk, differential privacy for LLM fine-tuning and releasing weights trained on personal data.

The privacy evidence to request before you sign

Ask for evidence that a named standard was met on this dataset, not a description of the supplier's general process.

  1. Standard and jurisdiction. The regime claimed and where the data subjects live.
  2. Method. Detector and version, entity types, masking or surrogate policy, date and location rules, and dropped fields.
  3. Measured misses. Recall per entity type on a labeled sample from this dataset, with sample size and labeler.
  4. Expert Determination report, if used. Methods, results, recipient assumptions and date [2]; HHS notes that because technology and information availability change over time, re-examination may be appropriate [1].
  5. Key custody. Who holds any re-identification code or pseudonymisation key; HIPAA permits such codes only under stated conditions [2].
  6. Residual risk. Quasi-identifiers left in and any re-identification test results; SP 800-188 recommends such studies and a disclosure review board [18].
  7. Your obligations. No re-identification, no linkage, onward-transfer limits and a duty to report found personal data (re-identification prohibition clauses).
  8. Source-side basis. Notice and consent records, and any Article 9 condition for special categories (consent and notice records).

Illustrative example: invented to show structure; it does not describe an available dataset.

dataset_id: support-tickets-r1
source_system: helpdesk export (ticket bodies and replies; attachments excluded)
standard_claimed: CCPA 1798.140(m) deidentified   # GDPR status: not claimed
data_subject_locations: [US-CA, US-TX, US-NY]
method:
  detection: NER model + regex rules (card numbers, SSNs)
  entity_types: [PERSON, EMAIL, PHONE, ACCOUNT_ID, STREET_ADDRESS, CARD_NUMBER]
  replacement: consistent surrogates within each ticket thread
  dates: shifted per customer; intervals preserved
  dropped_fields: [attachments, agent_email, ip_address]
qa_sample:
  records_reviewed: 2000
  labeling: two annotators, adjudicated
  recall_by_entity: {PERSON: 0.97, EMAIL: 1.00, PHONE: 0.99, ACCOUNT_ID: 0.95}
residual_risk_notes: product names and small-town names retained
reidentification_key: none retained by supplier
recipient_obligations: [no re-identification, no linkage to customer records, report found personal data]

For datasets sourced through SourceX, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded per dataset, and a sample is checked after processing; health records must meet HIPAA Safe Harbor or Expert Determination before a license is considered. Diligence materials on source, rights, preparation and allowed use are prepared per dataset. No method is perfect, so keep your own checks and describe the records and privacy standard you need. See the full de-identification evidence package checklist; suppliers can follow SourceX's playbook for de-identifying company data.

Controls that stay with the buyer after delivery

Supplier de-identification lowers risk at the source; your controls decide whether it stays low once the data meets other data and a model.

  • Scan before training. Check inputs and context fields, not only targets (scanning a corpus before fine-tuning).
  • Block linkage. Keep licensed sets out of joins with CRM data, logs or other licensed sets unless a fresh risk assessment covers the combination.
  • Contain access. SP 800-188 lists protected, non-public enclaves among its data-sharing models [18]; train inside one where you can.
  • Test the model. Run extraction probes or planted canary strings before any external release.
  • Rehearse the incident. Agree how found personal data is quarantined, reported and deleted (response playbook).

What de-identification costs the model, and how to specify around it

Every method removes signal, so list the fields your model needs before choosing a standard. Safe Harbor removes all elements of dates tied to the person except the year and geographic units smaller than a state (three-digit ZIP codes for geographic units containing more than 20,000 people may be retained), and groups ages over 89 [2]. That can break a readmission or seasonal-demand model; Expert Determination has no fixed list, so an expert may allow more detail if the risk stays very small [1][2] (what Safe Harbor dates and ZIP codes cost).

In vision, a CVPR 2023 workshop study found that face-only anonymization of detection datasets caused a minimal accuracy drop, while traditional anonymization hurt noticeably, especially when whole bodies were masked [27] (blurring evidence). Privacy constraints also narrow coverage: a paper in Stanford's GRACE journal argues that de-identified health datasets used for foundation models are often outdated and demographically limited [28]. For high-risk systems, EU AI Act Article 10(3) requires data that is sufficiently representative of the people the system serves [29], so check coverage after de-identification; the compliance hub gives application dates as amended in 2026.

Start here: privacy guides by question

Each question below has its own guide.

Your questionRead
Business associate agreement or data license?BAA vs data license for AI developers
How is re-identification risk measured?Risk assessment methods to request
Is synthetic data from licensed records still personal data?Synthetic data privacy risk
US-sourced data with EU or UK subjects?Transfer mechanisms for AI teams
Which US state standards apply?State deidentification standards compared
Can anonymized business data be licensed at all?Can anonymized data be licensed?

For the rest of the purchase, start at the AI data buyer's hub.

Mistakes that turn a de-identified dataset into a liability

Most privacy failures in data deals come from trusting a label instead of the test behind it.

  • Treating HIPAA de-identification as GDPR anonymisation. Safe Harbor answers a US question; any records in the same file that fall under the GDPR still face the Recital 26 test [3].
  • Assuming pseudonymised data is non-personal for you. After EDPS v SRB the question is whether you could re-identify, including with data you already hold [4].
  • Signing flow-down terms you cannot honor. Joining records to customer data puts the CCPA no-re-identification commitment at risk [5].

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Sourcing data that contains personal information?

Describe the records you need, where the people in them are located, and the privacy standard your reviewers require. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery through private, access-controlled workflows. Share your data and privacy requirements.

Guides in this section

Sources

  1. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  2. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (current text). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  3. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)". https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  4. Bird & Bird, "EU: The SRB decision - a new era for personal data and data processing agreements" (2025). https://www.twobirds.com/en/insights/2025/eu-the-srb-decision-a-new-era-for-personal-data-and-data-processing-agreements
  5. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  6. Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/
  7. Information Commissioner's Office, "Anonymisation guidance: About this guidance" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/about-this-guidance/
  8. Information Commissioner's Office, "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  9. Information Commissioner's Office, "Pseudonymisation" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
  10. U.S. Government Publishing Office / U.S. Department of Education, "34 CFR 99.31 - Under what conditions is prior consent not required to disclose information?" (CFR 2018 edition). https://www.govinfo.gov/content/pkg/CFR-2018-title34-vol1/pdf/CFR-2018-title34-vol1-sec99-31.pdf
  11. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  12. U.S. Department of Health and Human Services, Federal Register, "Confidentiality of Substance Use Disorder (SUD) Patient Records (Final Rule)" (2024). https://www.govinfo.gov/content/pkg/FR-2024-02-16/html/2024-02544.htm
  13. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  14. Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
  15. Latanya Sweeney, "k-Anonymity: A Model for Protecting Privacy" (2002). https://dataprivacylab.org/people/sweeney/kanonymity.html
  16. Rocher, Hendrickx and de Montjoye, "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
  17. National Institute of Standards and Technology (Garfinkel), "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
  18. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  19. Microsoft presidio project, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  20. arXiv:2109.00281, "Benchmarking and challenges in security and privacy for voice biometrics" (2021). https://arxiv.org/pdf/2109.00281
  21. arXiv:2512.16086, "Privacy Blur: Quantifying Privacy and Utility for Image Data Release" (2025). https://arxiv.org/pdf/2512.16086
  22. NEMA / DICOM Standards Committee, "DICOM PS3.15 Security and System Management Profiles, Annex E: Attribute Confidentiality Profiles" (current edition). https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html
  23. Carlini et al., "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  24. Nasr et al., "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
  25. CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  26. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  27. Hukkelås and Lindseth, "Does Image Anonymization Impact Computer Vision Training?" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
  28. Stanford University Open Journal Systems (GRACE), "Article on de-identified datasets in foundation-model training (title to confirm)". https://ojs.stanford.edu/ojs/index.php/grace/article/download/3837/1799/11712
  29. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data