Privacy, de-identification and sensitive data
De-identified, anonymized, pseudonymized, aggregated: one dataset, different legal status in the US and EU
Quick answer
"De-identified" and "anonymized" are not synonyms. De-identified is a US statutory status defined separately by HIPAA, the CCPA and other state laws, each with its own test and conditions. Anonymous is the EU threshold under GDPR Recital 26, where data leaves the regulation only if a person can no longer be identified by means reasonably likely to be used. The same licensed file can be de-identified under HIPAA, still personal information under a state law, and pseudonymized personal data under GDPR, so classify it per regime before signing.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Four labels, four different legal questions
Each label answers a different question, so a supplier's single word on a data sheet tells you little until you know which regime it was measured against. "De-identified" asks whether a statutory test was met by the party holding the data. "Anonymous" asks whether anyone, using reasonably likely means, can still identify a person. "Pseudonymized" describes a technique that keeps identity recoverable with separately held information. "Aggregated" describes a shape (counts, rates, group statistics) rather than a legal status.
The practical consequence for an AI buyer is that the label travels with the file but the status does not. Status depends on who holds the data, what auxiliary information they have, what contract binds them and where they process it. For a broader orientation on privacy-safe sourcing, start with the privacy and de-identification buyer's guide; for the plain-language term, see the data anonymization glossary entry.
HIPAA: de-identified means outside PHI, but only by one of two routes
Under HIPAA, health information is de-identified when it does not identify an individual and there is no reasonable basis to believe it can be used to identify one, and once that standard is met the data is no longer protected health information [1]. The rule allows exactly two implementation routes in 45 CFR 164.514(b): Expert Determination, where a qualified expert documents that re-identification risk is very small for the anticipated recipient, and Safe Harbor, which removes 18 listed identifiers and requires that the covered entity has no actual knowledge the remainder could identify someone [1][2].
Three details matter for licensing. Safe Harbor still permits the year of dates and, in some cases, the first three digits of a ZIP code, so free-text clinical notes and rich longitudinal tables often fail it in practice [2]. Section 164.514(c) lets the covered entity keep a re-identification code, provided the code is not derived from patient information and the mechanism is not disclosed [1]. And a limited data set under 164.514(e) is not de-identified at all: it remains PHI, strips 16 direct identifiers, and can only be shared under a data use agreement [1].
If you need to choose between the two methods, the Safe Harbor vs Expert Determination comparison covers the trade-off for training data, and the Expert Determination review guide explains how to read the report itself.
CCPA and other state laws: deidentified is a status with ongoing conditions
California treats information as "deidentified" only when it cannot reasonably be used to infer information about, or be linked to, a particular consumer, and the business holding it meets three conditions [3]. The business must take reasonable measures to prevent association with a consumer or household, publicly commit to keep and use the data in deidentified form without attempting re-identification, and contractually require any recipient to comply with the same obligations [3]. Those conditions run with the data, which means a buyer that receives California deidentified data inherits a contractual no-re-identification duty and should expect a public commitment requirement of its own.
The CCPA defines pseudonymization and aggregate consumer information separately from deidentified information [3]. Pseudonymized data, where identity is recoverable with additional information kept apart, is not carved out of personal information by those definitions, so a buyer should treat a pseudonymous California file as personal information unless counsel concludes otherwise. Aggregate consumer information relates to a group from which individual identities have been removed and that is not linked or reasonably linkable to any consumer or household [3].
Other state comprehensive privacy laws borrow the structure but not the exact words, and the differences in conditions, public-commitment duties and pseudonymous-data treatment are real [4]. Health-specific state laws, the GLBA for financial records and FERPA for education records each add their own definitions; FERPA, for example, permits release of education records without consent only after all personally identifiable information is removed and the institution reasonably determines a student's identity is not personally identifiable [10]. For the California specifics, read CCPA/CPRA deidentified data: the obligations a buyer inherits.
GDPR: anonymous is a high bar, and pseudonymized data is still personal
Under GDPR, data is outside the regulation only if it is anonymous, meaning it no longer relates to an identified or identifiable person, taking into account all the means reasonably likely to be used, by the controller or by another person, to identify someone [5]. Recital 26 states that data which has undergone pseudonymization and could be attributed to a person using additional information should be considered information on an identifiable natural person [5]. Article 4(5) defines pseudonymisation as processing so that data can no longer be attributed to a specific person without additional information that is kept separately and protected by technical and organizational measures [5].
The EDPB's Guidelines 01/2025 on pseudonymisation, adopted in January 2025, restate that pseudonymised data remains personal data and present pseudonymisation as a safeguard rather than an exit from the GDPR [6]. The Court of Justice then added nuance in EDPS v SRB (C-413/23 P, 4 September 2025): sufficiently strong pseudonymised data may be personal data for the original controller but not for a recipient that has no means reasonably likely to re-identify [7]. That recipient-relative view is the basis for the recipient-perspective analysis after EDPS v SRB, and the GDPR anonymised vs pseudonymised classification guide applies it to training files.
Two status notes as of October 2026. The Commission's Digital Omnibus proposal would narrow the personal data definition along recipient-relative lines, but as of September 2026 the GDPR amendments remained proposals and are not law [9]. EDPB Opinion 28/2024 remains the reference for when a trained model itself can be considered anonymous, a question it says must be decided case by case [8].
Aggregated data: personal or not depends on cell size and linkability
Aggregated data is non-personal only when no individual can be singled out or inferred from it, which is a property of the specific tables, not of the word "aggregate." Under GDPR, aggregates fall under the same Recital 26 test, so small cells, rare combinations and repeated releases over time can make an aggregate identifying [5]. Under the CCPA, aggregate consumer information is excluded only if it is not linked or reasonably linkable to any consumer or household [3].
For AI training, true aggregates are rarely what buyers license: model training usually needs row-level records, conversation turns or documents. When a supplier offers "aggregated" data that is actually per-account monthly rollups, treat it as row-level and apply the de-identification analysis. Differencing attacks across overlapping releases and joins with other licensed sources are the common failure modes, covered in linkage and mosaic risk when combining datasets.
One dataset, three regimes: a worked classification
The same file can carry three statuses at once, so map each regime separately and record which obligations attach. The matrix below follows one hypothetical dataset from a US provider network to a buyer with an EU training team.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Dataset state | HIPAA (US covered entity source) | CCPA (California consumers) | GDPR (EU-established buyer processes it) | Obligations a buyer should expect |
|---|---|---|---|---|
| Raw encounter notes with names, MRNs, full dates | PHI | Personal information (if not exempt as PHI) | Personal data; special category (health) | Not licensable for training without authorization, BAA or other legal basis |
| Limited data set: direct identifiers removed, dates and ZIP kept | Still PHI under 164.514(e) [1] | Treated as personal information | Personal data | Data use agreement; no re-identification or contact |
| Safe Harbor output; supplier keeps a 164.514(c) re-identification code | De-identified, not PHI [1][2] | Deidentified only if the three 1798.140 conditions are met [3] | Pseudonymised personal data for the supplier; for the buyer, possibly not personal under EDPS v SRB if it lacks means to re-identify [7] | Public no-re-identification commitment, contractual flow-down, GDPR analysis per recipient |
| Expert Determination output with free text and coarsened dates | De-identified for the anticipated recipient named in the report [2] | Deidentified if conditions met [3] | Anonymous only if Recital 26 test is met for the buyer's actual means [5] | Recipient restrictions from the expert report; linkage limits |
| Published aggregate tables, minimum cell size enforced | Not PHI if the tables meet the 164.514(a) standard [1] | Aggregate consumer information if not reasonably linkable [3] | Anonymous if no singling out [5] | Release controls on repeated queries |
Three lessons follow from the table. First, "HIPAA de-identified" does not answer the GDPR question: GDPR applies to processing in the context of an EU establishment regardless of the data subjects' nationality, and a retained re-identification code at the source is exactly the additional information Recital 26 contemplates [5]. Second, an Expert Determination is scoped to an anticipated recipient and environment, so moving the data to a new team or joining it to another dataset can invalidate the determination [2]. Third, every applicable regime applies at once, so in practice the most demanding one sets the working controls for that processing activity.
Classification checklist for a cross-border license
A short, repeatable checklist keeps the classification honest, and it belongs in the diligence file next to the license. Ask the supplier for the documents behind each answer rather than a yes or no; the de-identification evidence package checklist lists what those documents look like.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Origin regime: Which laws governed the data at the source (HIPAA covered entity or business associate, CCPA business, GLBA financial institution, FERPA institution, 42 CFR Part 2 program)?
- Method and date: Safe Harbor, Expert Determination, CCPA reasonable measures, or a custom method; who applied it and when.
- Keys and codes: Does anyone retain a crosswalk, salted hash secret, token vault or 164.514(c) code? Who controls it?
- Recipient scope: Does the expert report or risk assessment name the anticipated recipient, environment and permitted joins?
- Flow-down terms: No-re-identification, no-contact, onward-transfer and public-commitment duties the buyer must accept.
- Buyer means: What auxiliary data does your organization already hold that could link to this file (CRM exports, public records, other licensed sources)?
- EU processing: Will an EU establishment, EU contractor or EU-hosted infrastructure touch the data? If so, run the Recital 26 and EDPS v SRB analysis for that recipient.
- Free text and media: Have notes, transcripts, audio or screen recordings been tested for residual identifiers, not just structured columns?
- Re-test triggers: Which events (new joins, new recipients, model release) require reassessment?
Where obligations attach in the training lifecycle
Status can change after delivery, so the classification needs revisiting at joins, transfers and model release. Joining a de-identified file to another licensed source, or to your own user data, can make records identifiable again, which restores personal-data status under the CCPA and GDPR and can breach the no-re-identification terms that came with HIPAA-derived data. Moving data to an EU team can bring GDPR into play for a file that was outside it in the US.
The trained model is a separate question. EDPB Opinion 28/2024 says model anonymity is assessed case by case, considering whether personal data can be extracted or obtained through queries [8]; see training-data extraction and memorization risk for the technical side and the re-identification prohibition clause guide for the contract terms you will be asked to sign.
How SourceX handles de-identification in licensed datasets
SourceX sources operational datasets from US companies on request and rights-reviews every dataset for ownership and consents before delivery under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect; health records require HIPAA de-identification by Safe Harbor or Expert Determination. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which gives your counsel the inputs for the regime-by-regime classification above. Buyers can describe the data they need on the SourceX buyers page; a request does not guarantee a match. For the supplier-side question, see can anonymized data be licensed and is de-identified data truly anonymous.
Sourcing de-identified data that holds up across jurisdictions
SourceX looks for US businesses that hold the data you describe and serves AI teams wherever they are based. Nothing is contracted until a supplier agrees, every release is approved by the supplying company, and delivery runs through private, access-controlled workflows after an executed agreement. To start, tell SourceX what de-identified data you need.
Sources
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- National Law Review, "Finding the Delta: Understanding the Differences in State Deidentification Standards". https://www.natlawreview.com/article/finding-delta-understanding-differences-state-deidentification-standards
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- VPH Institute, "The European Data Protection Board (EDPB) adopts pseudonymisation guidelines" (2025). https://www.vph-institute.org/news/the-european-data-protection-board-edpb-adopts-pseudonymisation-guidelines.html
- Bird & Bird, "EU: The SRB decision - a new era for personal data and data processing agreements" (2025). https://www.twobirds.com/en/insights/2025/eu-the-srb-decision-a-new-era-for-personal-data-and-data-processing-agreements
- CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
- Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/
- U.S. Government Publishing Office / U.S. Department of Education, "34 CFR 99.31 - Under what conditions is prior consent not required to disclose information?" (2018). https://www.govinfo.gov/content/pkg/CFR-2018-title34-vol1/pdf/CFR-2018-title34-vol1-sec99-31.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.