Skip to content

Definitions and comparisons

Anonymized vs de-identified vs pseudonymized data: what's the difference?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Pseudonymized data swaps identifiers for codes that can be linked back, so GDPR generally still treats it as personal data. De-identified data, a US legal term used in laws such as the CCPA, has identifiers removed plus safeguards against re-identification. Anonymized data cannot reasonably be linked to anyone. Licenses should name the standard met and the method used.

Key takeaways

  • The three terms describe distance from a person: pseudonymized data is linkable, de-identified data is unlinked with safeguards, and anonymized data is unlinkable.
  • Pseudonymized data generally remains personal data under GDPR for anyone able to link it back, so most obligations that apply to names and email addresses still apply.
  • De-identification in US privacy law typically pairs technical removal with commitments not to re-identify and contract terms for recipients.
  • Free text is where most identifiers survive, so automated scanning needs sampling and human review.
  • A license should name the standard the delivered records meet, not simply call the data anonymous.

What does each term mean?#

Pseudonymized, de-identified and anonymized describe three different distances between a record and the person it came from: linkable with a key, unlinked with safeguards, and unlinkable. The words are often used loosely in contracts and sales material, which is how companies end up promising more than their method delivers.

The practical test is who could put a name back on a record. If you can, the data is pseudonymized. If nobody reasonably could and there are rules against trying, it is de-identified. If nobody could even with effort and outside data, it may be anonymous.

  • Pseudonymized: direct identifiers such as names and emails are replaced with codes or tokens, and someone holds a key or table that maps them back. GDPR Article 4(5) defines it as data that can no longer be attributed to a person without additional information kept separately and protected.
  • De-identified: identifiers are removed or transformed so the data cannot reasonably be linked to a person, and the holder adds safeguards. Under the CCPA, those safeguards include reasonable measures against re-association, a public commitment not to re-identify and contract terms binding every recipient.
  • Anonymized: the data cannot be linked to a person by anyone using means reasonably likely to be used, the test in GDPR Recital 26. European regulators assess it against three risks: singling out, linkability and inference. It is the hardest standard to meet and to prove.

Anonymized vs de-identified vs pseudonymized compared#

The three categories differ on reversibility, legal status and what a license has to cover. The table gives the general position; the exact standards depend on the law and on the facts of each dataset, and counsel should confirm which one your records meet.

Anonymized vs de-identified vs pseudonymized compared
QuestionPseudonymizedDe-identifiedAnonymized
Reversible?Yes, with the key or mappingNot reasonably, and re-identification is prohibitedNo, by anyone, using reasonably likely means
Still personal data under GDPR?Yes, at least for anyone who holds or can obtain the keyOften yes, unless it also meets the anonymity standardNo
Still personal information under the CCPA?Generally yesGenerally no, if the law's de-identification conditions are metGenerally no
Typical techniquesTokenization, keyed hashing, encryption of identifiersRemoval, generalization and masking of free text, plus contract termsAggregation and suppression of rare values, tested against re-identification
What licensing needsTreatment as personal data: legal basis, notices and request handlingA documented method plus no-re-identification and onward-transfer termsEvidence the standard is met, since the claim is easy to overstate

Is pseudonymized data still personal data?#

Pseudonymized data is generally still personal data under GDPR, because Recital 26 treats data that could be attributed to a person using additional information as information about an identifiable person. EU courts have considered whether a recipient with no realistic means of re-identification is in the same position, so EU records need case-by-case advice. US state privacy laws generally treat data as personal information while it can still be associated with someone.

Pseudonymization is still worth doing. It reduces exposure inside your own company and in a breach, and consistent tokens keep a customer's support history connected, which buyers value. The problem is the key: while anyone keeps it, the delivered records remain linkable.

Mainstream de-identification services offer pseudonymization methods such as format-preserving encryption, deterministic encryption and date shifting; Google's Sensitive Data Protection API supports all three. These change how data looks without, on their own, changing its legal status.

Why free text decides which category you are in#

Free text decides the category because operational records hide identifiers in prose. A ticket can have an empty name field and still carry a signature block, a callback number, a street address and a description that fits only one customer.

Quasi-identifiers add to the risk. Latanya Sweeney's 2000 research estimated from 1990 Census data that 87% of Americans were likely unique on five-digit ZIP code, gender and date of birth, and Philippe Golle's 2006 re-analysis of 2000 Census data put the figure at about 63%. A job title, a small town and a date can do the same in a support ticket. Tests such as k-anonymity and l-diversity, which Google's Sensitive Data Protection API includes, measure this risk, but they suit structured fields better than paragraphs.

Automated detection helps with volume, not certainty. Presidio, an open-source detector of personal information, says as much in its own documentation: automated detection carries no guarantee of finding every sensitive detail, so other protections are needed alongside it. Budget for sampling and for a person reading the notes.

Which technique fits which operational record?#

The right technique depends on where identifiers sit in each record type and what the buyer needs to keep. The table lists common starting points; the final method is set after sampling the actual records.

Which technique fits which operational record?
Record typeWhere identifiers hideCommon treatment
Support tickets and chatSignatures, greetings, pasted logs and attachmentsRemove contact details, replace names with consistent tokens, drop attachments
CRM activity and emailContact fields, email headers and meeting notesDrop headers and contact fields, tokenize accounts, mask names in notes
Job and dispatch recordsService addresses, gate codes and technician notesGeneralize addresses to an area, remove access details, use role codes for technicians
Call transcriptsSpoken names, card and account numbers read aloudRedact transcripts and review them; leave out audio unless it is required
Engineering tickets and code reviewsAuthor handles, customer names, credentials in logsTokenize authors, remove customer references, scan for secrets

Illustrative: a plumbing company prepares its job history#

Illustrative: a fictional plumbing and drain company runs Housecall Pro and wants to license its job history: requests, technician diagnoses, estimates, completed work and callbacks. The records are valuable because they connect a symptom to a diagnosis and a fix.

Its first plan was to hash customer names and keep everything else. Counsel pointed out that hashed names are pseudonymized and that the notes still held addresses, gate codes and phone numbers. The company switched to random tokens with no stored mapping, generalized addresses to service area, removed access details, replaced technician names with role codes and had a reviewer sample the notes.

The license describes the records as de-identified, names the method, bars re-identification and linking with other data, and requires notice if an identifier turns up. Repeat-service history survives, because each household keeps the same random token across its jobs.

How should a license describe the standard?#

A license should describe the standard by naming the category the records meet and the method behind it, because loose labels are where disputes start. The table shows words that often appear in early drafts and how counsel usually tightens them.

The recipient's promises then keep the category true after delivery: no attempts to re-identify or link the records, the same terms for any permitted recipient, notice if an identifier turns up and destruction of copies when the license ends. Those clauses only work if they match the category the license names.

How should a license describe the standard?
Word in a draftWhat a reader may assumeSafer drafting
AnonymousNobody can identify anyone by means reasonably likely to be used, the GDPR testUse it only with evidence such as re-identification testing; otherwise say de-identified and name the method
Scrubbed or cleanedNothing specific; the words have no legal meaningList the fields removed and describe how free text was reviewed
De-identifiedThe legal conditions are met, which under the CCPA include a public commitment and recipient contract termsName the law or standard and add no-re-identification and onward-transfer clauses
Pseudonymized or tokenizedIdentifiers were replaced, possibly reversiblyState whether a key exists, who holds it, and that it never travels with the records
AggregatedOnly counts or summaries were deliveredState the minimum group size and how small groups were suppressed

How SourceX approaches de-identification#

De-identification is the core of Preparation, the third step of the SourceX five-step transaction, and it starts only after the Rights step has settled what may be licensed at all. Personal and confidential details come out before delivery, no re-identification key travels with the package, and the supplier approves the result.

The method, the review results and any remaining risk go into the privacy record of the SourceX Evidence Packet, so the category named in the license matches what was actually done to the records.

Frequently asked questions

Is hashing an email address enough to anonymize it?

No. A hashed email is pseudonymized: anyone holding a list of addresses can hash them the same way and match the results, and the hash still links every record about that person. Use random tokens with no stored mapping, or remove the identifier altogether.

Does HIPAA use the same terms?

HIPAA has its own de-identification standard for protected health information, with two methods: Expert Determination, in which a qualified expert finds the re-identification risk very small, and Safe Harbor, which removes 18 specified identifiers with no actual knowledge that the rest could identify someone. If your records include health information, that standard may apply alongside state privacy laws, so assess it with counsel.

Who decides whether data is de-identified enough?

The company licensing the data decides, usually with counsel and someone experienced in re-identification risk, and it should document the reasoning. For higher-risk records some companies commission an independent expert review. Expect buyers to ask to see the method.

Can anonymized data be licensed without telling customers?

Anonymized data generally falls outside privacy laws, but contracts and privacy notices can still restrict how you use customer records. Check customer agreements and your notice before licensing. Many companies also describe the practice in their notice so customers are not surprised.

Is aggregated data the same as anonymized data?

Not always. Aggregation lowers risk, but small groups can still single someone out, such as a count of one customer in a rare category. Aggregated data is anonymous only if no individual can be picked out of it.

Sources

  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source
  • Google's Sensitive Data Protection API offers re-identification risk metrics including k-anonymity and l-diversity, and de-identification transforms including format-preserving encryption, deterministic encryption and date shifting. Source
  • GDPR Recital 26 treats pseudonymised data that could be attributed to a person using additional information as information on an identifiable person, judges identifiability by all means reasonably likely to be used, such as singling out, and excludes anonymous information from data protection principles. Source
  • GDPR Article 4(5) defines pseudonymisation as processing personal data so it can no longer be attributed to a specific data subject without additional information kept separately and protected by technical and organisational measures. Source
  • Under Cal. Civ. Code 1798.140(m), information is deidentified only if the business takes reasonable measures, publicly commits to keep it deidentified and not reidentify it, and contractually obligates recipients to comply. Source
  • The Article 29 Working Party's Opinion 05/2014 on Anonymisation Techniques assesses techniques against three risks: singling out, linkability and inference. Source
  • Latanya Sweeney estimated from 1990 U.S. Census data that 87% of the U.S. population was likely unique on five-digit ZIP code, gender and date of birth. Source
  • Philippe Golle's 2006 paper found that gender, ZIP code and full date of birth uniquely identify about 63% of the U.S. population on 2000 Census data. Source
  • Under 45 CFR 164.514(b), protected health information can be de-identified by Expert Determination or by Safe Harbor, which requires removing 18 specified identifiers and having no actual knowledge that the remaining information could identify the individual. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify