Privacy and preparation
Pseudonymization vs anonymization: which does a data license need?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
A data license should usually ship anonymized, or de-identified, records rather than pseudonymized ones. Pseudonymization swaps identifiers for codes that whoever holds a key or lookup table can reverse, so GDPR treats the result as personal data and US state laws generally do too. Anonymization removes the link to a person for good. Keep every key out of the deal.
Key takeaways
- Pseudonymization is reversible by whoever holds the key, so pseudonymized records remain personal data.
- Anonymization, called de-identification in most US laws, aims to make re-identification not reasonably possible for anyone.
- Consistent random tokens with no stored mapping give buyers linked records without creating a key.
- Hashing an email, phone number or employee ID is pseudonymization at best, and the result can often be reversed.
- Recurring deliveries that need the same token for the same person across shipments require a kept key, which changes the legal analysis.
What is the difference between pseudonymization and anonymization?#
Pseudonymization swaps direct identifiers such as names, emails and employee numbers for codes, while keeping a way back: a lookup table, an encryption key or a hash anyone can recompute. Anonymization removes identifiers and reduces indirect clues until the records can no longer reasonably be tied to a person by anyone, including the company that created them.
The difference is not how the output looks. A record that says [EMPLOYEE_17] might be pseudonymized or anonymized depending on whether a mapping from that token to a real employee still exists somewhere.
European law draws the line explicitly. GDPR Article 4(5) defines pseudonymisation as processing data so it can no longer be attributed to a specific person without additional information that is kept separately and protected. Recital 26 says pseudonymized data that could be attributed to a person with that information should be treated as personal data, while truly anonymous information falls outside the regulation altogether. Case law on whether pseudonymized data is personal data for a recipient with no realistic access to the key is still developing, so do not build a deal on that argument without counsel.
US privacy laws mostly use the word de-identified rather than anonymized, and attach conditions to it. California's definition, for example, requires reasonable measures so the data cannot be associated with a consumer or household, a public commitment to keep it de-identified and not try to re-identify it, and contract terms obliging recipients to comply. For a data license, the US term is usually the one counsel works with.
Pseudonymization vs anonymization compared#
Pseudonymization and anonymization differ on reversibility, legal status and how much risk the recipient carries. The table compares them on the questions that come up in a licensing review.
| Question | Pseudonymization | Anonymization or de-identification |
|---|---|---|
| Can someone reverse it? | Yes, whoever holds the key or table | Not reasonably, by design |
| Status under GDPR | Personal data for anyone who can attribute it using the additional information | Outside GDPR if truly anonymous |
| Status under US state laws | Personal data; some laws relax a few duties for pseudonymous data | Generally outside personal data when the statutory conditions are met |
| Linked records for the buyer | Yes, through stable codes | Yes, if tokens are consistent within the package |
| Harm if the dataset leaks | High if the key leaks too | Lower, though indirect clues still matter |
| Typical techniques | Lookup tables, keyed hashes, reversible encryption | Removal, consistent tokens with no mapping, generalization |
| Fit for a data license | Treat as personal data, with all that entails | The usual target for licensed operational records |
Consistent tokens without a key: the middle path#
Consistent tokens without a stored key give buyers most of what pseudonymization offers while avoiding its main legal problem. Each person in a package gets a random placeholder such as [CUSTOMER_4] that stays the same across their tickets and emails, and the mapping used to assign it is deleted once preparation and QA finish.
Linkage is what makes operational records useful for AI training: the same customer reporting a fault, the same technician returning, the same engineer reviewing several pull requests. Consistent tokens preserve that story without preserving identity.
Linkage also accumulates clues. A long thread about one customer can mention their town, their equipment and their job title, and together those can point to a person. The Article 29 Working Party's 2014 opinion on anonymisation techniques framed this as three risks: singling out a person, linking records about them and inferring facts about them. Consistent tokens deliberately allow linking within a package, so review the longest token histories first and generalize or remove indirect details where they pile up. For structured exports, Google's Sensitive Data Protection API offers re-identification risk metrics such as k-anonymity and l-diversity.
Pseudonymization methods and their weak points#
Pseudonymization methods vary in how easily they reverse, and most are weaker than they look. The common ones below show up in exports that teams believe are already anonymized.
A quick test sorts most fields. Ask who could turn the code back into a person using information they already hold or could easily obtain. If the answer includes anyone, even your own HR or finance team, the field is pseudonymized and needs a different treatment before it ships.
- Lookup tables: a spreadsheet mapping names to codes. Anyone with the spreadsheet can reverse every record, and spreadsheets travel.
- Plain hashes: hashing an email or phone number produces the same output every time, so anyone can hash a list of likely values and match them.
- Keyed hashes and deterministic encryption: stronger, because reversal needs the key, but the data stays pseudonymized while the key exists, and whoever holds it can reverse every record.
- Internal IDs: badge numbers, employee numbers and customer account numbers look anonymous but map straight back through HR, ERP or CRM systems.
- Date shifting: moving dates by a random offset hides exact timing but keeps intervals, and the offset table is itself a key.
When a license might accept pseudonymized data#
A license might accept pseudonymized data when the buyer needs the same person to carry the same code across several deliveries, such as quarterly refreshes of a support archive. Matching codes across shipments requires keeping a key, so the records stay personal data from your side.
That changes the deal. Counsel then treats the license as a disclosure of personal data, which can bring in sale definitions, assessments and notice obligations under state laws, and data processing terms under GDPR where it applies. Some sellers accept that; many prefer fresh tokens per delivery and accept the loss of cross-delivery linkage.
If a key is kept, it should never travel with the data, never be shared with the buyer, and sit under named access in your own systems.
Illustrative: a manufacturer's quality records#
Illustrative: a fictional precision machining company plans to license nonconformance reports and corrective action records from its QMS. Each report names the operator by badge number, the inspector by name and the supplier contact by email. Customer-owned drawings are excluded from the start.
The first plan hashes badge numbers and keeps the HR mapping. Counsel points out that the hash is reversible by anyone who can read the badge list, so the package would remain personal data. The revised plan assigns random tokens per person for this package only, removes names and emails from free text, generalizes shift dates to the week, and deletes the mapping after QA.
The outcome: buyers can still follow how one inspector's rejections led to a corrective action, and no one at the company or the buyer can turn a token back into a person.
How SourceX handles keys and tokens#
Token design is settled in Preparation, the third stage of the SourceX five-step transaction. Packages carry consistent placeholders where linkage adds value, and any working mapping stays in the supplier's environment and out of the delivery and the license. Where a buyer asks for the same codes across several deliveries, that request goes back to the supplier and counsel as a rights question rather than being settled quietly during preparation.
The method, the token scope and the fate of any mapping are recorded in the privacy record of the SourceX Evidence Packet. The license itself prohibits re-identification attempts and linking with outside data to identify people.
Frequently asked questions
Is hashed data anonymized?
Usually not. A plain hash of an email or phone number can be matched by hashing a list of likely values, and a keyed hash stays reversible while the key exists. Hashing is a pseudonymization technique. For licensed records, a random token with no stored mapping is usually the better choice.
Can we keep the mapping table after licensing, just in case?
You can, but the data then remains pseudonymized from your side, and counsel may treat the license as a disclosure of personal data. If you keep it, store it apart from any copy of the dataset, restrict access by name, and decide in advance when it will be destroyed.
Does anonymization make operational records useless for AI?
Rarely. The value of support tickets, quality records or code reviews sits in the problem, the reasoning and the outcome, not in who the customer was. Consistent tokens keep the thread readable. Identity details add little to training and a lot to risk.
What should the license say about de-identification?
The license should name the de-identification method or standard, prohibit re-identification and linking with other data to identify people, require prompt notice if identifiable content is found, and require its deletion. These terms also help satisfy the contract condition in many state definitions of de-identified data.
Do the data standards for AI datasets record which method was used?
Some do. The Data and Trust Alliance's Data Provenance Standards include a code list of privacy-enhancing tools, with pseudonymization, anonymization, tokenization and k-anonymity among the options, so a dataset's metadata can state the method applied.
Sources
- GDPR Article 4(5) defines pseudonymisation as processing personal data so that it can no longer be attributed to a specific data subject without the use of additional information, provided that information is kept separately and protected by technical and organisational measures. Source
- GDPR Recital 26 states that personal data which have undergone pseudonymisation, which could be attributed to a natural person by the use of additional information, should be considered information on an identifiable natural person, and that data protection principles do not apply to anonymous information. Source
- The Article 29 Data Protection Working Party's Opinion 05/2014 on Anonymisation Techniques (WP216), adopted in April 2014, assesses anonymisation techniques against three risks: singling out, linkability and inference. Source
- Post-CPRA, Cal. Civ. Code 1798.140(m) treats information as deidentified only if the business takes reasonable measures so it cannot be associated with a consumer or household, publicly commits to keep it deidentified and not reidentify it, and contractually obligates recipients to comply. Source
- Google's Sensitive Data Protection API offers re-identification risk-analysis metrics including k-anonymity and l-diversity. Source
- The Data Provenance Standards' Privacy Enhancing Tools code list includes data anonymization, masking, redaction, k-anonymity, l-diversity, pseudonymization and tokenization, among others. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.