Privacy and preparation
Hashed emails and IDs: why hashing is not anonymization
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Hashed data is not anonymous. A hashed email or customer ID still links every record about the same person, and anyone holding a list of emails can hash it and match the results. Before licensing records, replace identifiers with random tokens whose mapping is destroyed, or remove them, and treat salted or keyed hashes as pseudonymization.
Key takeaways
- Hashing the same email always produces the same value, so hashed identifiers keep records linkable.
- Emails and phone numbers are guessable, so a hashed list can be matched back by hashing candidate values.
- Salting and keyed hashing slow outsiders down, but the key holder can still link records, which is pseudonymization.
- Random tokens with a destroyed mapping keep sequences intact without leaving a path back to the person.
- Under GDPR, pseudonymized data generally remains personal data.
Why hashing looks like anonymization but is not#
Hashing looks like anonymization because the output is unreadable: an email becomes a long string of letters and digits that cannot be decoded directly. But a hash function is deterministic, so the same email produces the same hash every time, in every system that uses the same function.
Determinism is the reason hashing is popular in marketing and analytics. Advertising platforms match customer lists by comparing hashed emails, which only works because the hash still points to the person. If a hashed value can be matched, it has not been anonymized.
For licensing, the practical question is whether someone holding the dataset could link a record to a person. A hashed identifier makes that easier, not harder, because it is a stable key across every record in the dataset and potentially across other datasets too.
Three ways hashed identifiers get linked back#
Hashed identifiers get linked back in three common ways, and none requires breaking the hash function itself.
The third route is the most common in business records. A support ticket keyed by a hashed customer ID may still contain the customer's signature, and the hash then quietly links every other ticket that person ever filed.
Hashed identifiers also travel. Marketing tools, analytics platforms and warehouse views often hold the same hashed email under different column names, so a scan for one field name misses the copies. Search by value pattern as well as by field name.
- Guessing: emails, phone numbers and customer numbers come from small, predictable spaces, so an attacker hashes a list of candidates and looks for matches.
- Joining: the same unsalted hash appears in other datasets, such as advertising audiences or partner data feeds, so records can be joined across them.
- Context: the hash sits beside a job title, a city, a company name or ticket text that mentions the person, which identifies them without touching the hash.
Salted, keyed and tokenized identifiers compared#
Salted, keyed and tokenized identifiers differ in who can link records and whether anyone can reverse the process. The table compares the common options for a dataset being prepared for licensing.
Keyed approaches are legitimate preparation tools. Deterministic encryption, offered in mainstream tooling such as Google's Sensitive Data Protection API, maps each input to a consistent output so tables can still be joined while preparation is under way. The decision that matters is whether the key survives delivery.
| Technique | Same input, same output? | Who can link back | Fit for licensing |
|---|---|---|---|
| Plain hash of an email or ID | Yes, everywhere | Anyone who can guess or obtain the input | Not suitable as a privacy measure |
| Hash with a fixed secret salt | Yes, inside your systems | Anyone holding the salt | Pseudonymization; keep the salt secret or destroy it |
| Keyed hash or deterministic encryption | Yes, for the key holder | The key holder | Pseudonymization; useful for joins during preparation |
| Random token with mapping kept | Yes, through the mapping table | Whoever holds the mapping | Pseudonymization; mapping stays with the supplier |
| Random token with mapping destroyed | Consistent inside the dataset only | No one through the token itself | Preferred when sequences matter |
| Removal or generalization | Not applicable | No one through the identifier | Preferred when identity adds nothing |
What privacy rules generally say about hashed data#
Privacy rules generally treat hashed identifiers as personal data when the data can still be linked to a person. Under GDPR, pseudonymized data, which includes keyed hashes and tokens with a retained mapping, generally remains personal data, and the same may apply to data transformed with deterministic or format-preserving encryption.
US state privacy laws tend to ask whether data can reasonably be linked to an individual, and a hash anyone can recompute from a known email is hard to describe as unlinkable. Which definitions apply, and whether a prepared dataset meets them, is assessed deal by deal with counsel.
The practical rule for documents is simple: avoid calling hashed data anonymous in contracts, privacy notices, security questionnaires or dataset cards. Describe the technique precisely instead, so a reader knows who could link the records and how.
A decision rule for identifiers before licensing#
A decision rule for identifiers starts with what the buyer actually needs from them. Most AI uses of business records need to know that two records involve the same customer or agent, not who that person is.
Check free text after tokenizing. Replacing the customer ID column does nothing for the email address typed into the third comment of a ticket, and that leftover email undoes the tokenization for every record linked to it.
Run the tokenization inside your own environment, so raw identifiers never reach anyone doing later preparation or review, and keep the tokenization code and configuration with the redaction log.
| If the dataset needs | Use | Why |
|---|---|---|
| To follow one customer or employee across records | Random tokens per person, mapping destroyed after QA | Keeps sequences without a path back |
| Repeat deliveries with the same tokens | Random tokens with a mapping held only by the supplier | Consistent across versions; still pseudonymized, so document it that way |
| No link between records | Remove the identifier | Nothing left to protect or leak |
| Location or time patterns | Generalize, such as region or month | Keeps the signal at a coarser level |
| To join with the buyer's own data | Decline | Joining on identifiers is re-identification by design |
Illustrative: an email marketing software company replaces hashed IDs#
Illustrative: a fictional email marketing software company wants to license product usage sequences showing how account admins build, test and send campaigns. Its event warehouse keys every event by a SHA-256 hash of the user's email, which the data team had long described internally as anonymized.
A preparation review points out that any customer's admin list could be hashed and matched with little effort. The company replaces each hash with a random token per user and per account, keeps the mapping only until QA passes, then destroys it. It also strips campaign subject lines that contained recipients' names.
The redaction log records the token method and the date the mapping was destroyed. The usage sequences remain intact for analysis, and no value in the dataset can be recomputed from an email address.
How SourceX treats hashed identifiers#
SourceX treats hashed identifiers as personal data during the Preparation step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Identifiers are replaced with random tokens or removed, and the choice is recorded in the privacy record of the SourceX Evidence Packet.
The buyer receives the method, never the keys or mapping tables. Where repeat deliveries need consistent tokens, the mapping stays with the supplier, and the packet describes that arrangement as pseudonymization rather than anonymization.
Frequently asked questions
Does adding a salt make hashed emails anonymous?
No. A salt stops outsiders from matching against precomputed tables, but anyone holding the salt can still hash candidate emails and match them, and records stay linked to each other. A salted hash is pseudonymization, not anonymization.
Is SHA-256 safer than MD5 for this purpose?
For anonymity, the difference barely matters. Both are deterministic, so guessable inputs such as emails can be hashed and matched either way. The choice of algorithm matters for security uses such as integrity checks, not for whether hashed identifiers point to people.
Can we keep hashed IDs if the buyer promises not to re-identify?
A contractual ban on re-identification is a useful control, and many licenses include one, but it does not change what the data is. Replace hashes with random tokens anyway, and keep the contract clause as a second layer of protection.
Are hashed IP addresses safe to include?
Generally not. The range of possible IP addresses is small enough that every value can be hashed and compared. Remove IP addresses, or generalize them to a coarse network or region if location matters for the use case.
Do internal customer numbers need the same treatment?
Yes. Internal IDs are not hashed, but they behave the same way: anyone with access to your systems, or a document that cites the number, can link them. Replace them with tokens too, especially where invoices, emails or tickets mention the number in free text.
Sources
- Google's Sensitive Data Protection API supports de-identification transforms including deterministic encryption (CryptoDeterministicConfig) and format-preserving encryption. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.