Skip to content

Privacy and preparation

Hashed emails and IDs: why hashing is not anonymization

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Hashed data is not anonymous. A hashed email or customer ID still links every record about the same person, and anyone holding a list of emails can hash it and match the results. Before licensing records, replace identifiers with random tokens whose mapping is destroyed, or remove them, and treat salted or keyed hashes as pseudonymization.

Key takeaways

  • Hashing the same email always produces the same value, so hashed identifiers keep records linkable.
  • Emails and phone numbers are guessable, so a hashed list can be matched back by hashing candidate values.
  • Salting and keyed hashing slow outsiders down, but the key holder can still link records, which is pseudonymization.
  • Random tokens with a destroyed mapping keep sequences intact without leaving a path back to the person.
  • Under GDPR, pseudonymized data generally remains personal data.

Why hashing looks like anonymization but is not#

Hashing looks like anonymization because the output is unreadable: an email becomes a long string of letters and digits that cannot be decoded directly. But a hash function is deterministic, so the same email produces the same hash every time, in every system that uses the same function.

Determinism is the reason hashing is popular in marketing and analytics. Advertising platforms match customer lists by comparing hashed emails, which only works because the hash still points to the person. If a hashed value can be matched, it has not been anonymized.

For licensing, the practical question is whether someone holding the dataset could link a record to a person. A hashed identifier makes that easier, not harder, because it is a stable key across every record in the dataset and potentially across other datasets too.

Three ways hashed identifiers get linked back#

Hashed identifiers get linked back in three common ways, and none requires breaking the hash function itself.

The third route is the most common in business records. A support ticket keyed by a hashed customer ID may still contain the customer's signature, and the hash then quietly links every other ticket that person ever filed.

Hashed identifiers also travel. Marketing tools, analytics platforms and warehouse views often hold the same hashed email under different column names, so a scan for one field name misses the copies. Search by value pattern as well as by field name.

  • Guessing: emails, phone numbers and customer numbers come from small, predictable spaces, so an attacker hashes a list of candidates and looks for matches.
  • Joining: the same unsalted hash appears in other datasets, such as advertising audiences or partner data feeds, so records can be joined across them.
  • Context: the hash sits beside a job title, a city, a company name or ticket text that mentions the person, which identifies them without touching the hash.

Salted, keyed and tokenized identifiers compared#

Salted, keyed and tokenized identifiers differ in who can link records and whether anyone can reverse the process. The table compares the common options for a dataset being prepared for licensing.

Keyed approaches are legitimate preparation tools. Deterministic encryption, offered in mainstream tooling such as Google's Sensitive Data Protection API, maps each input to a consistent output so tables can still be joined while preparation is under way. The decision that matters is whether the key survives delivery.

Salted, keyed and tokenized identifiers compared
TechniqueSame input, same output?Who can link backFit for licensing
Plain hash of an email or IDYes, everywhereAnyone who can guess or obtain the inputNot suitable as a privacy measure
Hash with a fixed secret saltYes, inside your systemsAnyone holding the saltPseudonymization; keep the salt secret or destroy it
Keyed hash or deterministic encryptionYes, for the key holderThe key holderPseudonymization; useful for joins during preparation
Random token with mapping keptYes, through the mapping tableWhoever holds the mappingPseudonymization; mapping stays with the supplier
Random token with mapping destroyedConsistent inside the dataset onlyNo one through the token itselfPreferred when sequences matter
Removal or generalizationNot applicableNo one through the identifierPreferred when identity adds nothing

What privacy rules generally say about hashed data#

Privacy rules generally treat hashed identifiers as personal data when the data can still be linked to a person. Under GDPR, pseudonymized data, which includes keyed hashes and tokens with a retained mapping, generally remains personal data, and the same may apply to data transformed with deterministic or format-preserving encryption.

US state privacy laws tend to ask whether data can reasonably be linked to an individual, and a hash anyone can recompute from a known email is hard to describe as unlinkable. Which definitions apply, and whether a prepared dataset meets them, is assessed deal by deal with counsel.

The practical rule for documents is simple: avoid calling hashed data anonymous in contracts, privacy notices, security questionnaires or dataset cards. Describe the technique precisely instead, so a reader knows who could link the records and how.

A decision rule for identifiers before licensing#

A decision rule for identifiers starts with what the buyer actually needs from them. Most AI uses of business records need to know that two records involve the same customer or agent, not who that person is.

Check free text after tokenizing. Replacing the customer ID column does nothing for the email address typed into the third comment of a ticket, and that leftover email undoes the tokenization for every record linked to it.

Run the tokenization inside your own environment, so raw identifiers never reach anyone doing later preparation or review, and keep the tokenization code and configuration with the redaction log.

A decision rule for identifiers before licensing
If the dataset needsUseWhy
To follow one customer or employee across recordsRandom tokens per person, mapping destroyed after QAKeeps sequences without a path back
Repeat deliveries with the same tokensRandom tokens with a mapping held only by the supplierConsistent across versions; still pseudonymized, so document it that way
No link between recordsRemove the identifierNothing left to protect or leak
Location or time patternsGeneralize, such as region or monthKeeps the signal at a coarser level
To join with the buyer's own dataDeclineJoining on identifiers is re-identification by design

Illustrative: an email marketing software company replaces hashed IDs#

Illustrative: a fictional email marketing software company wants to license product usage sequences showing how account admins build, test and send campaigns. Its event warehouse keys every event by a SHA-256 hash of the user's email, which the data team had long described internally as anonymized.

A preparation review points out that any customer's admin list could be hashed and matched with little effort. The company replaces each hash with a random token per user and per account, keeps the mapping only until QA passes, then destroys it. It also strips campaign subject lines that contained recipients' names.

The redaction log records the token method and the date the mapping was destroyed. The usage sequences remain intact for analysis, and no value in the dataset can be recomputed from an email address.

How SourceX treats hashed identifiers#

SourceX treats hashed identifiers as personal data during the Preparation step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Identifiers are replaced with random tokens or removed, and the choice is recorded in the privacy record of the SourceX Evidence Packet.

The buyer receives the method, never the keys or mapping tables. Where repeat deliveries need consistent tokens, the mapping stays with the supplier, and the packet describes that arrangement as pseudonymization rather than anonymization.

Frequently asked questions

Does adding a salt make hashed emails anonymous?

No. A salt stops outsiders from matching against precomputed tables, but anyone holding the salt can still hash candidate emails and match them, and records stay linked to each other. A salted hash is pseudonymization, not anonymization.

Is SHA-256 safer than MD5 for this purpose?

For anonymity, the difference barely matters. Both are deterministic, so guessable inputs such as emails can be hashed and matched either way. The choice of algorithm matters for security uses such as integrity checks, not for whether hashed identifiers point to people.

Can we keep hashed IDs if the buyer promises not to re-identify?

A contractual ban on re-identification is a useful control, and many licenses include one, but it does not change what the data is. Replace hashes with random tokens anyway, and keep the contract clause as a second layer of protection.

Are hashed IP addresses safe to include?

Generally not. The range of possible IP addresses is small enough that every value can be hashed and compared. Remove IP addresses, or generalize them to a coarse network or region if location matters for the use case.

Do internal customer numbers need the same treatment?

Yes. Internal IDs are not hashed, but they behave the same way: anyone with access to your systems, or a document that cites the number, can link them. Replace them with tokens too, especially where invoices, emails or tickets mention the number in free text.

Sources

  • Google's Sensitive Data Protection API supports de-identification transforms including deterministic encryption (CryptoDeterministicConfig) and format-preserving encryption. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify