Skip to content

Privacy, de-identification and sensitive data

Pseudonymisation techniques for training data: keyed hashing, tokens and consistent surrogates

Quick answer

Pseudonymisation replaces direct identifiers with values that cannot be linked back to a person without separately held additional information, such as a secret key or a token vault. For training data, the workable pattern is a keyed hash (HMAC) or vault token for join keys, plus consistent, realistic surrogates in free text so the model still sees one customer across a whole conversation. Hashed data is not anonymous: under the GDPR it stays personal data wherever someone can reasonably re-link it [2][4].

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Is hashed data anonymous?

No: a plain hash of an identifier is pseudonymised data at best, and often trivially reversible. Email addresses, US phone numbers, account numbers and SSNs come from small or enumerable spaces, so an attacker can hash every candidate value with SHA-256 and match the output, a dictionary attack that takes minutes on commodity hardware. Even a well-protected hash remains personal data for anyone holding the means to re-link it; Recital 26 of the GDPR tests anonymity against "all the means reasonably likely to be used" [3], and the EDPB's pseudonymisation guidelines (adopted for consultation January 2025; no final version as of September 2026) say pseudonymised data remains personal data [1][2].

The legal question has a recipient-side twist as of October 2026. In EDPS v SRB (C-413/23 P, 4 September 2025) the Court of Justice held that pseudonymised data is personal data for a recipient if they have reasonable means to re-identify, including the possibility of obtaining the additional information from the sender [5][6]. That outcome depends on how strong the pseudonymisation is and on whether the recipient can obtain the additional information, so weak hashing forfeits it. The Digital Omnibus proposal to narrow the personal-data definition is not law as of October 2026 [11]; for the classification itself, see de-identified, anonymised, pseudonymised and aggregated data compared.

Which pseudonymisation technique fits which field?

The right technique depends on whether the field must join across tables, stay readable in text, or preserve a format for downstream parsers. NIST's de-identification survey treats pseudonymization as one tool among several, with re-identification risk driven mainly by what can be linked [8]. Use the table below to assign a method per field rather than one method per dataset.

Illustrative example: invented to show structure; it does not describe an available dataset.

TechniqueHow it worksKeeps joins?Reversible by recipient?Typical use in training dataMain failure mode
Unkeyed hash (SHA-256 of value)Deterministic digestYesYes, by dictionary attack on emails, phones, IDsAvoid for identifiersRainbow tables and enumeration
Salted hash, random salt per recordDigest with a unique saltNoNo, but salt is often shipped alongsideOne-off deduplication onlyBreaks entity consistency
Keyed hash (HMAC-SHA-256 with secret key)Digest with a key the supplier keepsYes, within one key scopeNot without the keycustomer_id, account_no, ticket requesterKey leakage, unnormalized inputs
Vault tokenizationRandom token, mapping table held by supplierYesNot without the vaultAccount and case IDs that need later lookup by supplierVault exported with the data
Format-preserving encryption (for example NIST FF1)Keyed encryption that keeps length and alphabetYesNot without the keyFields that downstream schemas validateSmall domains make guessing practical
Consistent realistic surrogatesMapped fake names, emails, phones generated from a keyed seedYes, within the mapping scopeNot without the mappingNames and contact details inside free textMissed mentions, gender or locale mismatch
Generalization or suppressionCoarsen or removeNon/aDates of birth, ZIP codes, rare job titlesOver-coarsening hurts utility

How to implement keyed hash pseudonymisation correctly

A keyed hash is only as strong as its key custody and its input normalization. Compute HMAC-SHA-256 over a normalized value, for example lowercased and trimmed emails, E.164 phone numbers, and account numbers stripped of separators, so "Jane.Doe@Example.com " and "jane.doe@example.com" map to the same pseudonym. Add domain separation by prefixing the input with the field type ("email|", "acct|") so the same string in two fields does not produce the same output and leak a cross-field link.

Truncation and encoding choices matter for training. Truncating to 64 bits keeps collision odds low at the scale of tens of millions of entities, while very short tokens collide and silently merge two customers. Render pseudonyms with a type prefix such as CUST_7f3a9c21 so the model learns a stable placeholder pattern, and so tokens never look like real account numbers that might trip downstream PII scanners.

Decide the key scope deliberately, because it sets who can link what. One key per delivery prevents linkage across licenses but breaks longitudinal joins; one key per buyer relationship keeps joins across refreshes, which is what stable record IDs across dataset deliveries need, but also means a single key compromise exposes every delivery. Never use one global key across buyers: identical tokens in two buyers' datasets would let them be joined.

Consistent pseudonyms across documents and conversations

Entity consistency is the reason to pseudonymise rather than redact for SFT and agent training. If "Jane Doe" becomes [NAME] in turn 1 and [NAME] again for the agent in turn 3, the model cannot learn who said what, which customer the refund belongs to, or how to carry a case across a handoff. Consistent surrogates, where every mention of one person maps to the same fake name throughout a conversation or case, preserve coreference.

Choose the consistency scope per use case. Conversation scope is safest and fits single-session support chats; case scope fits multi-email threads and ticket histories; account scope fits longitudinal CRM or billing histories; global scope is rarely justified and maximizes linkage risk. Derive surrogates from a keyed seed (for example HMAC of the normalized identity, used to index a name list) so the same identity yields the same surrogate within scope without storing a lookup table.

Surrogates fail in predictable ways that you should test for:

  • Partial mentions: "Jane", "Ms. Doe", "JD" and the local part of jane.doe@ must all map to the surrogate family, or the original leaks beside the replacement.
  • Structured and unstructured drift: the requester field says Maria Lopez while the body still says "Hi Jane", which teaches the model a false link.
  • Attribute mismatch: swapping a name across gender or locale can break pronoun agreement and language-specific forms.
  • Surrogate collision with real people: a surrogate list containing real employees or customers recreates personal data; generate from vetted lists.

Detection is the weak link, because surrogates are applied only where an identifier was found. Open-source tools such as Microsoft Presidio say plainly that ML-based detection gives no guarantee of finding all sensitive information [10]; measure recall as described in PII redaction for LLM training data. Indirect identifiers such as rare job titles or small towns survive pseudonymisation entirely; see indirect identifiers in business text.

Pseudonymisation secret key management and separation of additional information

The key, salt, vault or mapping table is the "additional information" that defines pseudonymisation, so it must stay with the supplier and be kept apart under technical and organizational controls. The GDPR definition requires that this information be "kept separately" and protected, and the EDPB guidelines connect pseudonymisation to data minimization, data protection by design and security of processing under Articles 5, 25 and 32 [2][3]. The ICO likewise describes the additional information as held separately and treats pseudonymised data as personal data (under review after the DUAA 2025; draft guidance published August 2026) [4].

For a buyer, the practical rule is simple: you should never receive the key, the salt, the token vault or the surrogate mapping. Ask the supplier to keep keys in a KMS or HSM-backed service, restrict use to the pseudonymisation job's service account, log every key access, and document rotation. Rotation is a trade-off: rotating the key re-pseudonymises everything, which breaks joins with earlier deliveries unless the supplier re-issues history.

US health data has its own version of this rule. Under HIPAA, a covered entity may assign a re-identification code to de-identified data only if the code is not derived from information about the individual and the mechanism for re-identification is not disclosed [9]. A keyed hash of an MRN is derived from the MRN, so for health records the stronger pattern is a random vault token, with the de-identification route chosen as Safe Harbor or Expert Determination.

Intake checklist for pseudonymised training records

Before loading a pseudonymised delivery, verify the method, the scope and the residual risk rather than trusting the label. The checklist below is what a data engineer can run at intake; pair it with the de-identification evidence package checklist for documents to request.

Illustrative example: invented to show structure; it does not describe an available dataset.

pseudonymisation_manifest:
  delivery_id: dlv-0003
  key_scope: buyer_relationship        # delivery | buyer_relationship
  key_custody: supplier_kms            # buyer must never hold
  fields:
    customer_email:   {method: hmac_sha256, normalize: lower_trim, prefix: "EMAIL_", bits: 64}
    account_number:   {method: vault_token, prefix: "ACCT_"}
    agent_name:       {method: surrogate, scope: conversation, list: vetted_names_v2}
    customer_name:    {method: surrogate, scope: case, aliases: [first, last, initials, email_local]}
    date_of_birth:    {method: generalize, to: year}
    free_text_body:   {method: detect_then_surrogate, detector: ner_plus_regex}
  qa:
    sample_size_reviewed: 500
    detector_recall_on_sample: 0.97
    residual_findings_logged: true
  • Hash test: hash a list of plausible emails and phones with plain SHA-256 and MD5 and confirm none match delivered tokens.
  • Consistency test: sample 50 multi-turn conversations and confirm each person keeps one surrogate across turns, quotes and signatures.
  • Cross-field test: confirm the same raw value in two fields yields different tokens when domain separation is specified.
  • Leakage scan: run your own detector over the delivered text and log any real-looking names, emails or account numbers.
  • Linkage review: list quasi-identifiers that remain (job title, location, dates) and decide whether to generalize them.
  • Memorization planning: tokens and surrogates can be memorized like any string; see training-data extraction and memorization risk.

What pseudonymisation means for the trained model

Pseudonymising inputs lowers risk but does not by itself make a model anonymous. EDPB Opinion 28/2024 says whether a model trained on personal data is anonymous must be assessed case by case, including the likelihood of extracting personal data from it [7]. Strong pseudonymisation with supplier-held keys, consistent surrogates and measured residual risk is the evidence that makes that assessment defensible.

SourceX works from the same premise. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked, while no method is perfect. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and diligence materials on source, rights, preparation and allowed use are prepared per dataset. Teams that want to discuss a specific record type can describe the records they need. Background on terminology is in the pseudonymization glossary entry and the privacy guide for de-identified training data, within the wider AI data hub.

Request pseudonymised operational records for training

SourceX sources operational datasets such as support and sales histories, engineering records and finance and legal workflows from US companies on request, and every release is approved by the supplying company. Personal details are removed or replaced before delivery, and data moves through private, access-controlled workflows only after an executed agreement. Describe the records and the pseudonymisation you need at SourceX for buyers.

Frequently asked questions

Can a buyer ask for the pseudonymisation key to fix bad joins?

Holding the key would let the buyer re-identify people and would weaken any argument that the data is non-personal for the buyer after EDPS v SRB [5][6]. Ask the supplier to rerun the join on their side and re-deliver instead.

Is format-preserving encryption better than HMAC for training data?

Only when a downstream schema validates length or character set. For small domains such as four-digit codes, any deterministic keyed method can be guessed by trying all values, so generalize or suppress those fields instead.

Do surrogates need to look realistic?

Realistic surrogates keep text fluent and help models learn natural dialogue, while bracketed placeholders are easier to audit. Many teams use realistic names in free text and prefixed tokens in structured fields, documenting both in the manifest.

Sources

  1. European Data Protection Board, "Guidelines 01/2025 on Pseudonymisation" (2025). https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2025/guidelines-012025-pseudonymisation_en?page=4
  2. VPH Institute, "The European Data Protection Board (EDPB) adopts pseudonymisation guidelines" (2025). https://www.vph-institute.org/news/the-european-data-protection-board-edpb-adopts-pseudonymisation-guidelines.html
  3. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  4. Information Commissioner's Office, "Pseudonymisation" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
  5. EUR-Lex (Official Journal of the European Union), "Judgment in Case C-413/23 P, EDPS v SRB" (2025). https://eur-lex.europa.eu/eli/C/2025/5551/oj/eng
  6. Bird & Bird, "EU: The SRB decision: A new era for personal data and data processing agreements?" (2025). https://www.twobirds.com/en/insights/2025/eu-the-srb-decision-a-new-era-for-personal-data-and-data-processing-agreements
  7. European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
  8. National Institute of Standards and Technology, "NISTIR 8053: De-Identification of Personal Information" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
  9. Electronic Code of Federal Regulations (HHS), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  10. Microsoft (microsoft/presidio project), via pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  11. Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data