Privacy and preparation
Consistent pseudonyms: keeping one person one placeholder across a dataset
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Consistent pseudonymization replaces each person with the same placeholder everywhere they appear, so a support thread, an email chain and a Slack discussion still show who said what without revealing who they are. Resolve name variants to one entity first, generate placeholders with a secret key kept apart from the data, then destroy or escrow that key.
Key takeaways
- Random tags per mention break conversation structure; consistent placeholders keep it.
- Entity resolution comes first: nicknames, email addresses, handles and misspellings must map to one person.
- Generate placeholders with a keyed hash or a mapping table that never travels with the dataset.
- Use a different key for each recipient so two licensees cannot join their copies.
- Consistent placeholders are not anonymization; the surrounding text still needs review.
Why does consistency matter in conversation data?#
Consistency matters in conversation data because the value of a support thread or project discussion lies in who responded to whom, who escalated and who resolved the problem. If every mention of a person gets a fresh random tag, a thread with a customer, two agents and an engineer turns into a crowd of strangers, and a model cannot learn the handoffs.
Consistency also keeps records joinable. When the same customer placeholder appears in a Zendesk ticket, the linked Jira issue and the Slack channel where engineers discussed the fix, a buyer can follow the path from complaint to code change. That linkage often separates a useful operational dataset from a pile of disconnected messages.
Four ways to replace a name, compared#
The replacement style decides whether the conversation survives de-identification. The comparison below uses one sentence from a support thread to show the trade-offs.
| Approach | Example output | Thread still readable? | Main risk |
|---|---|---|---|
| Blanket removal | [REDACTED] asked [REDACTED] to escalate | No | Meaning lost; roles unclear |
| Random tag per mention | PERSON_81 asked PERSON_02, then PERSON_55 replied as the same person | No | Looks consistent but is not |
| Consistent role-typed placeholder | Customer_0412 asked Agent_17 to escalate | Yes | Patterns under one tag can profile a person |
| Consistent synthetic name | Dana Whitfield asked Sam Ortiz to escalate | Yes, and reads naturally | Fake names can collide with real people |
The five-step method for consistent placeholders#
The five-step method for consistent placeholders works the same for tickets, email, chat and project tools. What changes between systems is where identifiers hide, not the logic.
Treat the steps as one run across every source in scope. Pseudonymizing Zendesk this month and Slack next quarter with a different setup is an easy way to lose consistency.
- Inventory: list every identifier type that refers to a person across sources, including names, email addresses, phone numbers, usernames, chat handles, account IDs and signature blocks.
- Resolve: group the variants that belong to one person into a single entity with a stable internal ID.
- Generate: derive each placeholder from that internal ID with a keyed hash (an HMAC) or a lookup table. The secret key, sometimes called a global salt or pepper, stays the same for the whole run and is stored apart from the data; a random salt per record would break consistency.
- Apply: replace every variant in one pass across all sources, including subject lines, headers, metadata fields, quoted replies and free text.
- Verify and retire: check that threads and joins survive and that no two people share a placeholder, then destroy the key and mapping or place them in escrow under the supplier's control.
Entity resolution: the step pipelines often skip#
Entity resolution decides whether Robert Kim, Bob, rkim at the company domain, the handle bobk and R. Kim are one person. A pipeline that hashes each string separately gives each variant its own placeholder, which breaks the thread as surely as random tags.
Start from structured sources: the user directory, the help desk agent list and the CRM contact table, which already tie names to emails and handles. Then handle free text: first-name-only mentions, nicknames, misspellings and pronouns that refer back to a named person, which is the coreference problem. Where two people share a first name, such as two agents named Maria, use the surrounding record, such as the assignee or message author, to choose, and fall back to a generic tag when the text cannot decide.
Some cases need a rule written in advance: a person who is both an employee and a customer, someone who changed surnames, and shared mailboxes that several people answer from. Decide each rule once and apply it everywhere.
Keyed hashes, lookup tables and a key per recipient#
A keyed hash or a lookup table is what keeps placeholders stable without exposing the original. An unkeyed hash of an email address is not enough, because anyone can hash a list of likely addresses and compare. The same logic applies to the key itself: whoever holds it can recompute placeholders for a list of real people, so it is a re-identification key and needs the same protection as a mapping table.
Cloud tooling supports the same idea. Google's Sensitive Data Protection API includes deterministic encryption among its de-identification transforms, and its API definition recommends it over format-preserving encryption wherever the original character set need not be preserved, because format-preserving encryption carries significant latency costs.
Use a different key for each recipient. The same customer then becomes Customer_0412 in one licensee's copy and Customer_7730 in another, so two licensees cannot join their copies. The per-recipient key also gives each copy a fingerprint, which is the basis of the recipient markers covered in the guide to canary records.
| Method | How it works | Strength | Watch for |
|---|---|---|---|
| Keyed hash | Placeholder derived from the entity ID plus a secret key | Same input, same output, no table to store | Shortening the hash into a readable tag can cause collisions; losing the key ends consistency |
| Lookup table | Each entity assigned a placeholder in a stored mapping | Easy to audit and to assign readable tags | The table is a re-identification key and must be secured |
| Deterministic encryption | Reversible encryption with a fixed key | Allows authorized re-linking | Output stays closer to personal data while the key exists |
Illustrative: a support thread that stays readable#
Illustrative: a fictional vertical software company prepared support tickets from Zendesk, linked issues from Jira and the Slack threads where engineers discussed fixes. Its first pass hashed each name string, so one customer administrator appeared under four placeholders: full name, first name, email address and Slack handle.
The CTO's team rebuilt the run around entity resolution from the agent list, the CRM contact table and the Slack user directory. They generated role-typed placeholders with a keyed hash and replaced variants in subject lines, signatures and quoted replies as well as message bodies. A reviewer then read complete threads end to end. Each person had one placeholder across all three systems, the escalation path was clear, and the key went into escrow with the security lead.
What consistent placeholders do not fix#
Consistent placeholders do not make a dataset anonymous. GDPR Recital 26 says pseudonymised data that could be attributed to a person with additional information should be considered information on an identifiable person. California's definition of deidentified data asks for more than technique: the business must also publicly commit not to reidentify the data and contractually bind recipients to the same terms. Whether a pseudonymized dataset may qualify as deidentified under a given law is a question for counsel.
Free text can still identify someone through a job title, a location, an unusual event or a signature, and a consistent tag gathers all of one person's activity in one place. Pair placeholders with the rest of the de-identification work: generalize quasi-identifiers, strip signatures and attachments with personal details, and review samples by hand.
Where long-term linkage is not needed, scope placeholders to a project or period so they cannot build a multi-year profile. Employee performance data is the clearest case: a stable tag on a warehouse picker turns months of scan events into a profile of one worker, so rotate it or drop the worker key altogether.
Who holds the key in a SourceX delivery#
The supplier holds the key and the mapping in a SourceX delivery; neither travels with the data. Placeholder rules, entity resolution decisions and the per-recipient key choice are settled in the Preparation step of the SourceX five-step transaction, and the privacy record in the SourceX Evidence Packet states whether the mapping was destroyed or escrowed. The supplier approves that record before Delivery.
Declaring the method also helps the buyer's own records. The Data & Trust Alliance's Data Provenance Standards include a privacy-enhancing tools code list that names pseudonymization, tokenization and masking, so a dataset card can state which technique was applied and to which fields, without revealing the key.
Frequently asked questions
Should placeholders look like names or like tags?
Tags such as Customer_0412 or Agent_17 are safer and make clear that text was changed. Synthetic names read more naturally, and some buyers prefer them for language work, but they can collide with real people. If you use synthetic names, draw them from a list checked against your own customer and employee tables.
Do we keep the mapping for future deliveries?
Only with a reason, such as adding later records under the same placeholders for the same recipient. In that case escrow the key under the supplier's control, with restricted access and a logged process. If no later delivery is planned, destroy it and record the destruction.
How should pronouns and indirect references be handled?
Pronouns rarely identify anyone, so most pipelines leave them. The risk is indirect references such as the owner's wife or our only electrician in Tulsa, which identify through context. Those need human review or rules that generalize roles and places.
Can a buyer ask us to re-identify a record later?
A license can allow narrow re-linking for agreed reasons, but every re-linking path keeps the data closer to personal data. Suppliers often decline it and keep any escrowed key for their own use, such as finding a person's placeholder to honor a deletion request.
Does the same method work for company names and account numbers?
Yes. Customer companies, vendors and account numbers can follow the same rules with their own placeholder types. Keep person and company placeholders visibly separate so a reader can tell a contact from an account.
Sources
- Google's Sensitive Data Protection API supports de-identification transforms including format-preserving encryption and deterministic encryption, and recommends deterministic encryption where the input alphabet need not be preserved because FPE incurs significant latency costs. Source
- The Data Provenance Standards' Privacy Enhancing Tools code list includes pseudonymization, tokenization, masking and redaction among other techniques. Source
- GDPR Recital 26 states that personal data which have undergone pseudonymisation, which could be attributed to a natural person by the use of additional information, should be considered information on an identifiable natural person. Source
- Under Cal. Civ. Code 1798.140(m), deidentified information requires reasonable measures against association with a consumer, a public commitment not to reidentify it, and contractual obligations on recipients to comply. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.