Privacy and preparation
Redaction tags, fake names or hashes: which replacement keeps data useful?
By SourceX Editorial · Updated
Short answer
Consistent typed tags such as [CUSTOMER_1] usually keep redacted business records the most useful, because they preserve who did what without naming anyone, and they are the safest default for licensed data. Fake names read naturally but let missed real names blend in, and plain hashes of emails or phone numbers can often be reversed. Choose a method per field.
Key takeaways
- Typed, consistent tags are the default for licensed operational records because they keep roles and conversation threads intact.
- Synthetic fake names read naturally but make a missed real name impossible to spot during review.
- Plain hashes of emails, phone numbers or IDs can be reversed by hashing lists of likely values.
- Partial masks such as a visible last four digits leak information and belong only where the tail carries meaning.
- Choose a method per field and document the placeholder vocabulary so buyers know what each tag means.
What are the options for replacing personal details?#
The options for replacing personal details range from deleting them outright to swapping in realistic substitutes. Each keeps a different amount of the record's meaning, and each fails in a different way.
Most preparation projects mix several methods. A support ticket might use typed tags for people, consistent tokens for order numbers and generalization for street addresses, all in the same paragraph.
- Generic redaction: every detail becomes the same marker, such as [REDACTED].
- Typed tags: the marker names the kind of detail, such as [NAME], [EMAIL] or [PHONE].
- Consistent typed tags: the same person or item keeps the same numbered tag, such as [CUSTOMER_1], within a defined scope.
- Partial masking: part of the value stays visible, such as the last few digits of an account number.
- Synthetic replacement: a realistic fake value stands in for the real one.
- Hashing or encryption: the value becomes a code computed from it, with or without a secret key.
- Generalization: a precise value becomes a broader one, such as a street address becoming a city.
Tags, masks, fake names and hashes compared#
Tags, masks, fake names and hashes trade readability against leak risk in different ways. The table compares them on what matters for a dataset someone will license and train on.
| Method | Readability | Leak risk | Usefulness to buyers | When to use |
|---|---|---|---|---|
| Generic redaction | Poor; who said what is lost | Low | Low; threads become hard to follow | Fields with no analytical value |
| Typed tags | Good | Low | Medium; the kind of detail survives | Contact details such as emails and phone numbers |
| Consistent typed tags | Good; roles and threads survive | Low, if the tag scope is limited | High | People, companies and case numbers in conversations |
| Partial masking | Good | Medium; visible digits narrow the search | Low | Only where the visible part carries meaning |
| Synthetic replacement | Natural | Medium; missed real names blend in | High for language tasks | When natural text matters and QA is strong |
| Plain hash | Poor | High for emails, phones and IDs | Low | Rarely, in a licensed dataset |
| Keyed hash or encryption | Poor | Depends on who holds the key | Medium; supports joins | Internal pipelines, not delivered packages |
Why consistent typed tags are usually the default#
Consistent typed tags are usually the default because operational records are conversations between roles. A ticket where [CUSTOMER_1] reports a fault, [AGENT_2] escalates and [ENGINEER_1] ships a fix still teaches the full workflow, even though nobody is named.
The scope of consistency is a real decision. Tags that stay consistent within one ticket or one case keep the story intact with little linkage risk. Tags that stay consistent across the whole dataset let a buyer follow one customer through years of records, which adds value and also lets clues accumulate into an identity. Most packages scope tags to the case or the account, and counsel signs off on anything broader.
Typed tags also make QA faster. Reviewers can skim for any capitalized name that is not inside brackets, which is the fastest way to spot a miss.
When fake names help and when they backfire#
Fake names help when the buyer needs natural-sounding text, for example to train a model that writes or reads customer messages. Brackets and underscores are not how people write, and some language tasks work better on fluent surface text.
Fake names backfire in review. Once every name in a ticket looks real, a reviewer cannot tell a missed real name from a substitute, so the leak becomes invisible. Substitutes can also collide with real people, break pronoun agreement or replace a name inconsistently across a thread.
There is a counterargument. A missed real name surrounded by substitutes is also harder for an outsider to pick out, an idea some privacy engineers call hiding in plain sight. That cover only helps once QA has already driven misses down, so it complements review rather than replacing it.
If synthetic replacement is used, draw names from a fixed, documented list, keep a findings log from the detection step, run QA on the tagged version before substitution, and state plainly in the dataset documentation that all names are synthetic.
What hashes and encrypted tokens actually protect#
Hashes and encrypted tokens protect values only as well as the key and the input space allow, and they read worst of all. A plain hash of an email address always produces the same output, so anyone with a list of likely emails can hash the list and match it; phone numbers and employee IDs fall faster because the set of possible values is small. Meanwhile a buyer reading a string such as 2cd5f1 cannot tell a driver from a carrier.
Keyed hashes and deterministic encryption resist that attack as long as the key stays secret, and Google's Sensitive Data Protection API, for one, supports deterministic and format-preserving encryption for this purpose. But while a key exists, the data is pseudonymized, not de-identified.
For a delivered package, hashes rarely beat a random consistent token. If a buyer needs to join two tables within a delivery, assign random tokens during preparation and delete the mapping afterward.
Illustrative: one freight ticket, four ways#
Illustrative: a fictional freight brokerage is preparing carrier support tickets from its TMS help desk. One ticket, written by a carrier dispatcher, shows how each method changes the same sentence. All names and numbers below are invented.
| Version | Ticket text |
|---|---|
| Original | This is Maria Okafor, dispatch at Okafor Trucking. Driver Luis Benitez is stuck at the Dayton dock, call him at 555-0142. PRO 88231 still shows no appointment. |
| Generic redaction | This is [REDACTED], dispatch at [REDACTED]. Driver [REDACTED] is stuck at the Dayton dock, call him at [REDACTED]. PRO [REDACTED] still shows no appointment. |
| Consistent typed tags | This is [CARRIER_CONTACT_1], dispatch at [CARRIER_1]. Driver [DRIVER_1] is stuck at the Dayton dock, call him at [PHONE]. PRO [SHIPMENT_1] still shows no appointment. |
| Synthetic replacement | This is Janet Price, dispatch at Price Freight Lines. Driver Tom Keller is stuck at the Dayton dock, call him at 555-0199. PRO 40517 still shows no appointment. |
| Plain hash | This is 7f3a9c, dispatch at b81e04. Driver 2cd5f1 is stuck at the Dayton dock, call him at e9a0d7. PRO 51c3b8 still shows no appointment. |
What the brokerage chose, field by field#
The brokerage chose consistent typed tags scoped to each shipment, because the value of these tickets lies in how a missed appointment moved from dispatcher to broker to warehouse and back. The carrier name became a tag too, since carrier relationships are commercially confidential even though a business name is not usually personal data.
Dock cities stayed, because lanes and facilities explain delays and a city alone does not identify a person. Phone numbers became a plain [PHONE] tag with no numbering, since no workflow depends on telling two numbers apart.
- People: consistent typed tags, scoped to the shipment or case.
- Emails and phone numbers: typed tags without numbering.
- Shipment, order and account numbers: consistent tokens with no retained mapping.
- Street addresses: generalized to city and state.
- Customer and carrier company names: tagged where contracts treat them as confidential.
How SourceX approaches replacement choices#
SourceX settles the replacement method field by field during the Preparation step of the SourceX five-step transaction, and the supplier reviews tagged samples before the Approval step. The placeholder vocabulary, the tag scope and any synthetic substitution are written into the dataset documentation and the privacy record of the SourceX Evidence Packet.
Frequently asked questions
What format should placeholders use?
Any format that cannot occur naturally in your records and is applied the same way everywhere. Square brackets with uppercase labels are common. In code datasets, avoid formats that look like template syntax or HTML, and document the full list of tags so a buyer can parse them reliably.
Should our own employees' names be replaced too?
Usually yes. Support agents, technicians and engineers are people whose names run through the whole archive. Role tags such as [AGENT_1] keep the workflow readable. Senior staff quoted in internal documents need the same treatment, and employee notices may also bear on what can be included.
Can we use different methods for different fields in the same dataset?
Yes, and most well-prepared datasets do. The important part is consistency within each field and clear documentation, so the buyer knows that [PHONE] means a removed number while [CUSTOMER_3] means the same customer throughout a case.
Does replacing names hurt model training?
Rarely in a meaningful way for operational records. What a model learns from a ticket is the problem, the steps and the resolution. Consistent tags preserve who did what. Where fluent names matter for a specific use, synthetic replacement can be discussed during scoping.
Sources
- Google's Sensitive Data Protection API supports de-identification transforms including format-preserving encryption and deterministic encryption. Source
Related resources
- QuestionDo AI labs buy financial data?
- QuestionDo AI labs buy spreadsheets?
- InsightCan financial advisors sell their data to AI companies?
- InsightCan data licensing affect my cyber insurance?
- InsightHow do I de-identify financial transaction data for AI training?
- SolutionData partnerships between businesses and AI developers
See if your company qualifies
A short company assessment. No data uploads are needed.