Privacy, de-identification and sensitive data
Masking vs surrogate replacement: how PII redaction style changes what a model learns
Quick answer
For supervised and chat fine-tuning, realistic surrogates that stay consistent within each conversation usually train better than placeholder tags such as <PERSON>, because the model never sees an unnatural token it can learn to emit. Tags remain the safer choice for evaluation sets, audit trails and highly sensitive data, because a residual real identifier stands out next to them. Surrogates hide leaks as well as protect them, so pair them with a residual-PII audit and a documented convention.
By SourceX Editorial · Updated
This page covers the output format of redaction, not which detector to use. For detector choice see LLM-based PII redaction vs NER and regex, and for the full pipeline see PII redaction for LLM training data. Both sit under the privacy and de-identification hub.
The four output styles and what each one writes into your text
Every redaction tool makes the same decision after detection: what string goes where the identifier was. Microsoft Presidio, an open-source de-identification SDK [1], exposes this choice through anonymizer operators: replace (which can fall back to the entity type name, such as <PHONE_NUMBER>), redact (delete the span), mask (overwrite characters), hash and encrypt, the only reversible one. LLM-based redactors make the output style a prompt-level setting, so the same model can emit tags, deletions or fluent substitutes [2]. Curation stacks such as NVIDIA's NeMo data-curation tooling treat PII removal as one configurable pipeline stage among many [3].
In practice buyers meet four styles:
- Deletion. The span disappears: "Call me at tomorrow." Grammar breaks and turn lengths shrink.
- Typed placeholder.
[NAME],<EMAIL_ADDRESS>,{{ACCOUNT_ID}}. Readable, auditable, unnatural. - Character mask or hash.
***-***-4421,a3f9c1…. Preserves some shape; hashes look like noise. - Realistic surrogate. "Dana Whitfield",
dana.w@example.net, a plausible but fake order number. Fluent, and indistinguishable from real text by design.
Why placeholder tags become fine-tuning artifacts
Placeholder tags teach the model that tags are valid output. In SFT the loss is typically computed on assistant turns, so every agent reply that says "I've sent the label to [EMAIL]" rewards the model for writing [EMAIL] in production. A support assistant trained this way can greet users as "Hi [NAME]" or quote "order [ORDER_ID]" when it should have asked for the number or used a tool result.
The damage shows up in three recurring failure modes:
- Tag leakage in generations. The model emits bracketed tokens verbatim, sometimes in formats that break downstream templating such as Jinja or Handlebars that uses
{{ }}. - Coreference collapse. When every person becomes
[NAME], a three-party escalation ("[NAME] asked [NAME] to call [NAME]") loses who did what, and the model learns weaker entity tracking. - Tokenizer and distribution oddities. Tags split into rare subword sequences that appear thousands of times, skewing token frequencies and occasionally colliding with special-token syntax in your chat template.
Mitigations exist if you must keep tags: index them ([NAME_1], [NAME_2]), mask them out of the loss on assistant turns, or rewrite tags to surrogates at training time while storing the tagged master. That last option is often the best of both worlds.
Why realistic surrogates also hide the identifiers your redactor missed
Realistic surrogates conceal residual leaks because a missed real name looks exactly like the substituted fake ones. Carrell and colleagues named this "hiding in plain sight" (HIPS) in a 2013 JAMIA pilot on clinical progress notes [7]: because de-identification leaves some identifiers behind, replacing the detected ones with realistic surrogates made most residuals hard for readers to pick out. Later work by the same group, in which readers actively hunted for leaked identifiers in surrogate-resynthesized text, still found some of them, so the protection is real but partial.
The flip side matters to a buyer. With tags, a QA reviewer scanning a sample spots "Jennifer Ortiz" immediately among [NAME] tokens. With surrogates, the same reviewer cannot tell a missed real name from a generated one without the surrogate log. Presidio itself warns that ML-based detection offers no guarantee of finding all sensitive information [1], and NIST SP 800-188 cautions that traditional de-identification has inherent limits compared with formal privacy methods [5]. Surrogates do not change the detector's recall; they change who can see its misses.
That makes the surrogate map itself a sensitive artifact. Whoever holds the record of which spans were replaced can separate real residuals from fakes, so it should stay with the data supplier or a restricted audit function, not ship with the training data. Our residual PII audit sampling guide covers how to test a delivery against that log.
Keeping surrogates consistent across a conversation or ticket thread
Consistency within a scope preserves coreference, and randomness across scopes prevents linkage. The practical rule is a deterministic mapping keyed on the original value plus a scope identifier: the same customer becomes "Dana Whitfield" in every turn of conversation 8812, but a different surrogate in conversation 9034. A keyed HMAC over (scope_id, entity_type, normalized_value) selecting from a surrogate dictionary gives reproducible output without storing a global lookup table.
Choosing the scope is a privacy decision, not just an engineering one:
- Per conversation or ticket. Best default for SFT. Multi-turn reasoning survives; cross-record linkage does not.
- Per customer account across tickets. Preserves longitudinal context such as repeat contacts, but recreates a stable pseudonym, which is pseudonymization rather than anonymization in most frameworks. See de-identified vs anonymized definitions and the pseudonymization glossary entry.
- Global. Avoid. A single mapping across the whole corpus turns surrogates into a join key.
Normalize before mapping. "Dana", "Ms. Whitfield" and "dana.whitfield@" should map to one surrogate identity with matching first name, surname and email local part, or the model learns that people change names mid-thread. Agent names are a separate decision: many teams keep a small, consistent pool of fake agent names so the assistant persona does not inherit a real employee's name.
Format-preserving surrogates for numbers, dates and structured values
Structured identifiers need surrogates that keep their format, checksum and arithmetic, or the model learns broken patterns. A support model that sees ORD-48A1 where real orders look like ORD-2291847 will hallucinate the wrong shape. Guidelines for common types:
- Card numbers: for chat data, mask all but the last four (
**** 4421) rather than generating Luhn-valid fakes, which can collide with real card numbers and make it hard to show the data is out of PCI DSS scope. - Phone numbers: use reserved fictional ranges where they exist, such as North American 555-01xx numbers, in the original format.
- Emails and URLs: use reserved domains such as
example.comandexample.net, keeping the local-part style consistent with the surrogate name. - Account, order and claim IDs: preserve length, prefix and character class; format-preserving encryption (for example NIST SP 800-38G FF1) produces a reversible, same-shape value if the supplier needs re-linkage.
- Dates: shift by a per-scope random offset so intervals survive ("refund issued 3 days after the complaint"). Under HIPAA Safe Harbor, dates directly related to an individual other than the year must be removed, so date shifting on health data generally relies on Expert Determination instead [4]. See Safe Harbor vs Expert Determination.
- Free-text addresses: replace with plausible addresses in the same state or region if geography matters to the task; otherwise generalize to city or state.
Choosing a style by use case
The right style depends on whether the data trains generation, measures it, or exists for review. Use this table as a starting point and record the decision.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Use case | Recommended style | Why | Main risk to manage |
|---|---|---|---|
| SFT / chat fine-tuning on support or sales conversations | Consistent per-conversation surrogates | No tag artifacts; coreference intact | Residual leaks hidden; keep surrogate log restricted |
| Preference data (DPO, RLHF comparisons) | Same surrogates in chosen and rejected responses | Model must not learn surrogate identity as a preference signal | Divergent surrogates between pair members |
| Held-out evaluation set | Indexed tags or surrogates matching the training convention | Mismatch between train and eval style distorts scores | Eval graders penalizing tags as "unnatural" |
| PII-detector training or benchmarking | Original spans with labels, or surrogates with span annotations | Detector needs ground-truth spans | See PII detection datasets |
| Human QA and audit copies | Typed, indexed tags | Residual real identifiers are visible | Tags must never flow into training |
| RAG corpora over tickets and documents | Surrogates or generalization | Retrieved text is shown to users verbatim | Fake names presented as real customers |
Write the convention into the dataset card and the delivery spec
A redaction convention that is not written down will be applied inconsistently across batches and misread by the next team. Hugging Face recommends a dataset card for every dataset to inform users of limitations and promote responsible use [6]; for redacted conversations that card should state the replacement style per entity type, the consistency scope, the surrogate sources and the known residual-risk estimate. The same fields belong in the license exhibit or delivery spec so a later batch cannot silently switch from surrogates to tags.
Illustrative example: invented to show structure; it does not describe an available dataset.
redaction_convention:
version: "2026-10-r2"
detector: "NER + regex + LLM second pass" # see detector evaluation report
default_style: surrogate
consistency_scope: conversation_id
surrogate_seed: "held by supplier; not delivered"
entities:
PERSON: { style: surrogate, source: "fictional name list", keep_gender_cue: true }
EMAIL_ADDRESS: { style: surrogate, domain: "example.net" }
PHONE_NUMBER: { style: surrogate, range: "555-01xx", keep_format: true }
CREDIT_CARD: { style: mask, keep_last: 4 }
ORDER_ID: { style: format_preserving, keep_prefix: true }
DATE: { style: shift, offset: "per conversation_id" }
STREET_ADDRESS: { style: generalize, to: "city_state" }
audit_copy_style: indexed_tag # e.g. [PERSON_1]
residual_audit:
sample_size_conversations: 400
method: "reviewers with surrogate log"
result_ref: "audit-report-2026-10.pdf"
Questions to ask before you license redacted conversations
Ask the supplier for the convention, the evidence and the masters policy before you sign. Useful questions:
- Which style was used per entity type, and was it the same in every batch?
- What is the consistency scope, and how are name variants normalized?
- Who holds the surrogate map or encryption keys, and is any re-identification path delivered to you?
- How was residual PII measured, and did reviewers have the surrogate log when they did it?
- Can the supplier produce a tagged audit copy of a sample for your own QA?
- For health data, which HIPAA method applies, and is date shifting covered by the expert's determination [4]?
The de-identification evidence package checklist lists the documents that back these answers. If you are describing a request for support, sales or engineering conversations, the SourceX buyer intake asks you to describe the data rather than name businesses.
Sourcing redacted conversation data with a documented convention
SourceX sources operational datasets such as support and sales histories from US companies on request, and every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Describe the conversations you need at sourcex.si/buyers.
Frequently asked questions
Can I convert tags to surrogates myself after delivery?
Yes, if tags are typed and indexed ([PERSON2]), you can render consistent surrogates at training time and keep the tagged copy as your audit master. Unindexed tags lose coreference permanently, so request indexed tags if you plan to do this.
Do surrogates make data anonymous?
No. Surrogates reduce the visibility of residual identifiers but do not remove them, and a stable cross-record pseudonym keeps data in pseudonymized territory under most frameworks. Treat surrogate output as de-identified with residual risk, and audit it.
Should evaluation data use the same convention as training data?
Usually yes. If training used surrogates and evaluation uses tags, scores can reflect the style mismatch rather than task skill, and LLM judges may react to tag-filled references differently from fluent ones. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Sources
- Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
- NVIDIA, "PII Identification and Removal (NeMo Framework data curation)". https://docs.nvidia.com/nemo-framework/user-guide/25.07/datacuration/personalidentifiableinformationidentificationandremoval.html
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
- JAMIA (Carrell et al.), "Hiding in plain sight: use of realistic surrogates to reduce exposure of PHI in clinical text" (2013). https://pmc.ncbi.nlm.nih.gov/articles/PMC3638183/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.