Privacy, de-identification and sensitive data
PII redaction for LLM training data: pipelines, tools and how to measure what they miss
Quick answer
PII redaction for LLM training data is a measured pipeline, not a single tool: layer deterministic patterns, an NER or transformer detector and, where budget allows, an LLM pass; replace findings with typed placeholders or consistent surrogates; then measure recall per entity type on a hand-labeled sample of your own corpus. No detector finds everything [2], vendor accuracy numbers rarely transfer, and the number that matters is how much personal data survives into the training shards.
By SourceX Editorial · Updated
Why training text needs a different redaction bar than logs or analytics
Training text needs a stricter bar because a model can memorize and later emit a phone number that appeared only a few times, and that exposure cannot be patched after the checkpoint ships. Research on fine-tuning shows that tuning on data containing personal information can amplify the risk that the information is recovered from the model [7]. NIST's generative AI profile lists data privacy among the risks that generative AI creates or worsens [8], which is how your security reviewers will frame it.
The asymmetry also changes the metric. In analytics, a few false negatives are an acceptable cost of keeping fields usable; in a pre-training or SFT corpus, a missed account number is a durable defect, while an over-redacted job title is a small quality loss. That is why recall, not F1, leads the scorecard below. Over-redaction still matters, and is covered in measuring over-redaction in a de-identified dataset.
The four-stage pipeline: detect, replace, sample, measure
A defensible pipeline has four stages, and each stage writes its own log so that the next stage can be audited. Treat the detector as one component, not the control.
- Normalize and segment. Decode HTML and quoted-printable email bodies, strip signatures and reply chains into separate fields, flatten JSON in ticket exports so that nested values (custom fields, tool-call arguments, attachment names) are scanned as text. Redactors frequently miss personal data inside structured payloads such as tool-call arguments [6].
- Detect in layers. Run regex and checksum validators first (Luhn for card numbers, ABA routing check digits, IBAN mod-97, US SSN invalid-range rules (000, 666, 900–999 area numbers), email and E.164 phone patterns). Add a statistical detector, such as Microsoft Presidio's analyzer with spaCy or transformer recognizers [2], for names, locations and organizations. Route low-confidence spans or high-risk document types to an LLM pass if the budget allows.
- Replace. Choose typed masks (
[PERSON],[EMAIL]) or consistent surrogates (the same fake name for the same real name within a conversation). The choice changes what the model learns; see masking vs surrogate replacement. - Sample and measure. Hand-label a stratified sample of redacted output, compute residual rates per entity type, and gate the shard on thresholds agreed before the run.
This page owns the pipeline and its measurement. Definitions live in the glossary entries on PII detection and redaction, and a plain-language overview is in how to remove sensitive details from business data for AI training.
Choosing detectors: Presidio, transformer NER and LLM passes
The right detector mix depends on your text type and volume, and published comparisons should only shortlist tools, never pick them. Presidio is a widely used open-source SDK for detecting and anonymizing PII, but the project itself says it cannot find all sensitive information and recommends additional protections [2]. Presidio publishes no official accuracy benchmark, so teams have to measure it on their own data [3].
LLM-based redaction can catch context-dependent identifiers that NER misses, such as "my manager in the Boise office" or a customer's rare medical device. In NVIDIA's own evaluation, the NeMo Curator LLM-based approach outperformed Presidio by 26% on core categories [1]; that is a vendor result on its chosen benchmarks, not a property of your tickets. Academic work on LLM redaction also measures risks, not only capability, including outputs that alter or leak content [5]. A deeper cost and leakage comparison is in LLM-based PII redaction vs NER and regex.
Watch for benchmark contamination and benchmark mismatch. A tool tuned on a public PII masking set may score well there and fall off on enterprise email or chat transcripts, and comparisons built on one public set can flatter the tool that trained near it [6]. For candidate evaluation sets, see datasets for training and evaluating PII detection models.
| Layer | Strong on | Weak on | Typical failure mode |
|---|---|---|---|
| Regex + checksum | Cards, SSNs, IBANs, emails, phones | Names, addresses, free-form IDs | Misses reformatted numbers ("four one one one...") |
| NER / transformer (e.g., Presidio recognizers) | Person, location, organization | Domain IDs, indirect identifiers | Drops names in lowercase chat or non-English text |
| LLM pass | Context-dependent and indirect identifiers | Throughput, determinism, cost | Paraphrases or truncates text it was asked to copy |
Measuring what the pipeline misses: per-entity recall on your own corpus
Measure recall per entity type on a labeled sample of your actual corpus, because an aggregate score hides the categories that cause real harm. A detector with 95% overall recall can still miss most account numbers if account numbers are rare in the sample. At least one published comparison of transformer models for PII redaction uses recall as its primary metric for this reason [4], and macro averaging across entity types keeps rare categories visible.
Build the gold set before tuning anything. Draw a stratified sample by source system (Zendesk tickets, Salesforce case comments, Gmail exports, call transcripts), language and document length; have two annotators label spans independently; and adjudicate disagreements. Score at span level with a lenient overlap rule for boundaries, but count a partially masked email (j***@acme.com left intact after the @) as a miss.
Then report residuals, not just detector recall. Residual rate is the share of sampled records that still contain at least one true identifier after replacement, broken down by entity type. That is the number procurement, privacy counsel and auditors can act on, and it maps cleanly to a sampling plan; see auditing residual PII in a delivered dataset.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Entity type | Gold spans in sample | Detected | Recall | Residual records (of 2,000) | Gate | Status |
|---|---|---|---|---|---|---|
| 1,412 | 1,409 | 99.8% | 2 | >= 99.5% | Pass | |
| PHONE | 868 | 851 | 98.0% | 11 | >= 98% | Pass |
| PERSON | 3,240 | 3,011 | 92.9% | 141 | >= 95% | Fail: add LLM pass on chat |
| ACCOUNT_ID | 214 | 171 | 79.9% | 38 | >= 97% | Fail: add custom recognizer |
| STREET_ADDRESS | 302 | 276 | 91.4% | 19 | >= 95% | Fail: review signature blocks |
Pre-training corpora versus fine-tuning sets
Scrubbing a pre-training corpus and redacting a fine-tuning set call for different trade-offs, mainly because of volume and how many times each example is seen. At billions of tokens, an LLM pass on every document is usually too expensive, so teams run regex and NER broadly and reserve LLM review for high-risk sources or sampled shards. Deduplicate first: near-duplicate removal shrinks the volume to scan and reduces repeated exposure of the same identifier, as described in near-duplicate detection with MinHash and LSH.
SFT and eval sets are smaller and seen more often per example, so each residual counts more. Scan both inputs and targets; an assistant turn that echoes a customer's address is as much a leak as the user turn. The detector choice for that stage is covered in scanning a training corpus for PII before fine-tuning.
Failure modes that survive a passing score
Several leak paths survive a good recall number because they sit outside what the detector was scored on. Common ones in licensed business text:
- Indirect identifiers. Job title plus small town plus incident date can single out a person with no name present; see indirect identifiers in business text.
- Inconsistent surrogates. "Maria" replaced by three different fake names in one thread breaks coreference and tempts reviewers to restore originals.
- Non-text carriers. File names, email headers, OCR layers and audio carry PII the text scanner never sees; scanned files need the approach in redacting PII in scanned documents.
- LLM rewriting. An LLM redactor may silently paraphrase, which corrupts the training signal; diff its output against the input and flag any edit outside detected spans [5].
- Re-identification by stronger models. De-identified text can be re-linked by capable LLMs, so test with an attack, not only a scan; see LLM-assisted re-identification.
Redaction is one layer among several; the Presidio project itself recommends additional systems and protections alongside it [2].
What to ask a supplier about their redaction run
When text arrives already redacted, ask for the evidence that would let you reproduce the measurement, not for a tool name. Useful requests: the detector stack and versions, the entity taxonomy, the replacement method and whether surrogates are consistent, the gold-set sampling design and its per-entity recall, the residual rate with confidence intervals, and the handling of attachments and structured fields. Process frameworks such as ISO/IEC 5259-4 give a vocabulary for documenting data quality steps for ML training data [9]. A full list of documents is in the de-identification evidence package checklist.
SourceX removes or replaces personal details such as names, emails, phone numbers and account numbers before delivery, records the method used and checks a sample, and no method is perfect. Health records require HIPAA de-identification by Safe Harbor or Expert Determination. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which gives your team a starting point for its own residual audit. You can describe the text you need on the SourceX buyer page. For the wider cluster, start at de-identified data for AI training or browse the AI data hub.
Sourcing licensed text with documented redaction
SourceX sources operational datasets, such as support and sales histories, documents and engineering records, from US companies on request, and every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Nothing is contracted until the supplying company agrees, and a request does not guarantee a match. Describe the corpus, entity types and redaction evidence you need at sourcex.si/buyers.
Sources
- NVIDIA, "PII Identification and Removal (NeMo Curator user guide 25.07)" (2025). https://docs.nvidia.com/nemo-framework/user-guide/25.07/datacuration/personalidentifiableinformationidentificationandremoval.html
- Microsoft (microsoft/presidio project), via pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- Grepture, "Presidio is not enough for PII redaction". https://grepture.com/blog/presidio-not-enough-pii-redaction
- Ready Tensor, "Comparison of transformer models for PII redaction". https://app.readytensor.ai/publications/8eAX8A1gfdkJ
- arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
- DEV Community, "Your PII redactor probably leaks tool-call arguments". https://dev.to/crp4222/your-pii-redactor-probably-leaks-tool-call-arguments-357e
- arXiv, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and ML, Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.