Privacy, de-identification and sensitive data
Datasets for training and evaluating PII detection and redaction models
Quick answer
A useful PII detection dataset pairs text that looks like your production traffic with span-level labels, an entity taxonomy that matches your redaction policy, and a held-out evaluation split that your detector has never seen. Public sets such as the Text Anonymization Benchmark and the Gretel PII masking data are good starting points, but synthetic or templated PII rarely reproduces the messy identifiers in real tickets, chats and documents, so most teams eventually need licensed, labeled operational text for evaluation.
By SourceX Editorial · Updated
Which public PII datasets are worth starting with
Public PII corpora are fine for bootstrapping a detector and comparing tools, but each one encodes a specific domain and labeling philosophy. NVIDIA's NeMo curation documentation, for example, evaluated its PII removal step on the Gretel PII masking dataset and the Text Anonymization Benchmark (TAB) [1]. TAB is the more demanding of the two: it is built from European court judgments and goes beyond tagging entity categories to record which spans must be masked to conceal the person being protected. That masking layer is what makes it a text anonymization benchmark rather than a plain NER corpus, and it tests quasi-identifiers that most NER sets ignore.
Clinical de-identification has its own lineage. The 2014 i2b2/UTHealth shared task covered de-identification of clinical narratives [6], and it remains a common reference point for PHI categories; see our guide to de-identifying clinical free text before reusing it. Tool comparisons also lean on Presidio's synthetic PII data [2], which is convenient but can share templates and recognizers with a Presidio-based tool being measured.
Why synthetic PII data stops being enough
Synthetic PII datasets measure whether a detector finds clean, well-formed entities in fluent sentences, which is the easy part of the problem. Generators fill templates like "My name is {first_name} {last_name} and my SSN is {ssn}", so entities sit in predictable syntactic slots with canonical formatting. Real operational text breaks those assumptions in ways that dominate production misses:
- Fragmented identifiers: account numbers split across lines, phone numbers with extensions, emails obfuscated as "jane at corp dot com".
- Context-dependent identity: "the night-shift supervisor at the Tulsa depot" identifies one person without a single name token, the quasi-identifier problem covered in indirect identifiers in business text.
- Structured noise: email signatures, quoted reply chains, CRM field dumps, JSON tool-call arguments and log lines inside free text.
- Ambiguous tokens: product names that are also surnames, ticket IDs that look like phone numbers, internal usernames.
Our working hypothesis, consistent with what practitioners report, is that recall measured on templated data overstates recall on this kind of text. The only way to know for your traffic is to evaluate on labeled samples that look like it.
Contamination: the benchmark your model already saw
A PII benchmark score is only meaningful if the detector was not trained on the same data family. Practitioners have warned that results can look optimistic when open models were fine-tuned on AI4Privacy-derived corpora and then tested on splits of that same family [3]. The same applies to Presidio-generated sets used to compare Presidio-based pipelines [2], and to LLM redactors, whose pretraining may include public benchmark text; recent research on LLM-based redaction examines both capability and risk [5].
Practical controls: record every training source in a data card [8], deduplicate evaluation text against training text with near-duplicate detection such as MinHash and LSH, and keep at least one private evaluation set that never leaves your environment. Also audit labels; widely used test sets carry an estimated average label error rate of at least 3.3% [7], and PII gold spans are no exception.
What a labeled PII evaluation record should contain
A redaction evaluation set needs character-offset spans, a stable label taxonomy and a masking decision, not just entity tags. Ask any supplier or annotation vendor for records shaped like the example below, with offsets computed on the exact delivered text encoding (UTF-8, normalized line endings).
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "tkt-000417",
"source_type": "support_ticket",
"text": "Hi, this is Dana R. from the Reno warehouse. Card ending 4417 was charged twice. Call me at 775-555-0142 x12.",
"spans": [
{"start": 12, "end": 19, "label": "PERSON", "identifier_type": "direct", "mask": true},
{"start": 29, "end": 43, "label": "ORG_LOCATION", "identifier_type": "quasi", "mask": true},
{"start": 57, "end": 61, "label": "PAYMENT_CARD_PARTIAL", "identifier_type": "quasi", "mask": true},
{"start": 92, "end": 108, "label": "PHONE", "identifier_type": "direct", "mask": true}
],
"annotation": {"guideline_version": "pii-v3.2", "annotators": 2, "adjudicated": true},
"split": "eval_private"
}
The identifier_type and mask fields mirror the distinction anonymization benchmarks such as TAB draw between detecting an entity and deciding to mask it. Without them, you can measure entity F1 but not whether the redaction actually protects the person.
How to choose between dataset options
The right mix depends on whether you are training, regression-testing or certifying a redaction model before it runs on sensitive data. Use this decision table as a starting point.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Need | Public benchmark (TAB, i2b2) | Synthetic PII (template or LLM generated) | Licensed real operational text, labeled |
|---|---|---|---|
| Bootstrap NER training | Domain-limited | Strong: cheap, unlimited, balanced labels | Useful for hard negatives |
| Compare tools fairly | Good if no tool trained on it | Weak: generator overlap [2] | Strong, if kept private |
| Estimate production recall | Weak: legal or clinical domain only | Weak: clean formatting | Strong when source types match traffic |
| Quasi-identifier masking | TAB covers it | Rarely modeled | Depends on guideline quality |
| Rights and access burden | License terms vary | Low | High: needs access controls and a license |
Measuring redaction quality, not just NER accuracy
Redaction models should be scored on recall per entity type and on residual identifiability, because a missed account number costs more than a false positive on a product name. Report span-level and token-level recall separately, since partial matches ("Dana" masked, "R." left) leak. Weight metrics by risk tier, track recall on each source type (tickets, chat, email bodies, OCR output) and sample residual PII after redaction, as described in auditing residual PII with sampling plans.
Treat no detector as complete. Presidio's maintainers state that because it relies on trained models, there is no guarantee it finds all sensitive information and other protections should be layered on [4]. For a deeper comparison of approaches, read LLM-based PII redaction versus NER and regex and our pipeline guide to PII redaction for LLM training data.
Sourcing real, labeled operational text for PII evaluation
Real-text evaluation data has to be licensed and handled as sensitive data even when your goal is to remove PII. That creates a tension: the most valuable evaluation records contain the identifiers you are trying to detect. Buyers usually resolve it in one of three ways, each of which belongs in the license and data card [8]:
- Labeled originals under strict access, held in a controlled environment for evaluation only.
- Surrogate-replaced text with span offsets preserved, so the model sees realistic formats without real identities; see masking versus surrogate replacement.
- Residual-PII test sets built from already de-identified text, used to measure what earlier pipelines missed.
SourceX sources operational datasets, such as support and sales histories, documents and finance or legal workflows, from US companies on request, and manages licensing. Personal details like names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. If you need labeled originals rather than surrogates, raise that requirement when you describe your data need to SourceX so it is assessed against supplier permissions. More context lives in the privacy hub and the PII detection glossary entry.
Find labeled PII data for your detector
SourceX finds US businesses that hold the operational text you describe, reviews rights and consents for each dataset, and delivers it under a license that defines records, uses, term and delivery. Data is sourced on request rather than held in stock, so a request does not guarantee a match. Describe the PII-labeled text you need.
Sources
- NVIDIA, "PII Identification and Removal (NeMo Framework User Guide 25.07)" (2025). https://docs.nvidia.com/nemo-framework/user-guide/25.07/datacuration/personalidentifiableinformationidentificationandremoval.html
- Ready Tensor, "PII detection tool comparison (Ready Tensor publication)". https://app.readytensor.ai/publications/8eAX8A1gfdkJ
- DEV Community, "Your PII redactor probably leaks tool call arguments". https://dev.to/crp4222/your-pii-redactor-probably-leaks-tool-call-arguments-357e
- Microsoft (microsoft/presidio), via pkg.go.dev, "Presidio - Data Protection API". https://data-privacy-stack.github.io/presidio
- arXiv:2508.05545, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
- ACL portal, "2014 i2b2/UTHealth Shared-Tasks and Workshop on Challenges in Natural Language Processing for Clinical Data" (2014). https://aclweb.org/portal/node/1851
- Northcutt, Athalye, Mueller, arXiv:2103.14749, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Pushkarna, Zaldivar, Kjartansson (Google Research), arXiv:2204.01075, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Pilan et al., "The Text Anonymization Benchmark (TAB): A Community Resource for Identifying and Handling Sensitive Information in Free Text" (2022). https://arxiv.org/abs/2202.00443
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.