Skip to content

Privacy, de-identification and sensitive data

LLM-based PII redaction vs NER and regex: accuracy, cost and leakage risk

Quick answer

LLM-based PII redaction finds context-dependent identifiers that regex and NER miss, such as a nickname tied to a rare job title, and one NVIDIA evaluation reported a 70B-class model scoring 26% better than Presidio on core categories [2]. It is also slower, costlier per token, nondeterministic and, if run through a hosted API, a disclosure of raw text [3]. For training corpora, the defensible design is a hybrid: regex for structured IDs, NER for names and places, and a constrained LLM pass for residual context, measured on your own labeled sample.

By SourceX Editorial · Updated

What each method actually detects in training text

Each method fails on a different class of PII, which is why choosing one is usually the wrong question. Regex and checksum validators (Luhn for card numbers, ABA routing checks, IBAN mod-97, SSN invalid-range rules) are near-perfect on well-formed structured identifiers and nearly blind to anything written in prose. Token-classification NER, whether spaCy pipelines, Presidio's analyzer recognizers or fine-tuned BERT-family encoders, labels PERSON, LOCATION, ORG and DATE spans well when the text resembles its training domain [4][6].

LLMs add the third layer: identifiers that only become identifying in context. Examples include "the only Spanish-speaking dispatcher on the Tulsa night shift," an account handle embedded in a URL path, a spelled-out phone number ("five five five, oh one..."), or a patient referred to by room and admission date. The PRvL study evaluates LLM redactors across architectures and adaptation strategies and documents both these capability gains and new risks introduced by using an LLM as the redactor [1]. For the end-to-end pipeline these methods plug into, see PII redaction for LLM training data.

Signal in the textRegex + validatorsNER (spaCy, Presidio, encoder)LLM pass
Card, SSN, IBAN, routing numbersStrong, deterministicWeak unless a recognizer is addedInconsistent; may miss digits
Emails, URLs, IPs, phone formatsStrongModerateStrong, including obfuscated forms
Person and place namesNoneStrong in-domain, weaker off-domainStrong, including nicknames
Quasi-identifiers (rare titles, events)NoneNonePartial; depends on prompt and schema
Code, logs, JSON payloadsStrong for known keysWeak; tokenization breaks spansModerate; long inputs truncate
Non-English and mixed-language textFormat-dependentModel-dependentGenerally stronger coverage

Where LLM redaction loses on accuracy

LLM redaction trades missed context for new error types that NER does not produce. The most damaging is span drift: the model rewrites or paraphrases surrounding text, so character offsets no longer align with the source and your audit trail breaks. Others are hallucinated entities (a redaction tag inserted where nothing was present), dropped content in long documents that exceed the effective context window, and run-to-run variance unless temperature is zero and outputs are schema-constrained.

Accuracy alone also hides what matters for privacy. A de-identification system with high token-level F1 can still leave the one surname that re-identifies a record, which is why end-to-end evaluation should weight recall on direct identifiers and check residual risk rather than report a single score [5]. Presidio's own maintainers state the tool cannot guarantee finding all sensitive information and recommend additional protections [4]; the same caution applies to any LLM.

Rules of thumb that hold in practice:

  • Ask the LLM to return spans (start, end, label) as JSON, not rewritten text, and apply replacements in code.
  • Reject outputs whose spans do not match the source substring exactly.
  • Chunk long documents with overlap so entities at boundaries are not lost.
  • Keep a deterministic regex layer even when the LLM appears to catch everything, because format-valid IDs deserve a validator, not a probabilistic guess.

Cost and throughput at corpus scale

NER and regex scale with CPU and are cheap enough to run on every document; a full LLM pass over a pre-training corpus is a GPU budget decision. A 70B-class model running inference over billions of tokens costs orders of magnitude more compute than an encoder doing token classification, and generation time grows with output length, which is another reason to request span JSON instead of rewritten documents.

The practical pattern is routing. Run regex and NER on everything, then send only the documents that need a context pass to the LLM: free-text fields with low NER confidence, records containing domain-specific risk terms (diagnosis codes, case numbers, employee IDs), or sources known to carry quasi-identifiers such as support tickets and incident reports. NVIDIA's NeMo Curator documents an LLM-based redaction module as one configurable stage in a curation pipeline, which fits this routing pattern [2].

For SFT and evaluation sets, which are small relative to pre-training data, running every record through an LLM pass is usually affordable and the accuracy gain is worth it. For continued pre-training on large domain corpora, routing is the only realistic design.

Leakage risk: the redactor can become the breach

Using a hosted LLM API to redact PII sends the unredacted PII to a third party, which is often the very disclosure you are trying to avoid. Vendor comparisons of managed redaction APIs note that these services receive the original data [3]. Before choosing one, check whether the provider retains prompts, whether data is used for its own training, the region of processing, and whether your data license or a HIPAA business associate agreement even permits the transfer.

Local or self-hosted models remove that transfer but introduce their own exposure. Prompts, raw inputs and model outputs land in inference logs, tracing tools and failed-job caches; those stores need the same access controls as the source data. PRvL also examines privacy risks that arise from the LLM redactor itself, not only its accuracy [1].

Residual PII matters more after training than before it. Research on fine-tuned models shows fine-tuning can amplify a model's tendency to reveal personal information from its training data [9], so a missed identifier in an SFT set is more likely to resurface than the same miss in a log archive. Pair redaction with a post-training leakage audit using canaries and membership inference.

Replacement style changes what the model learns

Detection is half the decision; what you write in place of the span shapes the trained model. Masking with typed tags such as [PERSON] is auditable but teaches the model an unnatural token distribution, while consistent surrogates ("Maria Lopez" becomes "Dana Whitfield" throughout a document) preserve coreference and fluency but can look like real data. LLMs are good at generating format-preserving surrogates; they are also capable of leaking the original value into the surrogate, so validate that no surrogate matches a source identifier. The trade-offs are covered in masking vs surrogate replacement.

A hybrid pipeline design for training corpora

The most robust design layers deterministic, statistical and generative detectors and merges their spans before any text is changed. Each layer covers the blind spots of the one before it, and the merge step resolves overlaps by preferring the longest span and the most specific label.

Illustrative example: invented to show structure; it does not describe an available dataset.

StageTool classInputOutputGate to pass
1. Structured IDsRegex + checksum validatorsAll recordsSpans with label and validator resultValidator unit tests pass
2. Named entitiesspaCy or Presidio analyzer, or fine-tuned encoderAll recordsSpans with confidence scoreRecall on labeled sample at or above target
3. RoutingRules on confidence, field type, risk termsStage 2 outputSubset flagged for context passRouting rate logged per source
4. Context passLocal LLM, temperature 0, JSON span schemaFlagged subset, chunked with overlapAdditional spansEvery span matches source substring exactly
5. Merge and replaceCode, not the LLMAll spansMasked or surrogate text plus span logNo overlapping or orphan spans
6. Residual checkSecond detector plus human review of a sampleRedacted outputMiss rate by entity typeMisses on direct identifiers reviewed and fixed

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "tkt-000417",
  "source_field": "ticket_body",
  "spans": [
    {"start": 18, "end": 29, "label": "PERSON", "detector": "ner", "score": 0.97},
    {"start": 64, "end": 76, "label": "PHONE", "detector": "regex", "validator": "nanp_ok"},
    {"start": 102, "end": 141, "label": "QUASI_ID", "detector": "llm", "rationale": "unique role + site"}
  ],
  "replacement_style": "surrogate",
  "pipeline_version": "redact-2026.10.1"
}

Logging the detector per span lets you measure what the LLM pass actually adds. If the context layer contributes almost nothing on a given source, drop it there and save the compute.

How to evaluate redactors on your own data

Published comparisons are directional; the only number that should drive your choice is measured recall on a labeled sample of your corpus. Build a gold set by double-annotating a stratified sample across sources and field types, adjudicating disagreements, and using an entity schema that separates direct identifiers from quasi-identifiers. Public benchmarks can help bootstrap, and datasets for training and evaluating PII detection lists options, but NER models evaluated on clinical text show why in-domain measurement matters, and domain shift is a common reason NER underperforms [6].

Report per-entity recall and precision, not one micro-averaged F1, and treat a missed direct identifier as more severe than a false positive [5]. Run the LLM configuration several times to measure variance. Then test residual risk on the redacted output, because LLMs can also re-identify people from text that passed redaction; see LLM-assisted re-identification.

For health text, detection quality alone does not settle compliance. HIPAA de-identification requires either Safe Harbor removal of the 18 listed identifiers or an Expert Determination that re-identification risk is very small [8], and NIST SP 800-188 cautions that traditional de-identification has inherent limits compared with formal privacy methods [7]. Clinical specifics are covered in de-identifying clinical free text.

What to ask when licensing already de-identified text

If you are buying text rather than redacting it yourself, ask the supplier which detector classes were used, whether an LLM ran on raw data and where, the replacement style, the measured miss rate by entity type, and whether a sample was reviewed by a person. SourceX removes or replaces personal details such as names, emails, phone numbers and account numbers before delivery, records the method used and checks a sample, while being clear that no method is perfect; health records require HIPAA Safe Harbor or Expert Determination. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which gives your privacy reviewer something concrete to test. You can describe the data you need on the SourceX buyer page, and the privacy cluster hub and the PII detection glossary entry cover related terms.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Source de-identified business text for training

SourceX sources operational text such as support and sales histories, engineering records and documents from US companies on request, with each dataset rights-reviewed, de-identified with a recorded method and delivered under a license that defines records, uses, term and delivery. Data is not held in stock and a request does not guarantee a match. Describe the corpus you need at sourcex.si/buyers.

Frequently asked questions

Is a local LLM safe enough to redact sensitive training data?

A local model avoids sending raw data to a third party, but the inputs and outputs still pass through logs, caches and tracing systems that must be access-controlled like the source data. PRvL also flags privacy risks that come from using an LLM as the redactor [1]. Treat the redaction environment as holding unredacted PII.

Can an LLM replace Presidio entirely?

It can outperform Presidio on some categories in reported evaluations [2], but it is costlier, nondeterministic and weaker as a validator of structured IDs. Presidio's maintainers say no detector can guarantee finding all sensitive information [4], so a practical design keeps Presidio-style recognizers and regex and adds an LLM pass where context matters.

Which metric should decide between redaction models?

Use recall on direct identifiers, broken down by entity type and source, measured on a labeled sample of your own corpus [5]. Precision matters for data utility, but a missed identifier is the failure that creates privacy exposure.

Sources

  1. arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
  2. NVIDIA, "PII Identification and Removal (NeMo Framework user guide 25.07)" (2025). https://docs.nvidia.com/nemo-framework/user-guide/25.07/datacuration/personalidentifiableinformationidentificationandremoval.html
  3. Grepture, "Best PII Redaction APIs for LLMs (2026)" (2026). https://grepture.com/compare/best-pii-redaction-apis-for-llms
  4. Microsoft presidio project (pkg.go.dev index), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  5. arXiv, "Automatic end-to-end De-identification: Is high accuracy the only metric?" (2019). https://arxiv.org/pdf/1901.10583
  6. arXiv, "A Comparative Evaluation Of Transformer Models For De-Identification Of Clinical Text Data" (2022). https://arxiv.org/pdf/2204.07056
  7. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  8. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  9. arXiv, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data