Privacy and preparation
How accurate is automated PII redaction on real business text?
By SourceX Editorial · Updated
Short answer
Automated PII redaction is accurate on identifiers with a fixed shape, such as email addresses, phone numbers and card numbers, and much less reliable on names, places and details that identify someone only in context. Published benchmarks rarely resemble your tickets or job notes, so measure recall on a labeled sample of your own records before trusting any tool.
Key takeaways
- Recall, the proportion of real identifiers a tool catches, is the accuracy measure that decides release risk.
- Patterned identifiers are detected far more reliably than person names, partial addresses and contextual clues.
- Business text breaks benchmark assumptions with lowercase chat, quoted replies, signatures and internal codes.
- Test on a labeled sample from each source system, read every miss, and re-test on fresh records after each change.
What does accuracy mean for PII redaction?#
Accuracy for PII redaction has two parts that pull against each other. Recall measures how many of the real identifiers in a set of records the tool found; precision measures how many of the items it flagged were really identifiers. A tool tuned to flag everything has high recall and poor precision.
For releasing data, recall matters most, because a missed name leaves the company exposed while an over-redacted product name only costs some usefulness. Precision still counts: a tool that replaces every capitalized word with a placeholder strips the operational detail that makes records worth licensing.
A third measure is the one a release decision rests on: the residual rate, meaning the identifiers still present in the final output after automated and human steps. Recall describes the tool; the residual rate describes the dataset.
Which PII types do tools detect reliably?#
PII types differ widely in how reliably tools detect them. The ranking below is general and qualitative, drawn from how detection methods work; your own measurements should replace it.
Patterned identifiers do well because tools can check them mechanically. Presidio, for example, combines named-entity recognition with regular expressions, rule-based logic and checksums with context. Names and places depend almost entirely on the recognition model, which is where errors concentrate.
| PII type | Example in business records | Typical reliability | Why |
|---|---|---|---|
| Email addresses | Contact in a support thread | High | Fixed pattern with an at sign and a domain |
| Payment card numbers | Number pasted into a ticket | High | Pattern plus checksum validation |
| Phone numbers | Callback number in a dispatch note | High to medium | Formats, extensions and spacing vary |
| Government ID numbers | Tax ID in a vendor form | Medium | Patterns collide with order and part numbers |
| Person names | Technician or customer named in a note | Medium to low | Lowercase, nicknames, surnames alone, names that are also words |
| Street addresses and places | Job site or delivery location | Low to medium | Partial addresses and local references |
| Contextual identifiers | The only night-shift supervisor at the north plant | Low | No pattern; identifies only through context |
Why benchmark results overstate accuracy on your records#
Benchmark results overstate accuracy because test sets are usually clean prose, news text or synthetic records, while business text is messy. Support tickets mix quoted replies, signatures and pasted tables; dispatch notes use shorthand and lowercase; engineering comments mention people by username.
Internal vocabulary causes both kinds of error. A product called Jordan or a crew named after a street gets flagged as a person or a place, while a customer called only the Hendersons in a job note may slip through. Long digit strings such as order, job or ticket numbers are often mistaken for phone or ID numbers.
Tool documentation is candid about this. Presidio's maintainers caution that automated detection comes with no promise of catching every sensitive item and advise layering further protections on top, a caveat that applies with equal force to commercial services.
How to measure recall on your own records#
Measuring recall on your own records takes a labeled sample, a fixed scoring method and the discipline to read every miss. The work is modest compared with the cost of releasing a dataset on an assumption.
Keep the labeled sample and the scoring script under version control. When a source system or a tool version changes, the same test shows whether recall moved.
- Draw a random sample from each source system and record type, including the messiest ones.
- Have two people label every identifier by type, then settle disagreements with a written rule.
- Run the tool with the exact configuration you plan to use in production.
- Score recall and precision separately for each PII type and each source system.
- Read every miss and sort it by cause: casing, format, context, unknown name or attachment.
- Change one thing at a time, then re-test on a fresh labeled sample rather than the one you tuned on.
What improves accuracy the most?#
Accuracy improves most from changes that teach the tool about your business, not from switching tools. The levers below are listed with their usual direction of effect, which your own tests should confirm.
Order the work by cost. Start with pre-processing and dictionaries, which are cheap and address common misses in business text, then add custom recognizers, and consider a second model only after that. Measure each step on its own so you know which change moved the results, and keep the ones that earned their place.
| Lever | Effect on misses | Effect on false positives |
|---|---|---|
| Dictionaries of known names from the CRM, HR roster and vendor list | Fewer misses for names in those lists | Slight rise where names are also common words |
| Custom recognizers for your ID formats | Catches account and job numbers that embed personal data | Lower, because generic number rules can be relaxed |
| Allow lists for product, crew and location names | No change | Clear drop |
| Pre-processing that strips signatures and quoted replies | Fewer repeated misses in long threads | Fewer |
| A second recognition model run alongside the first | Catches names one model misses | Rises unless outputs are reconciled |
| Human review of high-risk records | Catches contextual identifiers | Reviewers can restore wrongly redacted terms |
Illustrative: an engineering firm tests a redaction tool on RFI logs#
Illustrative: a fictional civil engineering firm wants to license RFI logs, submittal reviews and project correspondence exported from its project management system. Before choosing a tool, its IT lead labels a random sample of records from each source with help from a project manager.
The test shows strong recall on emails and phone numbers and weak recall on names. Site superintendents appear by surname only, and lot and parcel references identify homeowners on residential subdivisions. The tool also flags a road name used as a project code as a person.
The firm loads its contact list as a dictionary, adds a recognizer for its parcel reference format and an allow list for project codes. A re-test on fresh records shows the remaining misses concentrated in free-text site observations, so those go to human review, and correspondence with homeowners on residential jobs is excluded.
How SourceX measures redaction in Preparation#
In Preparation, the third step of the SourceX five-step transaction, redaction is measured on the supplier's own records rather than assumed from a tool's reputation. The labeled sample, the results by PII type and the changes made are all recorded.
That record becomes part of the privacy record in the SourceX Evidence Packet, alongside the residual review of the final output. The supplier sees the results before approving release, and the licensee sees how the dataset was prepared.
Frequently asked questions
Is there a recall level that counts as good enough?
No single level fits every dataset. The bar depends on how sensitive the records are, what a miss would expose and what other safeguards apply, such as human review and contract terms. Agree the standard with counsel before testing, and judge the final dataset by its residual rate, not by the tool's recall alone.
Do LLM-based detectors beat rule-based tools?
They often handle context better, catching a lowercase name or a description that identifies someone. They can also be inconsistent between runs, cost more at volume and raise questions about where the text is sent. Measure them on the same labeled sample as any other tool before deciding.
How large should the labeled sample be?
Large enough to contain many examples of each PII type you care about, from every source system. Rare types, such as government ID numbers, may need targeted sampling. A statistician or an experienced privacy engineer can size the sample for the confidence your release standard requires.
Does accuracy drop for records that mix languages?
Usually, yes. Many recognition models perform best in English, and names from other languages may be missed or misclassified. If your records mix languages, label enough of each to measure them separately, and consider excluding records in languages the tool handles poorly.
Should we re-test when we add a new source system?
Yes. A tool that performs well on support tickets can do poorly on dispatch notes or chat logs, because the writing style and the identifiers differ. Label a fresh sample from each new system and score it before including that system in a release.
Sources
- Presidio combines named-entity recognition, regular expressions, rule-based logic and checksums with context to identify PII. Source
- Presidio's documentation warns that because it is using automated detection mechanisms, there is no guarantee that Presidio will find all sensitive information, and that additional systems and protections should be employed. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.