Skip to content

Privacy, de-identification and sensitive data

Indirect identifiers in business text: job titles, rare events and small places

Quick answer

Indirect identifiers in text are details that name nobody on their own but single out a person when combined or read against outside knowledge. Examples include a unique job title, a rare incident, a small branch location, a date tied to an event, or a named B2B customer account. Named-entity recognition (NER) and regex redaction routinely miss them [1]. Treat them with a separate pass that inventories quasi-identifier types, measures how rare each combination is, and then generalizes, suppresses, perturbs or drops the record.

By SourceX Editorial · Updated

Why NER-based redaction leaves people identifiable

Redaction pipelines catch direct identifiers well but miss contextual ones, because identity in business text often comes from a combination of facts rather than from any single token. Tools such as Microsoft Presidio label spans like PERSON, EMAIL_ADDRESS and PHONE_NUMBER. The project itself warns that its models cannot guarantee finding all sensitive information [10]. Recent research shows some PII types evade both pattern matching and NER because the identifying signal sits in context, not in a recognizable entity [1].

The classic structured-data result carries over directly. Sweeney's k-anonymity work showed that linking a few attributes against an outside dataset can re-identify people in "anonymous" releases [8]. A later model-based estimate by Rocher and colleagues estimated that 15 demographic attributes would correctly re-identify 99.98% of Americans in any dataset, though that figure is a model estimate, not an observed rate [9]. Legal and policy scholarship keeps documenting re-identification through auxiliary information [2], and NIST's 2015 survey documents re-identification of data that had been released as de-identified [5].

Free text is worse than tables in one respect: nobody declared the quasi-identifier columns. "The only night-shift crane operator at the Dubuque yard" is one string, but it carries an occupation, a schedule, an employer site and a town. For the tabular counterpart of this problem, see de-identifying tabular and transactional data.

A quasi-identifier taxonomy for operational text

Operational corpora such as support tickets, incident reports, CRM notes, emails and HR or field-service logs leak identity through about seven recurring families. Each family needs its own detector and treatment, because each one combines with the others differently.

  1. Unique or rare roles. Titles such as "VP of Treasury", "sole HVAC tech for the north region" or "the new compliance hire". In a 40-person supplier, almost every title above individual contributor is a person.
  2. Rare events. A forklift injury, a $2.1M wire-fraud attempt, a data breach, a lawsuit, a termination for cause, an outage on a specific date. Rare events are often in local news, which gives an attacker the auxiliary data.
  3. Small places. Store numbers, branch codes, plant names, rural ZIP codes, building and floor ("3rd floor, Annex B"), or a sales territory that one rep covers.
  4. Dates plus events. "Out on leave since March 14", "joined two weeks before the audit". HIPAA Safe Harbor removes all date elements except year for this reason [3].
  5. Employer and B2B customer names. In B2B support data, the customer account name ("Acme Logistics, 12 seats") usually narrows the speaker to a few admins. HIPAA Safe Harbor's identifier list explicitly covers identifiers of the individual's employers, relatives and household members, not only the individual's own [3].
  6. Relational references. "The CFO's assistant", "Mike's replacement", "her husband also works in receiving". These describe a person through another identifiable person.
  7. Distinctive attributes and free-form numbers. Rare languages spoken, disabilities, unusual tenure ("32 years on the line"), vehicle fleet numbers, internal asset tags, ticket IDs reused in other systems, and Slack or Jira handles that NER labels as plain text.

Special-category signals such as health, union membership or immigration status often ride along with these. Handle them with the guidance in special category and sensitive data in operational training data.

Detecting indirect identifiers that NER misses

Detection works best as three layers: entity tagging extended with custom recognizers, corpus-level rarity statistics, and LLM-assisted or human review of the residue.

  • Custom recognizers. Add deny-lists and patterns for the supplier's own vocabulary: store and plant codes, territory names, internal product SKUs, job-title tables from the HRIS, and the CRM account list. Presidio accepts custom recognizers and spaCy accepts rule-based entity patterns; supplier reference tables beat any generic model here.
  • Rarity scoring. Extract candidate quasi-identifier values (normalized titles, locations, event types, month-year) per record and count them across the corpus. Any value or combination with a count below your threshold k is a candidate. This is k-anonymity applied to extracted features rather than declared columns [8].
  • Cross-record linkage. A ticket thread, an email chain or an incident timeline spreads facts across records. Group by thread, account or asset ID before computing rarity, or the per-record view will undercount exposure.
  • Contextual review. Send a stratified sample, oversampling low-count records, to a reviewer or an LLM classifier asked "could a coworker, customer or local reporter identify who this is about?" Recent work shows contextual PII needs this kind of reasoning rather than span labels [1]. LLMs also strengthen attackers, which is why LLM-assisted re-identification belongs in the threat model.

The ICO frames this as a "motivated intruder" test: could a reasonably competent person with access to public sources identify someone? [7] For operational text, the realistic intruders are a former colleague, a customer, a competitor and a journalist, and each brings different auxiliary knowledge.

Treatments: generalize, suppress, perturb or drop

Each detected indirect identifier gets one of four treatments, chosen by how much training signal the detail carries against how much risk it adds. NIST SP 800-188 lays out the same toolbox (suppression, generalization, perturbation and synthetic replacement) for government data releases [6].

Illustrative example: invented to show structure; it does not describe an available dataset.

Quasi-identifier familyExample spanDefault treatmentOutputSignal kept
Unique role"VP of Treasury"Generalize to a job-family ladder"[FINANCE_EXECUTIVE]"Seniority and function
Rare event"forklift tipped and broke his pelvis"Generalize severity; drop if public"[EQUIPMENT_INCIDENT, serious injury]"Incident class
Small place"Store #0417, Dubuque"Map to region or a consistent surrogate"[STORE_A] (Midwest)"Within-corpus linkage
Date + event"terminated on 2025-03-14"Shift dates per entity or coarsen to quarter"terminated in Q1"Sequence and interval
B2B customer"Acme Logistics, 12 seats"Surrogate plus size band"[CUSTOMER_117], 10-25 seats"Account behavior
Relational reference"the CFO's assistant"Generalize or suppress"an executive assistant"Role relation
Internal handle / asset tag"@jkowalski, asset FL-22"Consistent pseudonym"[USER_9], [ASSET_3]"Thread coherence
Record still unique after treatmentWhole ticketDrop record(removed)None

Rules of thumb for choosing a treatment:

  • Generalize when the attribute is the training signal (role seniority for an HR assistant, incident class for a safety model). Use a fixed hierarchy, such as an O*NET-style job family or a region map, so generalization stays consistent.
  • Suppress when the attribute adds little task value, such as a colleague's name in a sign-off.
  • Perturb dates with a consistent per-entity offset so intervals survive. Never shift each mention independently, because inconsistent dates break sequence tasks.
  • Drop the record when it stays unique after treatment. A corpus with 0.3% of tickets removed is usually a better asset than one carrying a single identifiable termination story.

Consistent surrogates matter for SFT and RAG. If "[CUSTOMER_117]" appears across 40 tickets, a model can still learn account-level patterns. Keep the surrogate map with the supplier, never in the delivered data.

Reviewer guidance with worked examples

Reviewers need concrete calls, not principles, so give them before-and-after pairs and an escalation rule.

Illustrative example: invented to show structure; it does not describe an available dataset.

BEFORE: "Escalated by Dana, the only bilingual dispatcher on Tulsa nights,
         after the 3/14 chlorine leak at the Peoria Ave plant."
RISK:   rare language + shift + small place + dated rare event = one person
AFTER:  "Escalated by [DISPATCHER] on the night shift after a
         [HAZMAT_RELEASE] at [PLANT_B] in [MONTH_1]."
ESCALATE IF: event was reported publicly -> drop record.

BEFORE: "Acme's IT admin (the one who keeps resetting SSO) opened a P1."
RISK:   B2B customer + behavioral descriptor narrows to one admin
AFTER:  "[CUSTOMER_117]'s IT admin opened a P1 about SSO resets."

Ask reviewers to record a reason code per edit (ROLE, EVENT, PLACE, DATE, ACCOUNT, RELATION, ATTRIBUTE). Reason codes let you measure which families the automated layer misses and retune it. That residual-miss rate is the evidence buyers should ask for; see how to measure what PII redaction misses.

What a buyer should request about indirect identifiers

Buyers should ask for evidence that indirect identifiers were treated, not only that names and emails were removed. A useful request list:

  • The quasi-identifier families the supplier's pipeline targets, and the custom dictionaries used (titles, sites, accounts).
  • The rarity threshold (k) and how records below it were handled, with counts of records generalized, suppressed and dropped.
  • Whether linkage was assessed at the thread or account level, not just per record.
  • The sampled manual review: sample size, oversampling of rare records, and the residual-miss rate by family.
  • The surrogate and date-shift method, confirming consistency and that keys stay with the data holder.
  • A motivated-intruder or re-identification assessment [7]. Our guide to re-identification risk assessment for licensed datasets covers methods to request.

For health text, the HIPAA bar is concrete. Under 45 CFR 164.514(a), information is de-identified only when there is no reasonable basis to believe it can identify the individual [4]. Safe Harbor also requires that the covered entity have no actual knowledge that remaining information could identify the person [3]. A rare-event narrative can fail that test even after all 18 listed identifiers are gone, which is why clinical free-text de-identification often relies on Expert Determination. The de-identification evidence package checklist lists the documents to collect, and the re-identification glossary entry and the insight on re-identification risk in business data give background for non-specialist reviewers.

Data type changes the profile. Job descriptions and interview notes are dense with role and employer signals; see licensing job descriptions and interview notes. Terse field notes compress place and event into a few tokens, as covered in operational free-text notes.

How SourceX handles personal details in operational text

SourceX sources operational datasets such as support and sales histories, engineering records, documents, and finance and legal workflows from US companies, on request rather than from stock. Before delivery, personal details such as names, emails, phones and account numbers are removed or replaced. The method is recorded and a sample is checked, though no method is perfect. Every dataset is rights-reviewed and comes with diligence materials covering source, rights, preparation and allowed use, which is where to raise the indirect-identifier questions above. You can describe the operational text you need, including your de-identification requirements. The broader privacy and de-identification hub and the AI data guides cover adjacent decisions.

Sourcing de-identified operational text with indirect identifiers treated

Describe the data, the model use and your de-identification requirements, and SourceX looks for US businesses that hold it. Every release is approved by the supplying company and delivered under a license that defines records, uses, term and delivery. A request does not guarantee a match. Start a buyer request.

Sources

  1. arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
  2. MJEAL, "Clarke – Spring 2026" (2026). https://www.mjeal-online.org/2026/04/12/clarke-spring-2026/
  3. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  4. eCFR, Office of the Federal Register, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  5. National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
  6. NIST Computer Security Resource Center, "NIST Publishes SP 800-188, De-Identifying Government Datasets" (2023). https://csrc.nist.gov/News/2023/nist-publishes-sp-800-188
  7. Information Commissioner's Office, "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  8. Data Privacy Lab, "k-Anonymity (De-identification Project)". https://dataprivacylab.org/projects/kanonymity/
  9. Rocher, Hendrickx, de Montjoye via IDEAS/RePEc, "Estimating the success of re-identifications in incomplete datasets using generative models (Nature Communications, 2019)" (2019). https://ideas.repec.org/a/nat/natcom/v10y2019i1d10.1038_s41467-019-10933-3.html
  10. Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data