Skip to content

Privacy, de-identification and sensitive data

Found personal data in a licensed dataset? A buyer's response playbook

Quick answer

If you find personal data in a dataset that was delivered as de-identified, stop new use, quarantine the affected files and every derived artifact, and scope the problem with a structured sample before anyone argues about severity. Then notify the supplier in writing with evidence, request a corrected delivery, and decide on the model separately: filter and continue, retrain from a clean checkpoint, or keep with documented controls. Record each decision, because counsel and auditors will ask for it later.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why residual personal data is an expected event, not a scandal

Residual PII is a normal failure mode of every redaction pipeline, so the right response is a rehearsed process rather than an improvised one. Microsoft's Presidio project states plainly that its ML-based recognizers cannot guarantee every sensitive item is found and that other protections should be layered on top [1]. Practitioner write-ups describe the same gap in production: names inside free text, IDs split across tokens, and identifiers in fields the pipeline was never configured to scan [2]. NIST SP 800-188 makes the governance point: traditional de-identification has inherent limits, and an organization should plan for them rather than assume perfection [10].

The typical places residuals hide in operational data are predictable:

  • Free-text fields: support ticket bodies, CRM notes, Jira comments, email quoted replies and signatures.
  • Structured fields that look harmless: customer_ref, external_id or invoice_no values that map one-to-one to a person.
  • Attachments and embedded layers: OCR text under redaction boxes in PDFs, EXIF and author metadata, spreadsheet hidden columns.
  • Media: names spoken in call audio that the transcript redacted, or faces and screens in recordings of hands-on work.
  • Quasi-identifiers in combination: a ZIP code, job title and date that single out one person even with no name present.

For deeper treatment of detection itself, see how to measure what PII redaction tools miss and auditing residual PII with sampling plans.

The first 72 hours: contain before you characterize

Containment comes first because every hour the data stays in active pipelines multiplies the number of copies you will later need to trace. Freeze, do not delete: deleting files destroys the evidence you need for the supplier conversation and for your own records.

  1. Freeze ingestion. Pause jobs that read the dataset: tokenization, embedding builds, fine-tuning runs, eval harnesses and RAG index refreshes.
  2. Quarantine by hash. Move affected files to a restricted bucket or prefix with read access limited to the incident team; keep a manifest of file paths and SHA-256 hashes.
  3. Inventory derivatives. List every artifact built from the dataset: shards, tokenized caches, vector indexes, synthetic data generated from it, checkpoints and adapters, eval sets and labeled subsets.
  4. Lock access logs. Preserve object-storage access logs and notebook histories so you can show who touched the data.
  5. Open a single incident record. One owner, one timeline, one place for decisions.

If the dataset already feeds a retrieval system, treat the vector index as live exposure: a RAG application can return the raw passage to any user, which is a faster leak path than model memorization.

Scoping: how much personal data, of what kind, and where

Scoping turns "we saw a phone number" into a defensible estimate of prevalence, category and spread. Draw a stratified random sample across files, time periods and field types rather than grepping only for the pattern you already found, and run a second detector with a different method (regex plus NER, or NER plus an LLM pass) so you are not measuring one tool against itself [1][2].

Classify each hit by category, because category drives the legal and commercial response:

  • Direct identifiers: names, emails, phone numbers, account numbers, SSNs, card numbers.
  • Sensitive categories: health information, financial account data, biometrics, children's data.
  • Quasi-identifiers: combinations that support re-identification by linkage.
  • Third parties: employees of the supplier, or customers of its customers.

Category changes the frame. Health data that was supposed to meet HIPAA Safe Harbor fails that method if any of the 18 listed identifiers remains, and the alternative route is a fresh Expert Determination [7]. Under GDPR, data that is identifiable by means reasonably likely to be used is personal data regardless of how it was labeled at delivery [8]; the UK ICO frames the same question through the motivated intruder test [9]. For the definitional differences, see de-identified vs anonymized legal definitions.

Notifying the supplier: what to send and what to ask for

Notify the supplier in writing promptly, with evidence the supplier can reproduce, because the supplier usually controls the raw source and the redaction configuration. Check your license first: many agreements contain notice duties, cure periods, replacement-delivery remedies and deletion obligations, and the notice should follow the clause the contract specifies. Do not send the personal data itself by email; reference record IDs and hashes, and share samples only through the access-controlled channel the parties already use.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample entry
incident_idDSI-2026-014
dataset / deliverysupport_tickets_v3, delivery 2 of 4, manifest hash a91f...
discovered2026-09-30, scheduled residual-PII audit
detection methodregex + NER pass, then manual review of flagged records
sample design2,000 records stratified by month and channel
findings37 records with customer phone numbers in body; 4 with full names in signature
categorydirect identifiers; no health or payment data found in sample
affected derivativestokenized shard set T3, RAG index kb-v7, SFT run ft-0912
containmentingestion paused 2026-09-30; files quarantined; RAG index offline
request to supplierroot cause in redaction config; corrected delivery for affected files; confirmation of method used
model decisionpending; see decision table
ownerAI governance lead

Ask the supplier for three things: the root cause (which field or recognizer failed), a corrected delivery processed with the fixed configuration, and an updated description of the de-identification method so your documentation stays accurate. If the supplier's own systems were breached, the situation is different; see licensing data after a security incident.

Retrain, filter or keep: deciding what happens to the model

The model decision depends on exposure, not on the existence of the defect. Memorization research shows that models can emit training data verbatim [4], that memorization rises with how often a sequence repeats in the corpus [5], and that leaked PII can surface in generation [3]. A phone number that appears once in a 50-billion-token pretraining mix is a different risk from a customer name repeated in 400 fine-tuning examples.

Illustrative example: invented to show structure; it does not describe an available dataset.

SituationTypical decisionWhat to record
Data only in RAG index or eval set, no trainingPurge index, rebuild from corrected deliveryPurge timestamp, rebuild manifest
Low prevalence, direct identifiers, pretraining mix, no duplicationFilter going forward; run targeted extraction probes on the modelProbe prompts, outputs, pass criteria
PII repeated across many fine-tuning examplesRetrain adapter or fine-tune from the last clean checkpointCheckpoint lineage, before/after probe results
Sensitive category (health, children, biometrics) in training dataEscalate to counsel; default toward retrainingLegal assessment, retention decision
Model already released externallyCounsel-led review of disclosure and output filteringRelease scope, mitigation timeline

Retraining from the last clean checkpoint is usually cheaper than a full run, and architectures such as SISA sharding exist precisely to make removal cheaper by limiting which shards must be retrained [6]. Output-side filters reduce exposure but do not remove data from the weights, so document them as mitigation rather than remediation. For EU deployments, whether a model trained on personal data can itself be considered anonymous is a separate test; see EDPB Opinion 28/2024 for data buyers.

Documentation regulators and auditors will expect

Write the record as if a regulator or customer auditor will read it, because for high-risk systems they may. EU AI Act Article 10 requires data governance and management practices for training, validation and testing data sets of high-risk AI systems [11]; as of October 2026, the high-risk application dates reportedly moved to 2 December 2027 (Annex III) and 2 August 2028 (Annex I) under Regulation (EU) 2026/1744, which also amends Article 10 [12]. A clean incident file is strong evidence of governance in practice.

Keep, at minimum:

  • the quarantine manifest and access logs;
  • the sample design, detectors used and hit counts by category;
  • the supplier notice, response and corrected-delivery manifest;
  • the model decision with its rationale and probe results;
  • deletion or retention confirmations for quarantined copies.

Whether the discovery is itself a reportable personal data breach under GDPR, a state breach statute or HIPAA is a question for counsel, and it depends on facts such as whether anyone outside the authorized team accessed the data.

Contract hooks that make the next incident easier

The cheapest time to plan residual-PII remedies is before signature. Buyers who negotiate these clauses rarely need to improvise:

  • Defined de-identification method and the identifiers in scope, by field.
  • Acceptance testing window with a stated sampling plan and threshold.
  • Notice duty in both directions, with a named contact and channel.
  • Replacement delivery as the primary remedy, with the corrected manifest.
  • Deletion and certification for quarantined files.
  • Model-treatment language stating what the buyer must do with models already trained.

Because no redaction is perfect, these clauses allocate a known risk instead of pretending it away [10]. When you request a dataset through SourceX, the license defines records, uses, term and delivery, so raise these points while terms are being agreed. For relationship-level terms, see sourcing data directly from operating companies, and return to the privacy cluster hub or the AI data hub for related guides.

Requesting de-identified datasets with a documented method

SourceX sources operational datasets from US companies on request and manages licensing; personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Every dataset is rights-reviewed and delivered under a license through private, access-controlled workflows after supplier approval. Describe the data you need at SourceX for AI data buyers.

Frequently asked questions

Should we delete the dataset immediately?

No. Quarantine it instead. Deletion removes the evidence needed to scope the problem, support the supplier notice and prove what was remediated; delete only after the corrected delivery is accepted and your record is complete.

Does finding one email address make the whole dataset personal data?

Not necessarily, but it makes the label untrustworthy until scoped. Prevalence, category and how reasonably the data can be linked to a person decide the answer, judged against the applicable legal test [8][9].

Can we keep a model trained on the affected data?

Sometimes. Low-prevalence direct identifiers in a large, deduplicated pretraining mix, with clean extraction probes, often support keeping the model with filters; repeated PII in fine-tuning data or sensitive categories usually push toward retraining [3][5].

Sources

  1. Microsoft (microsoft/presidio project), indexed on pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  2. Grepture, "Why Presidio is not enough for PII redaction". https://grepture.com/blog/presidio-not-enough-pii-redaction
  3. arXiv (Li et al., 2024), "LLM-PBE: Assessing Data Privacy in Large Language Models" (2024). https://arxiv.org/pdf/2408.12787
  4. arXiv (Carlini et al., USENIX Security 2021), "Extracting Training Data from Large Language Models" (2020). https://arxiv.org/pdf/2012.07805
  5. arXiv (Lee et al., 2021), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  6. arXiv (Bourtoule et al., IEEE S&P 2021), "Machine Unlearning" (2019). https://arxiv.org/abs/1912.03817v2
  7. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  8. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  9. UK Information Commissioner's Office, "How do we ensure anonymisation is effective?". https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
  10. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  11. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  12. European Parliament and Council of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data