Skip to content

Privacy, de-identification and sensitive data

Training-data extraction and memorization: privacy risk when training on licensed sensitive data

Quick answer

Language models can reproduce parts of their training data verbatim, and black-box extraction attacks have recovered names, phone numbers, email addresses and other personal data from deployed models, including aligned chat models [1][2]. When you train on licensed sensitive text, treat extraction as a residual risk to manage: de-identify and deduplicate before training, consider differential privacy for the most sensitive fine-tunes, filter outputs, and run extraction tests before every release.

By SourceX Editorial · Updated

This page is part of the privacy and de-identification guide for AI training data. It covers the buyer's risk analysis; the plain-language definition is in what model memorization is.

How extraction attacks pull personal data out of a trained model

An extraction attack elicits large volumes of model output, ranks it by signals that suggest memorization, and keeps the candidates that match real training text. Carlini et al. did this against GPT-2 with only black-box queries: they sampled text, ranked it with metrics such as the ratio of model perplexity to zlib compression entropy, and recovered verbatim sequences that included personal contact details [1]. Some extracted sequences appeared in only one training document, and larger models memorized more [1].

Nasr et al. formalized extractable memorization as training data an adversary can efficiently recover by querying a model without prior knowledge of the training set [2]. Their attacks recovered thousands of training examples from aligned production chat models, including a divergence attack that asks the model to repeat one word until it drifts out of chat behavior and starts emitting training text [2]. Alignment is therefore not a privacy control: a 2025 study found that aligned models still emit memorized text after spikes in token-level entropy [5].

Four terms cover most governance discussions:

TermWhat it meansWhy a buyer cares
Discoverable memorizationThe model completes a training record when prompted with that record's true prefixAn audit upper bound; needs access to the training data
Extractable memorizationAn adversary recovers training data without knowing it in advance [2]The realistic external threat to a deployed model
Partial or approximate memorizationThe model reproduces fragments or near-verbatim variantsUnder-detected by greedy-decoding tests [4]
Membership inferenceAn attacker infers whether a specific record was in trainingReveals that a person was in a sensitive dataset even without text output

Why fine-tuning on licensed sensitive data concentrates the risk

Fine-tuning gives a small, sensitive corpus far more weight per record than the same text would carry inside a web-scale pre-training mix, so most buyer-side extraction risk sits in supervised fine-tuning and continued pre-training on licensed data. The Janus Interface study showed that fine-tuning on a small amount of PII can amplify a model's leakage of personal information [3]. Smaller open models are not exempt: a 2025 paper reported extracting PII from a Llama 3 model through model-inversion attacks [7].

Third-party components add attack surface. The "Teach LLMs to Phish" work showed that a few benign-looking sentences inserted into training data can induce a model to memorize and later reveal secrets such as credit card numbers [6]. A 2025 paper showed that a backdoored open-weight base model can let its creator extract the data later used to fine-tune it [11]. If you fine-tune a third-party base model on licensed records, record the base model's provenance next to the dataset's.

Operational business data has predictable memorization hot spots:

  • Email signatures, legal disclaimers and ticket templates repeated across thousands of records, often carrying a name, direct line and job title.
  • Account numbers, order IDs, invoice numbers and claim numbers embedded in free text instead of structured fields.
  • API keys, connection strings and bearer tokens pasted into engineering tickets, incident postmortems and chat logs.
  • Rare events, such as a single lawsuit, outage or medical episode, that identify a person even after names are removed (see indirect identifiers in business text).
  • Employee chat and meeting transcripts in which people discuss customers and colleagues freely (see employee communications in training data).

Repetition is the memorization driver you can control

Duplicated text is memorized more readily, so deduplication is the cheapest model-side privacy control a buyer can require. Lee et al. removed exact repeated substrings with a suffix-array method and near-duplicate documents with MinHash, and the resulting models emitted memorized text ten times less frequently [8]. In operational corpora the duplicates are rarely whole documents; they are quoted email threads, auto-generated notifications and copied boilerplate, so substring-level deduplication matters more than document-level hashing.

Deduplication does not remove risk from unique records, because sequences seen once can still be extracted [1]. It also interacts with de-identification: a pipeline that replaces each person with the same surrogate everywhere creates a repeated string that may be memorized. That is harmless only if the surrogate mapping table never leaves the supplier, so ask how surrogates are generated and where any mapping key is held.

The mitigation stack to require, layer by layer

No single control removes extraction risk, so require controls at the data, training, output and release layers, with evidence for each. The table works as a review checklist; the last column is where residual risk usually hides.

LayerControlEvidence to requestWhat it misses
Source dataDirect-identifier removal or surrogate replacement; HIPAA Safe Harbor or Expert Determination for PHI [9]Method document, detector versions, sample-check resultsIndirect identifiers and rare events in free text
Source dataSecrets scanning of engineering records (known key prefixes plus high-entropy strings)Scan configuration and remediation logNovel credential formats
CorpusExact-substring and near-duplicate removal [8]Method, thresholds, before and after token countsUnique records memorized from one occurrence [1]
TrainingDP-SGD with stated epsilon, delta and privacy unitPrivacy accountant output; unit (example, document or person)Utility cost; weak bounds at large epsilon
TrainingCanary insertionCanary design and exposure resultsOnly measures what was planted
OutputPII detectors and n-gram match filters against the training setFilter recall test resultsParaphrased and partial reproduction [4]
ReleaseExtraction, divergence and PII-targeted prompt tests [2][10]Test plan, decoding settings, results, sign-offGreedy-only decoding understates exposure [4]

De-identification quality determines how much personal data reaches the weights at all. Health records need HIPAA de-identification by Safe Harbor, which removes 18 listed identifiers, or by Expert Determination [9]; for clinical free text, see de-identifying clinical notes for LLM training. For business text, compare detectors in LLM-based PII redaction vs NER and regex and request the full de-identification evidence package.

When SourceX sources operational text, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked. No method is perfect, so the model-side controls above still apply. You can describe the sensitive data you need and how you plan to train on it so the privacy review starts from the right facts.

Differential privacy is the only control here that bounds what one record can contribute, but the bound depends on the privacy unit and epsilon. If one customer contributes hundreds of tickets and the privacy unit is a single ticket, the protection for that customer is much weaker than the headline epsilon suggests. Implementation and utility trade-offs are covered in differential privacy for LLM fine-tuning, and differentially private synthetic text is an option when raw records should never reach your training cluster.

Making extraction testing a release gate

Extraction testing should be a release gate with written decoding settings and pass criteria, because a test that uses only greedy decoding will understate exposure. PoPETs 2026 research found that most partially memorized data goes unnoticed by greedy-decoding extraction and proposed membership decoding to surface it [4]. Run prefix-suffix tests on sampled training records, divergence and repeated-token prompts [2], PII-targeted prompts that supply a name and ask for contact details, and membership inference on canaries; the LLM-PBE toolkit packages several of these attacks and defenses [10].

Method detail for canaries, exposure metrics and membership inference is in auditing a fine-tuned model for leakage. The governance record below captures what was tested, against which threshold, and who signed off.

Illustrative example: invented to show structure; it does not describe an available dataset.

risk_id: MEM-014
dataset_ref: licensed-support-tickets-v3 (de-identified, surrogates)
use: supervised fine-tuning; internal assistant; API-only release
release_mode: api_only            # open_weights would require white-box tests
memorization_drivers:
  - repeated signatures and disclaimers in quoted email threads
  - order and account numbers in free-text ticket bodies
  - credentials pasted into engineering escalations
data_controls:
  deidentification: ner_plus_regex, surrogate_replacement, method_doc_v2
  sample_check: residual identifiers logged and remediated
  secrets_scan: known_key_prefixes + high_entropy_strings
  dedup: exact_substring_50_tokens + minhash_near_dup
training_controls:
  dp_sgd: not_used   # rationale recorded; revisit if PHI is added
  canaries: 200 synthetic records with random 16-digit strings
release_gate:
  tests: [prefix_suffix, divergence_prompts, pii_targeted_prompts, membership_inference]
  decoding: [greedy, sampled_temperature_1.0]
  pass_criteria: no canary recovered; any verbatim match >= 50 tokens triaged
  output_filter: pii_detector + training_ngram_match
owner: ai_governance_lead
review_trigger: each release and any retraining

If the model will ship as open weights, assume an attacker has logits, unlimited queries and the ability to fine-tune, and widen the tests accordingly. An API-only model with output filters and rate limits exposes less, but it still faces the divergence-style attacks shown against production chat models [2].

Where memorization meets regulation and the license

Regulators increasingly treat a model's ability to regurgitate personal data as a property of the model itself, not only of the training set. As of October 2026, EDPB Opinion 28/2024 remains the main EU reference: a model trained on personal data is not automatically anonymous, and anonymity requires that the likelihood of extracting personal data, directly or through queries, is insignificant [13]. The proposed GDPR "Digital Omnibus" changes are not law. In the US, the voluntary NIST AI 600-1 profile lists data privacy among the 12 risks it identifies for generative AI [12].

In the license and diligence file, make memorization an explicit, allocated risk:

  • Allowed uses that distinguish pre-training, fine-tuning, evaluation and retrieval, since each has a different exposure profile.
  • The de-identification method, residual-risk statement and sample-check results as diligence exhibits.
  • A prohibition on attempts to re-identify individuals, including through model outputs.
  • Who is notified, and what happens to affected checkpoints, if a test or user report surfaces a supplier record.
  • The planned release mode (internal, API or open weights) as a stated assumption of the risk assessment.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Sourcing sensitive training text with extraction risk in mind

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and finance and legal workflows, and manages the licensing process. Every dataset is rights-reviewed, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded, and nothing is contracted until the supplying company agrees. Describe the data you need, not the businesses that might hold it, at sourcex.si/buyers.

Frequently asked questions

Does de-identifying the data make extraction harmless?

It reduces the harm: if names, emails, phone numbers and account numbers were removed or replaced before training, an extracted sequence exposes less. Residual identifiers, rare events and combinations of quasi-identifiers can still point to a person, and language models are getting better at linking them (see LLM-assisted re-identification). Treat de-identification as shrinking the payload and extraction testing as measuring what remains.

Is retrieval-augmented generation safer than fine-tuning on the same records?

Retrieval moves the risk rather than removing it. A retriever can return an entire source passage to anyone who asks the right question, so access control and retrieval-time filtering replace extraction testing as the main controls; see personal data in RAG corpora.

What should we tell a supplier who asks whether their data could leak?

Share the planned controls concretely: deduplication method, the decision on differential privacy, output filters, release mode and the extraction test gate with its pass criteria. Suppliers weigh the same research summarized in can AI models memorize and leak my data?, and specific evidence shortens that conversation more than general assurances do.

Sources

  1. Carlini et al., USENIX Security 2021, "Extracting Training Data from Large Language Models" (2021). https://arxiv.org/pdf/2012.07805
  2. Nasr et al., ICLR 2025, "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
  3. arXiv, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
  4. PoPETs 2026, "Proceedings on Privacy Enhancing Technologies 2026, article popets-2026-0139" (2026). https://petsymposium.org/popets/2026/popets-2026-0139.pdf
  5. arXiv, "arXiv:2511.05518v1 (2025 preprint on extraction after entropy spikes)" (2025). https://arxiv.org/abs/2511.05518v1
  6. ICLR 2024, "Teach LLMs to Phish: Stealing Private Information from Language Models" (2024). https://proceedings.iclr.cc/paper_files/paper/2024/hash/f206871468dc89cb20fefbc75f2de861-Abstract-Conference.html
  7. arXiv, "Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models" (2025). https://arxiv.org/pdf/2507.04478
  8. Lee et al., arXiv / ACL 2022, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
  9. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  10. arXiv, "LLM-PBE: Assessing Data Privacy in Large Language Models" (2024). https://arxiv.org/pdf/2408.12787
  11. arXiv, "Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!" (2025). https://arxiv.org/pdf/2505.15656
  12. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
  13. CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models". https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data