Skip to content

Privacy, de-identification and sensitive data

Auditing a fine-tuned model for leakage: canaries, membership inference and extraction tests

Quick answer

A pre-release leakage audit for a fine-tuned model runs three complementary tests: canary exposure (did the model memorize secrets you planted on purpose), membership inference (can an attacker tell which licensed records were in training), and extraction (can prompting or decoding reproduce training text, fully or partially). Run them against held-out non-members drawn from the same distribution, use stronger-than-greedy decoding, and tie each pass/fail gate to what your data license and privacy review actually require.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why a fine-tuned model needs its own leakage audit

Fine-tuning on a small, sensitive, licensed corpus concentrates memorization risk, so a base model's privacy evaluation does not carry over. Many epochs over a few thousand support tickets, contracts or clinical notes give each record far more gradient weight than a single pass over a web-scale pretraining set. Research on fine-tuned models reports that fine-tuning can amplify recovery of personal information [7], and the foundational extraction work showed verbatim training sequences, including personal details, coming back out of GPT-2 [4].

Upstream controls reduce the risk but do not remove it. Redaction misses things (see PII redaction for LLM training data), and residual identifiers survive in free-text fields such as notes, email bodies and attachment OCR. The risk side is covered in training-data extraction and memorization risk; this page is the test plan an ML security reviewer signs before weights or an endpoint ship. For background, see what model memorization is.

Canary insertion: measuring memorization you control

Canaries give you ground truth: you plant unique synthetic secrets in the fine-tuning set and measure how strongly the trained model prefers them over random alternatives of the same format. Work on differentially private fine-tuning of language models uses inserted canary sequences in this way to measure memorization directly, without needing to know which real records are at risk [1].

Exposure compares a canary's model likelihood against a reference set of candidates that share its template. With n candidates, exposure is roughly log2(n) minus log2 of the canary's rank, so a canary ranked first scores the maximum and an unmemorized canary scores close to the bottom of the range. Exposure is not normalized, so report the candidate-space size alongside every number. Because exposure is a rank statistic, it can also be read as a membership signal, which helps when you explain results to a reviewer who thinks in differential-privacy terms.

Design rules that keep canaries informative:

  • Match the real formats. If the licensed data contains account numbers, MRNs or claim IDs, generate canaries in those exact formats inside realistic carrier sentences (for example, a ticket body or a note field), not as isolated strings.
  • Vary repetition. Insert some canaries once and some 2, 4, 8 or 16 times. Real duplicates exist in operational data (email threads quoted in replies, templated contract clauses), and the repetition curve shows where memorization starts.
  • Keep a held-out twin set. Generate canaries from the same template that never enter training; they are your non-member baseline.
  • Log everything. Record the canary ID, template, secret, insertion count, record ID it was injected into and the training run hash, so the audit is reproducible.

Canaries measure the training recipe, not the specific real records. A recipe that drives high canary exposure at four repeats tells you any real identifier repeated four times is at risk, which you then check with the tests below.

Membership inference: what attacks can and cannot tell you

Membership inference asks whether a specific record was in the training set, and for a fine-tuned model it is the most direct test of whether licensed records are distinguishable. Common signals range from simple loss thresholds to Min-K% Prob, which scores a text by its lowest-probability tokens [3], up to reference-model attacks such as LiRA that compare the target model's behavior against models trained with and without the candidate. A combined toolkit such as LLM-PBE packages membership inference alongside extraction tests [2].

Interpret results carefully, because evaluation design dominates the outcome. Published evaluations on large pretrained models have found membership attacks close to random guessing once members and non-members are drawn from the same distribution, while apparent successes often trace to distribution shift, such as non-members from a different time range. In an audit of a licensed corpus, the same trap appears if your non-members are tickets from a later quarter or contracts from a different business unit.

Practical rules for a defensible membership audit:

  • Draw non-members from the same delivery. Hold out a random slice of the licensed dataset before fine-tuning so members and non-members share source system, date range and preparation pipeline.
  • Report TPR at low FPR. Area under the ROC curve hides the cases that matter; report true-positive rate at 1% and 0.1% false-positive rate.
  • Use a calibrated attack where budget allows. Reference-model attacks are expensive but give a stronger signal than raw loss; loss and Min-K% are cheap first-pass screens.
  • Stratify by record type. Rare, long or highly repeated records leak first, so report results for short tickets, long documents and duplicated templates separately.

A near-random result on a well-constructed split is evidence, not proof, of low leakage. A strong result on a poorly constructed split is often an artifact.

Extraction tests: going beyond greedy decoding

Extraction tests ask whether the model will actually emit training text, and they must use more than greedy decoding or they will undercount partial memorization. Extractable memorization is training data an adversary can recover by querying the model without prior knowledge of the training set [5]; the original attacks generated many samples and ranked them by likelihood ratios to surface memorized text [4]. Recent work published at PoPETs 2026 reports that decoding strategies beyond greedy recover partially memorized data that greedy extraction misses [6].

Build the extraction suite in three layers:

  1. Prefix completion (discoverable memorization). Feed the first k tokens of real training records (k of 32, 64 and 128 is a reasonable spread) and measure exact-match and near-match suffixes.
  2. Untargeted sampling. Generate a large pool with temperature and top-k or nucleus sampling from empty or generic prompts, then scan outputs with the same PII detector used on the dataset and match against the training corpus.
  3. Targeted probing. Prompt with partial identifiers or entity context ("the customer at account ending...") and with adversarial instructions to repeat or continue documents.

Score partial matches, not just exact ones. Use normalized edit distance or longest common substring against the training records, and flag any output that reproduces a direct identifier (name plus account number, email, phone) even when the surrounding text differs. Running both test families from one toolkit [2] helps standardize reports across model versions.

The audit plan and release gates

A release gate converts test results into a go or no-go decision tied to the license and the privacy basis for the data. The thresholds below are placeholders: set yours with counsel, the data supplier's restrictions and your deployment exposure (internal tool, customer-facing API or open weights) in mind.

Illustrative example: invented to show structure; it does not describe an available dataset.

TestSetupMetricExample gate (internal API)Example gate (open weights)
Canary exposure200 canaries in ticket and note formats, repeats 1/4/16, 200 held-out twinsExposure per repeat level, candidate space 10^9Median exposure at 4 repeats below an agreed boundNo canary at any repeat ranked first
Membership inference2,000 members vs 2,000 same-delivery held-out records; loss, Min-K%, LiRATPR at 0.1% FPRNo better than an agreed multiple of baselineIndistinguishable from baseline within confidence interval
Prefix extraction1,000 records, prefixes of 32/64/128 tokens, greedy and beamExact and near-match suffix rateZero direct identifiers in suffixesZero direct identifiers; near-match rate reported
Sampled extraction100,000 sampled generations, PII scan plus corpus matchUnique training identifiers recoveredZero verified training identifiersZero verified training identifiers
Targeted probingRed-team prompts with partial identifiersSuccessful recoveriesZeroZero

Pair the table with an audit record so results stay traceable across retraining:

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "audit_id": "leak-audit-2026-10-ft-support-v3",
  "model": {"base": "base-7b", "adapter": "lora-r16", "train_run_sha": "9f1c..."},
  "dataset_ref": {"license_id": "LIC-0420", "delivery_id": "DLV-07", "deid_method": "replace-synthetic"},
  "canaries": {"inserted": 200, "held_out": 200, "repeats": [1, 4, 16], "candidate_space": 1000000000},
  "mia": {"members": 2000, "non_members": 2000, "split": "random-holdout-same-delivery", "attacks": ["loss", "min_k_20", "lira_8ref"], "tpr_at_0_1_fpr": null},
  "extraction": {"prefix_lengths": [32, 64, 128], "sampled_generations": 100000, "decoding": ["greedy", "beam4", "top_p_0.95", "membership_decoding"], "identifiers_recovered": null},
  "gate_decision": "pending",
  "reviewers": ["ml-security", "privacy-counsel"]
}

Map the work to the NIST AI RMF if your governance program uses it: these tests sit under MEASURE, and the release gate and remediation steps under MANAGE [9]. If the audit fails, the usual remedies are deduplication of repeated records, fewer epochs, stronger redaction of the fields that leaked, or training with differential privacy (see DP-SGD for LLM fine-tuning), followed by a full re-run of the suite.

How audit results connect to your data license and privacy position

Audit results matter because several obligations depend on whether the model itself still carries the licensed records. A data license may limit use to training and forbid redistribution of records; a model that regurgitates them can put you in breach through outputs alone. Before you plan to ship weights, read whether you can release model weights trained on licensed personal data.

In the EU, EDPB Opinion 28/2024 is the reference on when a model trained on personal data can be considered anonymous, and summaries highlight resistance to attacks that extract personal data as part of that assessment [8]. Membership inference and extraction results are the evidence a supervisory authority would expect to see; the buyer-side reading is in our page on EDPB Opinion 28/2024 and model anonymity. As of October 2026, no statute or regulator has published numeric pass thresholds for these tests, so document why you chose yours.

Leakage audits also feed back into procurement. Ask suppliers for the de-identification method, field-level redaction coverage and duplicate statistics in a de-identification evidence package, and run a residual PII audit on delivery before training. Every re-tune on new data needs a fresh audit, which belongs in your fine-tuning refresh cadence.

Where licensed data fits in the audit

Licensed data that arrives with a recorded preparation method makes the audit faster to design and easier to defend. SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents, and finance and legal workflows; every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, but no method is perfect, which is exactly why a model-side leakage audit still matters. Health records require HIPAA de-identification by Safe Harbor or Expert Determination. Describe the data you need through the SourceX buyer request.

For the full cluster, start at the de-identified data for AI training hub or the AI data buyer's guides.

Sourcing sensitive training data you can audit before release

SourceX looks for US businesses that hold the data you describe, prepares diligence materials on source, rights, preparation and allowed use for each dataset, and delivers only after an executed agreement and supplier approval through private, access-controlled workflows. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. Describe your dataset requirements to SourceX.

Sources

  1. Trinity College Dublin, School of Computer Science and Statistics, "Differentially-Private Fine-Tuning of a Small Language Neural Net Model (TCD SCSS project page)". https://projects.scss.tcd.ie/?p=4235
  2. arXiv, "LLM-PBE: Assessing Data Privacy in Large Language Models" (2024). https://arxiv.org/pdf/2408.12787
  3. arXiv (Shi et al.; ICLR 2024), "Detecting Pretraining Data from Large Language Models" (2023). https://arxiv.org/pdf/2310.16789
  4. arXiv (Carlini et al.; USENIX Security 2021), "Extracting Training Data from Large Language Models" (2020). https://arxiv.org/pdf/2012.07805
  5. arXiv (Nasr, Carlini et al.), "Scalable Extraction of Training Data from (Production) Language Models" (2023). https://arxiv.org/abs/2311.17035v1
  6. Proceedings on Privacy Enhancing Technologies, "PoPETs 2026 paper 0139 (partial memorization beyond greedy extraction)" (2026). https://petsymposium.org/popets/2026/popets-2026-0139.pdf
  7. arXiv, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
  8. CMS (summary of European Data Protection Board Opinion 28/2024), "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  9. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data