Skip to content

Data quality, coverage and contamination

Scanning a Training Corpus for PII Before Fine-Tuning: Inputs, Targets and Detector Choice

Quick answer

Run an automated PII scan over every record in the fine-tuning corpus, not a sample, as a blocking gate before the training job starts. Scan prompts, system messages, retrieved context and tool outputs as well as targets, because the model conditions on inputs and can still surface them [4]. Include credentials and confidential business data in scope [1], choose detectors by measuring per-entity recall on a labeled slice of your own data [2], and log a disposition for every finding.

By SourceX Editorial · Updated

This page covers the automated, full-corpus gate inside your own pipeline. The sample-based audit of a supplier delivery, with sampling math and acceptance thresholds, belongs on the acceptance sampling guide for dataset deliveries and the PII redaction guide for LLM training data. De-identification methods themselves live in the privacy cluster.

Why fine-tuning raises the cost of residual PII

Fine-tuning concentrates a small corpus into many gradient steps, so a residual phone number or account ID in that corpus is more likely to be memorized than the same string in a web-scale pretraining mix. Research on training-data extraction showed that language models can emit verbatim training sequences, including personal details, when prompted the right way [5]. The Janus work reports that fine-tuning can amplify recovery of personal information that a base model had only weakly absorbed [3].

Two practical consequences follow. First, one missed identifier repeated across 400 templated support tickets is a larger risk than 400 different identifiers seen once, so duplication and PII interact; run near-duplicate detection before the PII gate and scan the deduplicated set. Second, de-identification upstream does not remove the need for a downstream scan: NIST notes that traditional de-identification has inherent limitations [7], and supplier redaction is one control, not a guarantee. For the underlying mechanism, see what model memorization is.

Scan inputs, prompts and context, not just targets

Every text field the model sees during training is in scope, including fields excluded from the loss. Teams often scan only the assistant turn or the completion field because that is what the model learns to produce, but input text still creates exposure [4]: it sits in cached tokenized shards and data-loader logs, the model conditions on it, and in continued pre-training it is a target anyway.

Map the scan to the actual record schema rather than to a file:

  • SFT chat format (JSONL messages): system, every user turn, every assistant turn, and tool / function-call arguments and results. Tool outputs from CRM or ticketing lookups are a frequent leak path.
  • Prompt/completion pairs: both fields, plus any metadata object that a data loader might template into the prompt.
  • RAG-style or RAFT-style records: the retrieved passages and distractor documents, which often come from raw document stores that were never redacted.
  • Continued pre-training: the full document text, document titles, file paths and headers (email From: and To: lines survive many PDF and EML converters).
  • RAG indexing: chunk text and chunk metadata such as author, owner_email and source path, because retrieval returns them verbatim at inference time.

Placeholders matter too. If the supplier replaced names with tokens such as [NAME_1], scan for the replacement-token grammar so that broken or partial tokens ([NAME_1] Gonzal) surface as findings rather than passing.

What to scan for: entity scope beyond names and emails

The scan scope should cover direct identifiers, quasi-identifiers, credentials and confidential business data; one published sanitisation evaluation scores each fine-tuning sample for personal identifiers, credentials, health data and confidential business content [1]. Treat these as four detector families with different techniques.

FamilyExamplesDetector techniqueTypical failure mode
Direct identifiersNames, emails, phone numbers, postal addresses, SSNs, account and policy numbersNER model plus regex with checksum validation (Luhn for card numbers, SSN format rules such as no 000, 666 or 9xx area)Names in lowercase chat text, non-US phone formats, account IDs that look like order IDs
Quasi-identifiersDates of birth, ZIP codes, job title plus employer, rare diagnosesRegex and NER, then combination rules at record levelEach field passes alone; the combination identifies a person
Credentials and secretsAPI keys, bearer tokens, passwords, connection strings, private keysPattern and entropy scanners of the kind used in gitleaks or TruffleHogKeys pasted into tickets or logs; see the secrets and credentials scanning guide
Confidential business dataCustomer lists, contract values, internal hostnames, unreleased product namesCustom dictionaries and classifiers trained on your own labelsNo off-the-shelf model knows your internal code names

Health records need extra care. If any part of the corpus came from a HIPAA-covered source, confirm it was de-identified under Safe Harbor or Expert Determination before it reached your pipeline, and treat a scan hit on a health identifier as a provenance question, not only a redaction task; the HIPAA training data guide covers what to request.

Choosing a detector: benchmark on a labeled slice of your own data

Pick detectors by measured per-entity recall on your own records, because published scores come from other text distributions. One 2025 study on educational text reported PII detection recall of about 0.96 and 0.99 on two datasets, and roughly a threefold precision gain after fine-tuning the detector on in-domain data [2]. A separate evaluation of LLM-based redaction found both real capability and measurable risks [6]. Neither result tells you how a detector behaves on your support transcripts or engineering logs.

A workable benchmark protocol:

  1. Draw a stratified sample of 1,000 to 3,000 records across sources, record types and languages, oversampling free-text fields and tool outputs.
  2. Label spans by entity type with two annotators and adjudicate disagreements; measure agreement with the methods in the inter-annotator agreement guide.
  3. Run each candidate: a rules-and-NER framework such as Microsoft Presidio, a cloud DLP service, a fine-tuned token classifier, and an LLM-based redactor if latency allows.
  4. Score per entity type, not overall. Report recall and precision for PERSON, EMAIL, PHONE, ACCOUNT_ID, API_KEY and each custom type separately; an aggregate F1 of 0.95 can hide recall of 0.6 on account numbers.
  5. Prefer recall at the gate, then control precision with an allowlist (product names, public company addresses, test fixtures) so false positives do not drop useful records.
  6. Ensemble where families differ: regex with validation for structured IDs, NER for names, entropy scanners for secrets.

Keep the labeled slice as a frozen regression set. Re-run it whenever a detector version, model weight or allowlist changes, and fail the pipeline if per-entity recall drops.

Dispositions: drop, re-redact or return to the supplier

Every finding needs a recorded disposition, and the choice depends on the entity family, how many records are affected and whether you hold the rights to modify the data. Silent fixes break auditability; a finding without a logged action is indistinguishable from a missed one.

  • Drop the record when a credential, health identifier or confidential business value appears, or when the record's value depends on the identifier itself.
  • Re-redact in place when the identifier is incidental (a signature line, a greeting) and your license permits modification; use the same placeholder grammar as the supplier so the corpus stays consistent, and check for over-redaction afterwards.
  • Return to the supplier when findings cluster in one source, field or date range, which points to a broken upstream redaction step rather than random misses. This is also when the sample-based acceptance audit should be re-run.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "scan_run_id": "pii-gate-2026-10-09-r3",
  "corpus_version": "sft-support-v7@sha256:9c1e...",
  "detectors": [
    {"name": "presidio-analyzer", "version": "pinned", "entities": ["PERSON", "EMAIL_ADDRESS", "PHONE_NUMBER"]},
    {"name": "regex-validated-ids", "version": "2026.10", "entities": ["ACCOUNT_ID", "US_SSN"]},
    {"name": "secret-entropy-scanner", "version": "pinned", "entities": ["API_KEY", "BEARER_TOKEN"]}
  ],
  "finding": {
    "record_id": "rec_000418233",
    "field_path": "messages[3].content",
    "role": "tool",
    "entity_type": "ACCOUNT_ID",
    "span": [112, 124],
    "detector": "regex-validated-ids",
    "score": 0.99,
    "disposition": "re_redact",
    "replacement": "[ACCOUNT_1]",
    "actor": "pipeline:auto",
    "reviewed_by": "privacy-eng",
    "logged_at": "2026-10-09T14:02:11Z"
  },
  "gate_result": {"records_scanned": 412880, "findings": 37, "dropped": 9, "re_redacted": 26, "returned_to_supplier": 2, "blocking": false}
}

Store raw matched values only in an access-controlled findings store with short retention, or store a salted hash and offsets instead, so the scan log does not become a new PII repository.

Wiring the gate into the training pipeline

The gate belongs after deduplication and format normalization and before tokenization, and it should block the training job on any unresolved finding. Tokenized shards are hard to inspect and expensive to rebuild, so scanning text before sharding avoids re-tokenizing the whole corpus for a handful of records.

A minimal gate checklist:

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckPass condition
Coverage100% of records and every text-bearing field path scanned, including tool and system roles
Detector regressionPer-entity recall on the frozen labeled slice at or above the agreed floor
CredentialsZero unresolved secret findings
DispositionsEvery finding has a logged action, actor and timestamp
Supplier feedbackClustered findings sent back with record IDs and field paths
LineageScan run ID and corpus hash attached to the training run's metadata

Rerun the gate on every corpus change, including new supplier drops and synthetic augmentations generated from the corpus, since a generator can copy identifiers from its seed records. For RAG, run the same scan at indexing time and again when documents are re-ingested. Record the lineage alongside your provenance audit so a later deletion request can be traced to the affected runs.

How this connects to buying data

A strong in-house gate also tells you what to ask a supplier for: the redaction method, the entity list it covered, and how it was checked. When you source operational data through SourceX for AI data buyers, personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, but no method is perfect, which is why your own full-corpus scan still matters. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and diligence materials on source, rights and preparation are prepared per dataset. See how personal details are removed and the glossary entry on PII detection for terminology, and the quality cluster hub for related checks.

Request de-identified data for fine-tuning

SourceX sources operational datasets such as support and sales histories, engineering records and documents from US companies on request, with every release approved by the supplying company and personal details removed or replaced before delivery. Describe the data you need, including field paths and entity types you will scan for, and SourceX will look for businesses that hold it; a request does not guarantee a match. Start a buyer request.

Frequently asked questions

Should I scan before or after deduplication?

After. Deduplication shrinks the corpus and the review queue, and it removes the repeated copies that make memorization of a missed identifier more likely. Then scan every remaining record.

Is a sample-based audit enough if the supplier already redacted the data?

No, not for the training gate. A sample audit estimates the residual rate and supports acceptance of a delivery, while the full-corpus scan finds the specific records to drop or fix before training; use both.

Can an LLM replace regex and NER detectors?

It can add recall on unstructured mentions, but it should be benchmarked per entity type like any other detector [6], and it is usually slower and costlier per record. Most teams keep validated regex for structured IDs and entropy scanning for secrets.

Sources

  1. LatticeFlow AI Atlas, "Training Data Sanitisation (evaluation)". https://atlas.latticeflow.ai/evaluation/training_data_sanitisation
  2. arXiv, "Enhancing the De-identification of Personally Identifiable Information in Educational Data" (2025). https://arxiv.org/html/2501.09765v1
  3. arXiv, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
  4. Promptfoo LM Security DB, "LLM input PII leakage". https://promptfoo.dev/lm-security-db/vuln/llm-input-pii-leakage-75a0bd54/
  5. arXiv (Carlini et al., USENIX Security 2021), "Extracting Training Data from Large Language Models" (2020). https://arxiv.org/pdf/2012.07805
  6. arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
  7. NIST, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data