Skip to content

Privacy and preparation

Can you use an LLM API to redact PII without exposing the data?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

You can use an LLM API to redact PII only if the provider contractually commits not to train on or retain your inputs, a data processing agreement covers the work, and your customer contracts allow it. Otherwise run the model in your own environment. Either way, the model reads the raw data first, so treat redaction itself as a disclosure.

Key takeaways

  • The redaction step reads every unredacted record, so the model provider becomes a processor of your most sensitive data.
  • A hosted API is defensible only with no-training and no-retention terms, a DPA and a check of your customer contracts.
  • A model running in your own cloud account or on your own hardware avoids an extra external disclosure.
  • Have the model return spans to remove, apply them with ordinary code, and check results with a second detector and human review.

Why sending records to an LLM API is itself a disclosure#

Sending records to an LLM API discloses them to the provider before anything is redacted. The model has to read the names, emails, account numbers and credentials in order to find them, so the redaction step handles the rawest version of your data.

That makes the provider a processor of the very data you are trying to protect. Your customer agreements may limit which subprocessors can see their data, your privacy notice may describe where data goes, and any credentials in the text are exposed to one more system. A remote redaction service does not remove exposure; it moves it.

Consumer chat apps are the clearest no. Their terms can allow retention or use of inputs, and pasting customer records into them typically falls outside anything your customers agreed to.

Hosted API, your own cloud or a local model: how the options compare#

The deployment options differ mainly in where the raw text travels and which document governs it. Pick the option first, then the model; a strong model in the wrong place is still the wrong choice.

The cloud-hosted option still involves a provider, the cloud company, but usually under an agreement your security and legal teams have already reviewed for the same data. That is why it is often the practical middle path between a public API and your own hardware.

Hosted API, your own cloud or a local model: how the options compare
OptionWhere the data goesMain controlFits when
Consumer chat appThe provider's consumer serviceAccount settings onlyNot suitable for customer records
Commercial API under enterprise termsThe provider's infrastructureContract: no training, retention limits, DPAContracts allow the subprocessor and terms are verified
Model hosted in your cloud accountYour cloud tenancy and regionYour cloud agreement and access controlsYou already process this data in that cloud
Open-weight model on your hardwareStays inside your networkYour own security controlsStrict confidentiality or customer restrictions
Non-LLM detectors run locallyStays inside your networkYour own security controlsStructured identifiers and first-pass screening

The six-point contract and configuration checklist#

The six-point checklist covers training, retention, the DPA, location, your own pipeline configuration and customer contracts; confirm each point in writing or in configuration before any record goes to a hosted model. If any point cannot be confirmed, move the job to a model inside your own environment.

Keep the evidence with the project record: the relevant contract clauses, the DPA, the region setting and an export of the retention configuration. Then the decision can be shown later rather than remembered.

  • No training: the contract states that inputs and outputs are not used to train or improve the provider's models.
  • Retention: inputs and outputs are not retained, or are kept only for a defined window, including abuse-monitoring logs, and you know how deletion is confirmed.
  • Data processing agreement: a DPA covers the processing, lists subprocessors and sets security and breach-notice terms.
  • Location: processing happens in a region your contracts and policies allow.
  • Your own configuration: request logging, prompt tracing and observability tools in your pipeline do not store raw inputs, and API keys are scoped to this job.
  • Customer contracts: your agreements with customers permit sending their data to this provider for this purpose, or that data is excluded.

Why LLM redaction still needs a second check#

LLM redaction needs a second check because language models can miss identifiers, flag the wrong text, behave inconsistently across chunks of a long thread and quietly rewrite sentences they were asked only to clean. A rewritten resolution note can lose exactly the technical detail that made it worth keeping.

Design the job so the model returns the spans to remove, each with a label, and let ordinary code apply replacements to the original text. Run a pattern-based detector alongside it for structured identifiers such as emails, phone numbers and card numbers, and review a sample by hand.

Tool makers say the same about automated detection generally. Presidio's own documentation warns that, because it uses automated detection, there is no guarantee it will find all sensitive information, and that additional systems and protections should be employed.

What a sound LLM redaction pipeline looks like#

A sound LLM redaction pipeline treats the model as a detector, not an editor. The practices below keep the original text under your control and make every result checkable.

  • Send one record or one thread per request, with no unrelated records in the same call.
  • Ask for a structured list of spans with a category for each, never for rewritten text.
  • Give the model your allow list of product and module names so it leaves them alone.
  • Check that every returned span exists in the original text before applying it.
  • Log counts and categories, not raw text, in pipeline monitoring.

Alternatives that keep data in your environment#

Several non-LLM tools can run inside your environment or your existing cloud account and handle much of the work, especially for structured identifiers such as emails, phone numbers, card numbers and access keys.

Layer them rather than picking one. A common order of work keeps the expensive, least predictable step for the smallest share of the text:

  • Pattern rules first, for identifiers with a fixed shape: emails, phone numbers, card numbers, access keys and your own account ID format.
  • A named-entity detector next, for names, organizations and street addresses in free text, with your product allow list loaded.
  • A self-hosted model last, only for what the first two miss, such as a name written in lower case or a person described by role and location.
  • Managed cloud detection services still count as processors even inside your cloud account, so review their terms like any other vendor's.
Alternatives that keep data in your environment
ToolWhat it isNote
PresidioOpen-source, MIT-licensed SDK for PII detection and anonymization in text and imagesNow community-governed under the Data Privacy Stack organization
Google Cloud Sensitive Data ProtectionInspection, classification and de-identification for text, images and storageRuns in Google Cloud; review it as a processor like any cloud service
Amazon Comprehend PII detectionDetects PII entity types, including passwords and AWS access keys, per its 2023 developer guideRedaction ran as an asynchronous batch job; check current AWS documentation
Self-hosted open-weight modelsModels you run yourself, prompted to return PII spansQuality varies; test on your own records

Illustrative: a construction software vendor keeps redaction in-house#

Illustrative: a fictional vendor of construction project management software wants to prepare Intercom conversations and Jira issues for a licensing review. An engineer proposes pasting batches into a public chatbot to strip names.

The CTO stops the plan after checking the company's enterprise customer agreements, several of which restrict subprocessors. Instead, the team runs an open-weight model on a GPU instance in the company's existing cloud account, prompted to return labeled spans, plus a pattern detector for emails, phone numbers and project addresses.

Code applies placeholders to the original text, a reviewer reads a sample, and the CTO approves a package in which names and contact details are gone while RFI references, product terms and resolutions remain. No record leaves the company's environment during redaction.

How SourceX approaches redaction tooling#

SourceX treats the choice of redaction tooling as a supplier decision within the Preparation step of the SourceX five-step transaction. Records do not go to an outside service without the supplier's approval.

The privacy record in the SourceX Evidence Packet notes which tools ran, where they ran and how results were reviewed, so the buyer and the supplier's counsel can see how personal details were removed.

Frequently asked questions

Is an enterprise AI plan enough on its own?

Not automatically. Enterprise plans often include stronger data terms, but you still need to read the actual commitments on training, retention and subprocessors, sign or confirm a DPA, and check that your own customer contracts allow the provider to see their data.

Do we need to tell customers we used an AI tool for redaction?

It depends on your contracts and privacy notice. Some customer agreements require notice or consent before a new subprocessor is added. If redaction runs entirely inside your own environment, no new processor is involved, which simplifies the question. Confirm with counsel.

Can a small open-weight model do this well enough?

Sometimes, especially combined with pattern rules. Quality depends on the model, the prompt and how unusual your text is. Test it on a labeled sample of your own records, track misses and over-redactions separately, and keep a human review step.

Should credentials go through the same pipeline?

Detect them, but handle them differently. A password or API key found in a ticket or repository should be rotated at the source, not just masked. Dedicated secret scanners usually catch credential formats more reliably than general PII tools.

Does running redaction locally remove all privacy obligations?

No. It avoids an extra disclosure during redaction, but the decision to license, the rights review, the legal status of the prepared data and the license terms all still need their own review.

Sources

  • Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization in text and images, combining named-entity recognition, regular expressions, rule-based logic and checksums with context. Source
  • Presidio's documentation warns that because it uses automated detection mechanisms there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source
  • Presidio moved from a Microsoft-owned project to an independent, community-governed project under the Data Privacy Stack organization, recorded in release 2.2.363 dated 2026-06-28, and remains MIT-licensed. Source
  • Google's Sensitive Data Protection provides a sensitive data inspection, classification and de-identification platform that works on text, images and Google Cloud storage repositories. Source
  • Amazon Comprehend detects universal PII entity types including NAME, EMAIL, PASSWORD, AWS_ACCESS_KEY and AWS_SECRET_KEY, and redaction runs as an asynchronous batch job (June 2023 documentation snapshot). Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify