Skip to content

Privacy and preparation

Screenshots and attachments: the PII that text redaction misses

By SourceX Editorial · Updated

Short answer

Text redaction misses personal details in screenshots, photos, scanned forms, PDFs and log files, because it reads a record's text, never the pixels or embedded files. Treat attachments as a separate stream: exclude the highest-risk types, run OCR and image redaction only where an attachment adds real value, strip file metadata, and have a person check every image you keep.

Key takeaways

  • A text scanner that cleans the ticket body leaves every attached file untouched.
  • Scanned IDs, signed documents, customer spreadsheets and browser log captures are usually excluded rather than redacted.
  • OCR-based redaction catches printed text in images but not faces, handwriting or a house number in the background of a photo.
  • File metadata, such as photo location tags and PDF author fields, needs its own cleaning step and its own check.

Why does text redaction miss attachments?#

Text redaction misses attachments because an export stores them as separate files, usually referenced by a link or an ID, and a text scanner reads only the comment body and fields. A screenshot of a customer's account page can show a full name, email and billing address that never appear in the ticket text.

The same gap appears wherever records carry files: warranty claims with photos, quality reports with scanned inspection sheets, CRM activities with attached proposals, and engineering issues with log files. Attachments are also where people put the details they did not want to type, such as a photo of an invoice or a scan of a signed form.

Embedded images add a second layer. An email pasted into a ticket can carry inline images and a signature logo, and a PDF can hold scanned pages beneath a text layer.

Which attachment types carry the most risk?#

Attachment types differ sharply in how much personal information they carry and how hard it is to remove. Rank them before choosing tools, because the ranking decides which types are excluded outright and which are worth the effort of redaction.

Which attachment types carry the most risk?
Attachment typeWhat it often containsRiskDefault handling
Scanned IDs, signed contracts, handwritten formsSignatures, ID numbers, handwritingVery highExclude
Spreadsheets and CSV files sent by customersBulk lists of people, accounts or ordersVery highExclude, or handle as a separate structured dataset
Browser log and network capture filesSession tokens, cookies, IP addresses, emailsHighExclude, or scan as text for secrets and PII
Product screenshotsNames, emails and account details shown on screenHighOCR and redact, then review each image
Phone photos from job sites or customer premisesFaces, house numbers, plates, location metadataHighKeep only if the image is the point; review and strip metadata
Invoices, quotes and other generated PDFsNames, addresses, payment termsMedium to highExtract text, redact, re-render as a new file
Photos of parts, equipment or defectsSerial plates, labels, occasionally peopleMediumReview, crop, strip metadata
Logos and signature imagesNames, titles and phone numbers as imagesMedium risk, low valueRemove

A decision tree: exclude, OCR and redact, or keep#

A decision tree keeps attachment handling consistent across a large archive. Each attachment type passes through the same questions in order, and the answer is recorded so the same type is handled the same way in every system.

  • Does the attachment add information the text does not? If not, exclude it. Most signature images, logos and duplicate screenshots stop here.
  • Is it a type on the exclusion list, such as IDs, signed documents, bulk exports or log captures? If so, exclude it.
  • Is the text machine-readable, as in a generated PDF, or present only as pixels? Extract machine-readable text and redact it with the same rules used for the record body.
  • For images worth keeping, run OCR with PII detection, burn solid boxes into the pixels, and run a separate check for faces and number plates.
  • Strip metadata and rename the file to a neutral ID.
  • Send every kept image to a reviewer, and record the decision against the attachment ID.

How OCR-based PII detection works, and where it fails#

OCR-based PII detection reads printed text out of an image, runs that text through the same entity detection used for documents, and draws boxes over the matching regions. The open-source Presidio SDK includes a module that redacts PII in images, and Google's Sensitive Data Protection API describes itself as working on text, images and Google Cloud storage repositories.

The weak points are predictable. OCR struggles with low-resolution screenshots, rotated phone photos, handwriting, stylized fonts and text split across table cells, so a name broken over two lines can escape detection. OCR also finds only text: a face, a signature or the front of a house needs a separate detector or a person.

The redaction method matters as much as detection. A black rectangle added as an annotation over a PDF can leave the original text underneath and copyable, and light blur or pixelation can sometimes be partly reversed. Solid boxes written into the image pixels, saved as a new file, are the safer default.

Metadata and file details to clean before release#

File metadata carries personal details that neither text nor image redaction touches. Treat metadata cleaning as its own step, run it on every kept file, and spot-check the output with a metadata viewer.

  • Location and device tags embedded in phone photos.
  • Author, last-modified-by and company fields in PDF and Office document properties.
  • Tracked changes, comments and hidden sheets in Office files.
  • Embedded thumbnails that keep a preview of the image before redaction.
  • File names that contain a customer's surname, address or account number.
  • Attachment links in the export that still point to the live system and may open without a login.

Illustrative: a pump manufacturer sorts its warranty attachments#

Illustrative: a fictional maker of industrial pumps wants to license warranty and service records from its service desk and quality system. The claim narratives are rich, but most claims also carry attachments: photos of failed seals, nameplate photos, scanned packing slips, distributor invoices and the occasional spreadsheet of affected serial numbers.

The IT lead runs the decision tree. Scanned packing slips carry signatures and are excluded, along with invoices and spreadsheets. Nameplate photos are checked first, because a serial number can tie a pump to one end customer; where it does, the photo is dropped and the serial is generalized in the text. Photos of failed parts are kept, cropped and stripped of location tags after a reviewer confirms that no people or site signage appear.

The result is smaller than the full archive but clean: claim narratives, failure codes and part photos that show what went wrong, with the excluded attachment types listed in the release record.

How SourceX handles attachments in Preparation#

In the SourceX five-step transaction, attachments are inventoried separately during Supply, with their types listed alongside the text records. The fit check runs on descriptions of those files, never the files themselves.

During Preparation, each attachment type gets an explicit decision: excluded, redacted and reviewed, or kept. The SourceX Evidence Packet lists those decisions in its privacy record, so the supplier signs off on the list before Delivery and the licensee can see which file types the package leaves out.

Frequently asked questions

Can we simply drop every attachment?

Often, yes, and it is a sensible scope for a first release. Dropping attachments removes the hardest privacy work at the cost of some context. Keep a list of what was dropped, so a later release can add reviewed screenshots or part photos if a buyer shows interest in them.

Are inline images in emails treated as attachments?

They should be. Inline images, signature logos and pasted screenshots are stored as files even when they appear inside the message body. A text redaction pass leaves them in place, so include them in the attachment inventory and run them through the same decision tree.

Does OCR redaction work on handwriting?

Poorly in most tools. Handwritten notes, signatures and filled-in paper forms produce unreliable OCR output, so detection built on that output misses identifiers. Treat handwritten material as a candidate for exclusion or full human review rather than automated redaction.

What about audio and video files attached to records?

Treat them as a separate stream with its own rules. Voices and faces identify people even after names are removed, and recording consent questions may apply. Many first releases exclude audio and video and keep only transcripts that have gone through text redaction and review.

Should we redact attachments in the live system?

Usually not for licensing. Editing attachments in a help desk or quality system changes records your team still relies on and may conflict with retention duties. Work on an exported copy, leave the originals untouched, and document how the copy was prepared.

Sources

  • Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization in text and images, and includes a module that redacts PII in images. Source
  • Google's DLP API v2 definition states that Sensitive Data Protection provides access to a sensitive data inspection, classification, and de-identification platform that works on text, images, and Google Cloud storage repositories. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify