Skip to content

Privacy, de-identification and sensitive data

Redacting PII in scanned documents for document-AI training: pixels, OCR layers and metadata

Quick answer

A scanned document carries personal data in at least three places: the rendered pixels, any OCR or embedded text layer, and file-level metadata. A dataset is only de-identified when all three are cleaned consistently. Use solid, same-size masks or synthetic fill instead of blur, regenerate or scrub the text layer so it matches the redacted image, strip XMP, EXIF and document-info fields, and accept delivery only after a sampled visual review and a text-layer search both come back clean.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Where personal data hides in a scanned document file

Personal data in a document file sits in more layers than the visible page, and each layer needs its own check. Treating the image as the whole document is the most common reason a "redacted" corpus still leaks names and account numbers.

The layers a buyer should expect a supplier to account for:

  • Pixel layer. Typed and handwritten names, addresses, policy and account numbers, signatures, photos on ID cards, stamps with notary names, and barcodes or QR codes that encode identifiers.
  • Text layer. Searchable PDFs produced by scanners or OCR engines carry invisible text under the image. Drawing a black box over the image does not touch that text, so copy-paste or pdftotext still returns it.
  • Annotation and form objects. PDF AcroForm field values, comments, bookmarks and attached files can hold the original values even after the page is flattened visually.
  • File metadata. The PDF document-info dictionary (Author, Creator, Producer), XMP packets, TIFF and JPEG EXIF tags (device serials, GPS, operator names) and file names such as claim_SMITH_J_2024.pdf.
  • Sidecar annotations. Ground-truth JSON or hOCR/ALTO files that list every word with its bounding box. If the label file still contains the original string, the image redaction is cosmetic.

Health documents add a further concern. HIPAA Safe Harbor lists 18 identifier types, including full-face photographs and biometric identifiers, that must be removed for the individual and for relatives, employers and household members [4]. Medical imaging offers a useful parallel: DICOM's Basic Application Level Confidentiality Profile takes a deliberately conservative approach to removing patient-identifying attributes, while burned-in text in the pixel data is handled as a separate concern [8]. The same discipline applies to a scanned claim form or referral letter.

Why blur and pixelation fail as document redaction

Blur, pixelation and low-opacity overlays should be rejected outright, because they can be reversed. Research on blurring for image data release, mostly on faces and license plates, finds that the privacy blur provides varies with what is blurred and how [1], so it is not a dependable guarantee. Printed text is especially vulnerable: the font, size and character set are constrained, so a candidate string can be rendered, blurred the same way and matched.

Accept only these pixel treatments:

  • Opaque fill with a fixed color, applied to the flattened raster, not as a vector object on top of it.
  • Synthetic fill, where a realistic fake value (a surrogate name or a format-valid fake account number) is rendered into the box in a similar font. This keeps more training signal but needs a record that the value is synthetic.
  • Region removal for whole-page identifiers such as ID photos, where the box is replaced with a neutral background.

Ask how the redaction was applied. A box drawn in a PDF editor and saved without flattening often leaves the original image intact underneath the drawn object.

Keeping the text layer consistent with the redacted image

The text layer must be redacted or regenerated so that it agrees with the image, or a model will learn from text the image no longer shows. There are two acceptable approaches, and the buyer should know which one was used.

  1. Regenerate. Rasterize the redacted page, discard the original text layer, and run OCR again. This is the cleanest for privacy, but OCR output will now include whatever the mask or surrogate contains.
  2. Scrub in place. Delete or replace text objects whose bounding boxes intersect a redaction region, then confirm with a text extraction pass. This preserves the original OCR quality elsewhere on the page.

Either way, the ground-truth annotations must be updated too. For a key-value extraction set, a redacted policyholder_name field needs a decision: drop the label, keep the key with an empty value, or keep a surrogate value that matches the synthetic fill. Datasets like FUNSD, with 199 forms labeled for text detection, OCR, layout and entity linking [2], show how tightly words, boxes and links are coupled. A mismatch between label and pixels corrupts both training and evaluation; see document extraction evaluation ground truth for why field-level labels must match the image exactly.

Detection tooling helps but does not settle the question. Microsoft's Presidio project, a widely used PII detection and anonymization SDK, states that its ML models give no guarantee of finding all sensitive information and should be paired with other protections [6]. Handwriting, rotated stamps and low-contrast fax scans are where recall drops.

Preserving layout signal while removing identifiers

Good redaction preserves geometry: same-size masks, unchanged page dimensions and resolution, and consistent label boxes. Document-AI models learn from where fields sit, how tables align and how text density varies, so redaction should change content, not structure.

Practical rules buyers can put in a specification:

  • Same-size masks. The mask covers the original token's bounding box, padded by a small fixed margin, not a whole line or block unless the line is entirely personal data.
  • No resampling. Keep DPI, color mode and compression the same as the source scan, so the model does not learn redaction artifacts as a class signal.
  • Typed surrogates where layout matters. For receipt and invoice extraction tasks like SROIE [3], a masked total or date field removes a target the model needs. Surrogate values in the correct format keep the task intact while removing the real values.
  • Mark the redaction. Each redacted region gets an entry in a manifest with its box, entity type and treatment, so you can exclude it from loss or evaluation.

The trade-off between masks and surrogates is covered in masking vs surrogate replacement.

Signatures, handwriting and other hard identifiers

Signatures, handwritten names and stamps should be treated as identifiers by default. Handwriting recognizers miss cursive and mixed scripts, and a signature is often recognizable even when no name can be read from it.

Concrete handling choices:

  • Signatures: mask the entire signature block, including initials in margins. If signature detection is a training target, ask for synthetic signatures rather than real ones.
  • Handwritten free text: route pages with dense handwriting to human review, because automated NER on handwriting OCR output is unreliable.
  • Barcodes and QR codes: decode them during QA. A visually intact barcode can encode a member ID or tracking number.
  • Faces and ID photos: remove the full region, not just the eyes.

Re-identification risk also comes from combinations of fields that look harmless alone, such as date of service, ZIP code and a rare procedure. NIST SP 800-188 cautions that traditional de-identification has inherent limits compared with formal privacy methods [5]. Scanned forms rarely suit formal methods, so the residual risk must be argued case by case.

Acceptance tests a buyer should run on delivery

Acceptance should combine automated text-layer search, metadata inspection and a sampled visual review, and every test should be repeatable. Run them on every delivery, not only the first.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckMethodPass condition
Text-layer leakageExtract all text with pdftotext or a PDF library; search for known-format patterns (SSN, card PAN with Luhn check, email, phone)No matches outside documented surrogates
Hidden objectsList AcroForm fields, annotations, embedded files and optional content groupsNone present, or values cleared
MetadataDump document-info, XMP and EXIF with exiftoolNo personal names, device serials, GPS or source file names
Pixel residueSampled human review at 100% zoom of masked regions and marginsNo legible identifier; no blur or partial masks
Label consistencyCompare annotation strings to redaction manifestNo original values in labels for redacted boxes
Machine codesDecode all barcodes and QR codesNo real identifiers
Layout integrityCompare page size, DPI and box counts to pre-redaction statsWithin agreed tolerance

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_id": "clm-000417",
  "page": 2,
  "region_bbox": [412, 1180, 768, 1222],
  "entity_type": "MEMBER_ID",
  "detector": "pattern+human_review",
  "treatment": "synthetic_fill",
  "surrogate_value": "ZX4410093",
  "text_layer_action": "regenerated_ocr",
  "label_action": "replaced_with_surrogate",
  "qa_sampled": true
}

Size the review sample using the approach in auditing residual PII with sampling plans, and request the redaction manifest as part of the de-identification evidence package.

The legal bar depends on the data and jurisdiction, so the redaction specification should name the standard it targets. For health records, HIPAA de-identification uses either Safe Harbor or Expert Determination [4]. For California consumer data, the CCPA defines deidentified information and requires a business holding it to meet three conditions, including reasonable measures against re-identification and contractual commitments from recipients [7].

These standards are not satisfied by a tool output alone. They require documented methods, a residual-risk view and controls on the recipient. The broader framing for buyers is in the privacy cluster guide to de-identified data for AI training.

How SourceX handles scanned documents

SourceX sources operational datasets from US companies, including documents and finance and legal workflows, on request rather than from stock, so a request does not guarantee a match. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Health records require HIPAA de-identification through Safe Harbor or Expert Determination. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. You can describe the documents you need on the SourceX buyers page.

For related reading, see document AI datasets by task, the guide to de-identifying scanned forms and handwritten documents, and de-identifying invoices and receipts.

Request de-identified scanned document data

SourceX looks for US businesses that hold the scanned forms, claims or invoices you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Describe your document-AI data needs to SourceX.

Sources

  1. arXiv, "Privacy Blur: Quantifying Privacy and Utility for Image Data Release" (2025). https://arxiv.org/pdf/2512.16086
  2. Jaume, Ekenel, Thiran (arXiv:1905.13538), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  3. Huang et al. (arXiv:2103.10213), "ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction" (2021). https://arxiv.org/pdf/2103.10213
  4. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  5. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  6. Data Privacy Stack (GitHub Pages), "Presidio - Data Protection and De-identification SDK" (2026). https://data-privacy-stack.github.io/presidio
  7. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  8. DICOM Standards Committee, "DICOM PS3.15 2026c: Security and System Management Profiles, E.3.1 Clean Pixel Data Option" (2026). https://dicom.nema.org/medical/dicom/current/output/chtml/part15/sect_E.3.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data