Skip to content

Document AI data

Redaction Training Data: Documents with Human Redaction Decisions

Quick answer

Redaction training data, a specialized slice of document AI data, is a set of realistic documents in which trained reviewers marked every region they withheld, with the region's coordinates, the entity type, a reason code (a PII category, a FOIA exemption, a privilege basis) and the decision itself. The labels that matter most are the ones only humans supply: what was redacted, what was deliberately left visible, and why. Public sets rarely include all of this, so most teams combine public data, synthetic substitution and licensed reviewer decisions.

By SourceX Editorial · Updated

What a redaction label has to contain

A usable redaction label records geometry, semantics and a decision, not just a black box. A detection-only set (bounding boxes around names) trains a PII detection model, but a production redaction tool also has to decide whether a detected span should be withheld in this document, for this purpose. A name in a signature block may be released while the same name in a medical note is withheld.

At minimum, ask for these fields per region:

  • Geometry: page index, bounding box or polygon in pixel coordinates at a stated DPI, plus character offsets into the OCR text layer when one exists. COCO-style boxes, as used by layout sets such as DocLayNet, are a reasonable interchange format [7].
  • Entity type: person name, SSN, account number, address, date of birth, medical record number, signature, face, handwritten note.
  • Reason code: the legal or policy basis, such as a FOIA exemption code under 5 U.S.C. 552(b), a privilege basis (attorney-client, work product), or a HIPAA identifier category [5].
  • Decision: redact, release, or partial (for example, keep the last four digits), plus a flag for regions that were considered and deliberately left visible.
  • Provenance: reviewer role, review round, whether the region came from a tool suggestion that a human accepted, moved or deleted.

The last two fields are what separate redaction data from generic entity annotation. Without "considered and released" negatives, a model learns to over-redact, which in FOIA and discovery work produces its own complaints and rework.

Why redacted documents alone are weak training signal

Finished redacted documents show where the boxes went but not what was under them, so they cannot teach detection on their own. A USPTO patent on this problem notes that redacted documents may lack the signal needed to train models and describes generating training data from redacted information instead [2]. The underlying text is gone, so you cannot compute what the model should have found.

There are three practical routes around this:

  1. Before/after pairs from the producing organization. The unredacted original plus the reviewer's markup is the strongest signal, but the originals contain exactly the data the redaction was meant to protect. These pairs are rarely shareable, and when they are, they need their own de-identification and access controls.
  2. Synthetic substitution in original layouts. Replace the real values under each region with realistic synthetic values (a fake but well-formed SSN, a plausible name and address) while keeping the page layout, fonts, scan noise and the reviewer's region labels. This keeps the visual difficulty and the decision labels without shipping the original identifiers.
  3. Tool-correction logs. Where a redaction assistant proposes boxes and humans edit them, the edits themselves become labels. Axon describes using customer redaction edits to teach a redaction-assistant model where its detections were incomplete or mis-drawn [1]. This is the same pattern covered for extraction in document extraction correction logs and for decisions generally in human override and correction logs.

Fully synthetic documents are useful for pretraining detectors, and graph-based layout generation is one research approach [9]. They tend to miss the long tail that breaks redaction in practice: handwritten margin notes, faxed cover sheets, rotated stamps and identifiers embedded in tables. See where synthetic documents break before relying on them for evaluation.

Public FOIA releases: useful for reason codes, not detection

Public FOIA releases are a good source of reason-classification examples and a poor source of detection labels. FOIA lets agencies withhold information under nine exemptions, requires release of any reasonably segregable portion, and requires agencies to indicate where deletions were made, which in practice means an exemption code such as (b)(6) printed beside each box. That gives you thousands of real pages pairing visible context with a stated legal reason.

What you cannot get from them is the hidden text. A box marked (b)(6) tells you a privacy-based withholding happened near that context, not whether it covered a name, a phone number or a whole paragraph. Use these releases to train or test a model that suggests the right exemption for a region a human or detector has already found, and to learn agency-specific conventions such as citing (b)(6) and (b)(7)(C) together.

Two cautions apply. First, agency practice varies, and many exemptions permit rather than require withholding, so the same content may be withheld by one office and released by another. Second, check the reuse terms of any portal you collect from; government works are often unrestricted, but third-party material inside a release may not be. The public document datasets for commercial use page covers license checks for public corpora.

Each workflow has its own reason taxonomy, and a model trained on one transfers poorly to another. Decide which taxonomy your product ships with before you source a single page.

WorkflowTypical reason codesHard cases to requestMain source of truth
Public records / FOIA5 U.S.C. 552(b)(1)-(9), state public-records exemptionsPartial redaction of email headers; (b)(5) deliberative passagesAgency release letters and markings
Litigation and eDiscoveryAttorney-client, work product, PII, confidentiality designationsPrivileged passages inside otherwise responsive email threads; forwarded chainsPrivilege logs tied to document IDs
HealthcareHIPAA Safe Harbor identifier categories, Expert Determination rulesDates more specific than year, small geographic units, free-text notes [5][6]Covered entity's de-identification procedure
Financial servicesAccount numbers, SSN, card PAN, customer namesStatements with tables, check images, handwritten annotationsInternal data-handling policy

Privilege redaction is the hardest to source. The label depends on a lawyer's judgment about the communication, not on a pattern in the text, and privilege logs are themselves sensitive. Expect to need reviewer rationale fields and adjudication records rather than just boxes.

For healthcare, HHS describes two routes to de-identification under 45 CFR 164.514: Safe Harbor, which removes 18 listed identifiers, and Expert Determination [5][6]. A redaction model for clinical documents should carry the Safe Harbor category as its entity type so that evaluation can report recall per identifier class.

Visual PII on document images

Image-level redaction needs labels for content that never reaches the OCR layer. Signatures, faces on ID cards, handwritten names, barcodes and QR codes that encode identifiers, and stamps with names or license numbers all need polygon labels on the image itself. The signature, stamp and seal detection page covers that class of objects in more depth.

OCR quality bounds text-based detection. If the OCR misreads "J0hn Sm1th" on a faxed form, a text model never sees a name, and the box is missed. Pair redaction labels with OCR ground truth where possible, and keep the noisy-scan cases that public form sets such as FUNSD were built to represent [8].

How to build a redaction evaluation set

Measure recall first, per entity type and per reason code, because a missed identifier is the expensive error; then measure precision to control over-redaction. A single micro-averaged F1 hides the failure that matters, such as high recall on typed names alongside much lower recall on handwritten account numbers. TAB, the Text Anonymization Benchmark, is a useful reference here: it pairs court-case text with human masking decisions and proposes evaluation measures specific to anonymization rather than plain entity-recognition scores [3].

Practical rules for the eval split:

  • Score at the region level and the character level. A box that covers 80% of an SSN is a leak, so count partial coverage as a miss for high-risk types.
  • Hold out by source organization and template, not by page, so the model has not seen the layout before.
  • Double-review a sample. DocLayNet annotated a subset two or three times to measure inter-annotator agreement, and its baseline models scored roughly 10% below that agreement [7]. Your ceiling is reviewer agreement, not 100%.
  • Include LLM-based redactors as baselines. Recent work quantifies both the capabilities and the risks of LLMs used for PII redaction [4], and a buyer should know where a general model already stands before paying for labels.

The PII redaction for LLM training data page covers measuring leakage in text pipelines; this page is about the labels you need to train and score the redactor itself.

Request template for redaction training data

Illustrative example: invented to show structure; it does not describe an available dataset.

request: redaction-training-data
use: train and evaluate document-image redaction model
documents:
  types: [records-request responses, insurance claim files, clinical referral letters]
  formats: [PDF (scanned and born-digital), TIFF]
  min_dpi: 200
  languages: [en-US]
labels_per_region:
  geometry: polygon, page_index, char_offsets (if OCR layer)
  entity_type: [person_name, ssn, account_number, dob, address, mrn, signature, face]
  reason_code: [foia_b6, foia_b7c, privilege_ac, privilege_wp, hipaa_safe_harbor_<category>]
  decision: [redact, release, partial]
  considered_not_redacted: required
  provenance: [reviewer_role, review_round, tool_suggested, tool_edit_type]
identifier_handling: real values replaced with synthetic values in place; method documented
eval_split: held out by source organization; 10% double-reviewed
rights: license covering model training and evaluation; supplier approval per release

Ask suppliers how values were substituted, whether substitutions preserve format and length, and how they confirmed nothing real survived. The playbook on de-identifying company data for AI describes the controls to look for.

Rights and privacy checks before you license

Redaction data is unusual because the training signal and the privacy risk sit in the same pixels. Before you buy, confirm who owns the documents and the reviewer markup, whether the original data subjects' information has been replaced or removed, and whether the license allows training and evaluation for your specific product. Health records carry HIPAA de-identification requirements, and legal documents may carry protective orders or confidentiality designations that survive the matter.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

If you need reviewer decisions from real operational documents rather than public releases, you can describe the documents and labels you need to SourceX. SourceX sources operational datasets from US companies on request; documents and finance and legal workflows are among the kinds of data it looks for, and a request does not guarantee a match.

Sourcing redaction training data with SourceX

SourceX finds US businesses that hold the documents you describe, and every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Start a buyer request for redaction training data.

Frequently asked questions

Can I train a redaction model only on public FOIA releases?

You can train a reason-code classifier on them, but not a detector, because the withheld text is not visible. Pair them with substituted or licensed before/after data for detection.

Should redaction labels include regions that were not redacted?

Yes. "Considered and released" regions are the negatives that stop a model from over-redacting, and they are easy to omit when a set records only the boxes that were drawn.

How large should a redaction evaluation set be?

Size it per entity type rather than in total pages. Each high-risk type needs enough examples, across held-out templates, for its recall estimate to be stable.

Sources

  1. Axon, "Detailed use case description: redaction assistant model training". https://a.storyblok.com/f/198504/x/749b7c25f7/detailed-use-case-description-redaction-assistant-model-training.pdf
  2. United States Patent and Trademark Office, "US Patent 12468778 (generating training data from redacted information)". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/12468778
  3. Pilán et al., arXiv, "The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization" (2022). https://arxiv.org/pdf/2202.00443
  4. arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
  5. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  6. Electronic Code of Federal Regulations, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  7. IBM Research, arXiv, "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  8. Jaume, Ekenel, Thiran, arXiv, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  9. arXiv, "Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts" (2024). https://arxiv.org/pdf/2412.03590

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data