Skip to content

Privacy and preparation

Human review vs automated redaction: how to split the work

By SourceX Editorial · Updated

Short answer

A human review redaction workflow splits the work into four stages: an automated pass removes predictable identifiers across every record, a risk-ranked queue sends the records most likely to hide personal details to trained reviewers, a random sample tests what was released, and a named owner signs off. Machines carry the volume; people carry the judgment.

Key takeaways

  • Automate the identifiers that follow patterns, such as email addresses, phone numbers and account numbers, and reserve people for free text and context.
  • Route records to reviewers by risk signals, not at random: long notes, complaints, low-confidence detections and attachments go first.
  • Sample the released output, not the input, and fix the rule behind any miss before re-running the affected slice.
  • One named owner signs the release after reading a short package: sample results, the exclusion list, tool configuration and open escalations.

Why does redaction need both machines and people?#

Redaction needs both because automated tools and human reviewers fail in opposite ways. Tools are fast and consistent on identifiers with a fixed shape, but they miss a first name typed in lowercase, a nickname, or a phrase like the dispatcher who retired last spring. People catch that context, yet a reviewer reading years of customer emails tires, drifts and skips lines.

Tool makers concede the point. The open-source Presidio project tells its users plainly that automated detection cannot promise complete coverage and that other safeguards belong alongside it, a caveat that fits any automated detector.

The design question for an operations leader is therefore not manual vs automated redaction. It is where the hand-off sits, what moves a record from one side to the other, and who owns each stage.

The four stages of a hybrid redaction process#

A hybrid redaction process runs in four stages, and each stage has a single owner and a written exit condition. Skipping a stage usually shows up later as a miss nobody can explain.

  • Automated pass: detection tools run over every record and replace identifiers with typed placeholders, so a removed name reads as a name tag rather than a blank. Owner: data engineering or the preparation team.
  • Risk-ranked queue: records that trip a risk rule go to trained reviewers, highest risk first. Owner: the review lead.
  • Sample QA: a random sample of released records is checked line by line for anything that slipped through. Owner: someone who did not review those records.
  • Sign-off: a named executive approves release after reading the sample results, the exclusion list and any open issues. Owner: the COO or designated data owner.

Which records should go to the human review queue?#

The human review queue should hold the records where automated detection is weakest or where a miss would cost the most. Rank by signals that can be computed from the record itself, so the queue builds automatically and reviewers start with the worst cases.

Records about employees deserve their own rule. HR cases, disciplinary notes and payroll questions are often excluded outright rather than reviewed, because even a well-redacted complaint can point to the one person on a small team it describes.

Which records should go to the human review queue?
Risk signalWhy it raises riskQueue priority
Long free-text notes or email threadsNames, places and personal stories hide in narrative textHigh
Complaints, disputes and escalationsPeople describe health, money, family or legal troubleHigh
Low-confidence detections from the toolThe model was unsure, so a person decidesHigh
Attachments, screenshots and scanned formsText scanners do not read pixels or embedded filesHigh, or exclude
Custom fields the tool has not seen beforeFree-form fields such as job contact or special instructions often hold namesMedium
Short structured records with no free textIdentifiers follow patterns the automated pass handlesLow, sample only

Who does what on a redaction review team?#

A redaction review team works best with four clear roles and a firm line between doing the work and checking it. The person who reviews a record should never be the person who passes it in QA.

Reviewers who know the business catch more. A senior support agent can tell a product name from a customer name at a glance, and a dispatcher knows which driver nicknames show up in notes. Give reviewers a written guide with real examples from your systems, run a calibration round where two people label the same records, and settle disagreements before the main queue opens.

Reviewers should work in a controlled environment with downloads disabled, and log each decision with a reason code. That log becomes evidence that the release was reviewed, not just processed.

Who does what on a redaction review team?
RoleOwnsDoes not own
Data engineer or preparation teamExports, tool configuration, placeholders, re-runsDeciding what is acceptable to release
ReviewersQueue decisions: keep, redact more, or exclude the recordChanging detection rules on their own
Privacy lead or counselThe release standard, exclusion rules, rulings on edge casesDay-to-day queue work
COO or data ownerFinal sign-off and the exclusion listReviewing individual records

How should sample QA work, and what happens after a miss?#

Sample QA should test the released output, drawn at random across every source system and record type, by someone who did not review those records. Its job is to estimate what is left in the dataset, not to polish the records that happen to be sampled.

Agree the pass standard with counsel before the first sample is drawn. A team that sets the bar after seeing results will be tempted to set it wherever the results landed.

  • Draw the sample after redaction and review, from the release set itself.
  • Stratify by source system and record type so a small system is not drowned out by a large one.
  • Check every line, including signatures, quoted replies and field labels.
  • Log each miss with its category, such as person name, street address or account number.
  • When misses cluster in one category or system, fix the rule or the queue routing and re-run that whole slice.
  • Draw a fresh sample for the re-test instead of re-checking the same records.

What should the sign-off package contain?#

The sign-off package should let the approving executive decide on release without reading individual records. It is a short file assembled by the review lead and the privacy lead, and it travels with the dataset into the deal file.

Sign-off is a decision, not a formality. If the final sample missed the agreed standard, or escalations are still open, the right answers are to send the slice back or to narrow the scope. A signer who cannot explain in a sentence why each excluded record type was left out is not ready to sign.

  • The release standard agreed with counsel before testing, and whether the final sample met it.
  • Sample results by source system and category of personal detail, including each miss and the rule change that fixed it.
  • The exclusion list: record types, fields, attachment types and people left out, such as opt-outs and deletion requests, with the reason for each.
  • The tools, versions and configuration used in the final run.
  • A summary of the reviewer log: records reviewed, outcomes and any escalations still open.
  • Who performed each stage, showing that sample QA was independent of the review queue.

Illustrative: a wholesale distributor splits its redaction work#

Illustrative: a fictional regional electrical supply distributor plans to license several years of order exception records from its ERP, along with the customer service email threads linked to them. The COO wants the work done without pulling the service team away from customers for long.

The automated pass removes email addresses, phone numbers, account numbers and delivery addresses, and replaces contact names using the CRM contact list as a dictionary. Risk rules send credit holds, damage claims and unusually long threads to the queue, where two senior service reps review them.

The first QA sample finds truck drivers' first names in delivery notes, which neither the tool nor the CRM list covered. The team adds the driver roster to the dictionary, re-runs every delivery note and draws a new sample that comes back clean. Driver incident reports are excluded entirely, and the COO signs the release with the exclusion list attached.

How SourceX divides redaction work#

SourceX handles redaction inside Preparation, the third stage of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. The automated pass, the queue rules and the QA sample are agreed with the supplier before any records are processed, and nothing is shared during the initial assessment.

The results go into the privacy record of the SourceX Evidence Packet: which tools ran, which records people reviewed, what the samples found, what was excluded and who signed. The supplier sees that record and approves release in the Approval step before anything moves to Delivery.

Frequently asked questions

Can we skip human review if a tool reports high accuracy?

Not safely. Vendor accuracy figures come from test data that rarely looks like your support emails or job notes. Measure the tool on a labeled sample of your own records first, and even when it performs well, keep a human queue for categories where a single miss would matter, such as complaints and HR topics.

Should reviewers be employees or outside contractors?

Either can work. Employees know product names, customers and internal shorthand, which speeds decisions, but they may recognize the people in the records and need clear confidentiality rules. Outside reviewers bring capacity and distance and need a stronger written guide. Many teams mix the two, with employees taking the hardest queue.

Should redacted text be deleted or replaced with placeholders?

Placeholders are usually better. A typed tag such as customer name or street address keeps the sentence readable and tells an AI developer what kind of detail was removed, which preserves value. Use the same tags across every system, and avoid placeholders that leak information, such as initials or partial numbers.

What should a reviewer do when unsure?

Escalate rather than guess. Give reviewers three outcomes for every record, keep, redact more or exclude, plus an escalate flag that sends the record to the privacy lead. Unclear cases that recur become new written rules, so the guide gets sharper as the queue moves.

Does every record need a second reviewer?

No. Double review is expensive and best reserved for the highest-risk categories and the calibration round. For everything else, single review plus independent sample QA gives a clearer picture of residual risk than doubling every decision.

Sources

  • Presidio's own documentation warns that because it is using automated detection mechanisms, there is no guarantee that Presidio will find all sensitive information, and that additional systems and protections should be employed. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify