Privacy and preparation
Redaction QA sampling: how many records should you review?
By SourceX Editorial · Updated
Short answer
Review about 300 randomly chosen records per record type as a baseline. If none contains a missed identifier, the rule of three supports a miss rate below about 1 in 100 at 95 in 100 confidence, and 3,000 clean records support about 1 in 1,000. Sample each record type separately and resample after any miss.
Key takeaways
- Zero misses in n random records bounds the miss rate at roughly 3 in n, a result known as the rule of three.
- The bound depends on how many records you review, not on how large the dataset is, once the dataset is large.
- Sample each record type separately, because tickets, emails and transcripts fail in different ways.
- A miss in the sample means fix the cause, re-run redaction on the whole record type and draw a fresh sample.
- Define what counts as a miss before review starts, and track over-redaction separately.
What does a redaction QA sample measure?#
A redaction QA sample measures how often personal details survive automated redaction, by having people read a random selection of processed records and count the ones with a missed identifier. Nobody can read a full archive, so the sample is how a COO or privacy lead turns a sense of how good the redaction is into a number that can be written down.
Pick the unit before you start. The simplest unit is the record: a ticket, an email thread or a transcript counts as a miss if it contains at least one identifier that should have been removed. Counting individual identifiers is more precise but slower, and record-level counting is what most buyers and counsel ask about.
The sample must be random within each record type. Records chosen by convenience, such as the first page of an export or the shortest tickets, measure nothing useful.
The rule of three: sample sizes at a glance#
The rule of three says that if n randomly sampled records contain zero misses, you can state with 95 in 100 confidence that the true miss rate is below about 3 divided by n. It comes from asking how high the miss rate could be while zero misses in n records would still not be surprising. In acceptance-sampling terms it is a zero-defect plan: a single miss fails the round.
The table turns that into review plans. Choose the claim you need to make, then read across to the number of clean records required.
| Clean records reviewed | Miss rate you can claim, upper bound | Where it fits |
|---|---|---|
| 30 | About 1 in 10 | Smoke test of a new configuration |
| 100 | About 3 in 100 | Tuning rounds before the full review |
| 300 | About 1 in 100 | Standard review of one record type |
| 600 | About 1 in 200 | Record types with sensitive content |
| 1,000 | About 3 in 1,000 | High-risk record types such as call transcripts |
| 3,000 | About 1 in 1,000 | Packages where the license or counsel demands a tighter claim |
Why dataset size barely changes the sample#
Dataset size barely changes the sample because the bound depends on the number of records reviewed, not on the share of the dataset they represent. A clean sample of 300 supports the same claim whether the record type is a modest archive or a very large one.
Small record types are the exception. When a record type holds only a few hundred records, the sample becomes a large share of the whole, and reviewing every record is often simpler than sampling. Do that for small but sensitive sets such as HR-adjacent email or executive threads.
Higher confidence costs more records. At 99 in 100 confidence, the rule becomes roughly 4.6 divided by n, so the same claim needs about half again as many clean records.
Dataset size does change what the bound means in absolute terms. In a hypothetical record type of 200,000 tickets, a bound of 1 in 100 still allows for up to about 2,000 tickets with a surviving identifier. That is why counsel may ask for a tighter claim on large free-text record types, or for scope changes that drop the hardest fields.
Stratify by record type and risk#
Stratifying means drawing a separate random sample for each record type, because a single pooled sample lets an easy record type hide a hard one. Redaction that works well on structured CRM notes can fail on call transcripts, and a pooled sample dominated by notes will not show it.
Keep targeted review separate from random sampling. Searching for known hard patterns, such as signatures or spelled-out emails, is a good way to find problems, but those records do not count toward the random sample that measures the miss rate.
- Record family: tickets, email threads, chat, call transcripts, CRM notes, code comments.
- Source system and era: records from before and after a system migration often format names differently.
- Content type: body text, attachments converted to text, and metadata fields.
- Risk: oversample strata that are likely to mention health, finances, minors or home addresses.
What counts as a miss?#
A miss is any identifier that the redaction rules say should be removed but that survives in the processed record. Write the rules down before review starts so reviewers do not decide case by case.
| Finding | Counts as a miss? | Action |
|---|---|---|
| Full name, email or phone number left in text | Yes | Fix the rule, re-run, resample |
| Identifier left in an attachment or metadata field | Yes | Extend redaction to that field |
| Unusual first name alone in a small team's chat | Usually yes | Treat by the written rule |
| Same person given two different tags | No, but a quality defect | Fix tag consistency |
| Product or place name tagged as a person | No, over-redaction | Track and tune separately |
What to do when the sample finds a miss#
When the sample finds a miss, the measurement is over for that round, and the job becomes finding the cause. Patching the one record and continuing to count would make the final claim meaningless.
One miss also weakens what the sample can support. With one miss in n records, the upper bound rises to roughly 4.7 divided by n; with two misses, roughly 6.3 divided by n. Fixing the cause and drawing a fresh full sample is usually faster than reviewing more records to recover the claim.
- Stop counting and record the miss with its record type and pattern.
- Identify the cause: missing pattern, text extraction error, unusual format or reviewer disagreement.
- Fix the configuration or add a preprocessing step.
- Re-run redaction across the entire record type, not just the sampled record.
- Search the whole record type for the same pattern.
- Draw a new random sample of full size and start the count again.
Illustrative: a restoration contractor plans its QA#
Illustrative: a fictional water and fire restoration company is preparing job notes, adjuster email threads, estimate line items and photo captions from its job management system. Homeowner names, addresses and insurance claim numbers appear throughout.
The COO sets four strata, plans 300 clean records for estimates and photo captions, and 1,000 for job notes and adjuster emails because those carry the most free text. Two reviewers read an overlapping portion to check that they agree, and the team seeds a handful of records with known identifiers to confirm reviewers catch them.
The first email sample turns up a homeowner's name inside a forwarded signature block. The team adds signature stripping, re-runs the email stratum, and draws a fresh sample that comes back clean. The QA log records both rounds.
How SourceX uses QA sampling#
SourceX builds QA sampling into the Preparation step of the SourceX five-step transaction. Sample sizes per record type, the definition of a miss, each round's results and any fixes are written into the privacy record of the SourceX Evidence Packet, and the supplier reviews that record before giving approval. The claim that goes into the license is the one the samples support, no stronger.
Frequently asked questions
Can an automated tool do the QA instead of people?
A second detector is a useful screen and can surface records for targeted review. It cannot measure the first tool's miss rate on its own, because both tools may share the same blind spots. Human review of a random sample is what supports the claim.
Who should review the samples?
People who know the records and the rules: often a trained operations or support lead, with privacy counsel on call for borderline cases. Reviewers should not be the same people who configured the redaction, and a portion of each sample should be read by two reviewers to check agreement.
What does the rule of three assume?
It assumes records are drawn at random and independently, that reviewers catch every miss in the records they read, and that the redaction configuration does not change during the review. Reviewer misses matter most: a reviewer who overlooks identifiers makes the bound look better than it is, which is why seeded records and double review help.
Should the buyer run its own QA?
Buyers may sample on receipt, and the license should require them to report and delete any identifiable content they find. That is a backstop, not a substitute. The supplier's QA should be complete before delivery, because once data leaves, a miss is already a disclosure.
What if no realistic sample size can support the claim counsel wants?
Then change the scope rather than the arithmetic. Excluding the hardest record type, dropping free-text fields or removing attachments can make a tighter claim achievable with the same review effort.
Related resources
- IndustryBPO & contact centers data
- QuestionDo AI labs buy medical data?
- QuestionDo AI labs buy video of people working?
- InsightHow do I de-identify customer support transcripts for AI training?
- InsightHow do I de-identify IT service tickets for AI training?
- InsightPurpose limitation: can records collected for one purpose be licensed for AI?
See if your company qualifies
A short company assessment. No data uploads are needed.