Definitions and comparisons
Redaction vs de-identification: what's the difference?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Redaction removes or masks specific details, such as a name or an account number, while de-identification is the broader outcome of making records no longer reasonably linkable to a person. Redaction is one technique; de-identification is a standard reached with several techniques plus checks. AI data buyers usually ask for the standard, not just the technique.
Key takeaways
- Redaction targets listed identifiers; de-identification also deals with combinations of details that point to one person.
- Business text leaks identity through context, such as a rare job title, a small town or a unique incident, even with every name removed.
- De-identification combines removal, replacement, generalization and suppression, then checks the result.
- Several US state privacy laws define deidentified data by conduct as well as content, including commitments and contract terms.
- A written method and a sampled review are what let a buyer accept records as de-identified.
What is the difference between redaction and de-identification?#
Redaction is the removal or masking of specific pieces of information inside a record, such as replacing a customer's name with a tag or blacking out a phone number. It works on items you can list in advance, and it is what most helpdesk redaction features and PII detection tools do.
De-identification is the end state in which records can no longer reasonably be linked to an individual, along with the process that gets there. Redaction is usually part of that process, but de-identification also asks what remains after redaction: whether job titles, dates, locations, writing style or unusual events still single someone out.
The difference shows up in how each is described. A redaction setting says which fields were masked. A de-identification method says what risk was assessed, which techniques were applied, how the output was checked and what commitments bind the recipient.
What each removes and what risk remains#
Redaction and de-identification leave very different levels of risk in business records. The comparison below reflects how privacy teams and data buyers usually see them.
| Aspect | Redaction | De-identification |
|---|---|---|
| What it targets | Listed identifiers: names, emails, phone numbers, account numbers | Direct identifiers plus combinations of details that point to a person |
| Context in free text | Usually left in place | Reviewed and generalized or removed where it identifies someone |
| Quasi-identifiers such as role, place and date | Usually untouched | Generalized, shifted or suppressed |
| Re-identification risk left | Often meaningful in narrative records | Reduced to a level the method and review support |
| Documentation | A list of masked fields | A written method, review results and recipient commitments |
| When buyers accept it | Structured fields with no narrative content | Narrative records such as tickets, emails and notes |
Why removing names is not enough in business records#
Removing names is not enough because operational records describe people through their circumstances. A dispatch note about the one certified crane operator at a named yard, a CRM note about a regional manager who just returned from medical leave, or an incident report about a well-known outage can identify someone without a single name.
These are the identifiers redaction most often leaves behind in business text.
- Email signatures and footers with titles, direct lines and office addresses.
- Quoted reply chains that repeat the original sender's details.
- Ticket, order and invoice numbers that link back to a live system.
- Rare job titles, small locations and exact dates used together.
- Screenshots, PDFs and other attachments that text redaction never reads.
- Spelled-out or spoken contact details in call transcripts.
- Customer company names, which may be confidential even when no person is named.
Techniques de-identification draws on#
De-identification draws on a family of privacy techniques, of which redaction is only one. The Data and Trust Alliance's Data Provenance Standards list redaction, masking, minimization, pseudonymization, tokenization, k-anonymity and others as separate privacy-enhancing tools, which reflects how distinct they are in practice.
For business text, a handful of techniques do most of the work. Mainstream tooling supports them: Google's Sensitive Data Protection API, for example, includes deterministic encryption for consistent replacement values and date shifting by a random number of days.
| Technique | What it does | Typical use in business records |
|---|---|---|
| Redaction or masking | Removes or hides a value | Phone numbers, card numbers, passwords in tickets |
| Consistent replacement | Swaps a value for the same token every time | Customer names across a ticket thread |
| Generalization | Replaces detail with a broader category | City to region, exact date to month, title to role family |
| Date shifting | Moves dates by a consistent offset | Keeps intervals between events while hiding real dates |
| Minimization | Drops fields that add risk but no value | Billing addresses, internal user IDs, IP addresses |
| Suppression | Removes whole records that cannot be made safe | Tickets about a single unusual incident |
How US state laws frame deidentified data#
Several US state privacy laws, the CCPA among them, define deidentified data by what the holder does as well as by what the data contains. In general terms, a business is expected to take reasonable technical measures, commit publicly not to re-identify, and require recipients by contract not to re-identify either.
That framing matters for licensing. Data that meets the definition generally falls outside personal information, which can affect whether a license counts as a sale or sharing of personal information. Which laws may apply, and whether the conditions are met, is assessed deal by deal with counsel.
Email, chat and audio need more than field masking#
Email, chat and audio need more than field masking because identity is spread through the content rather than stored in labeled fields. An email archive repeats the same people in headers, greetings, signatures and quoted replies, so a name masked in one place often survives in another.
Chat logs add nicknames, emoji reactions tied to user handles and shared files. Call recordings are harder still: a voice can identify a speaker on its own, so many projects license transcripts rather than audio, then treat spoken numbers, spelled-out emails and names said aloud as their own redaction problem.
For each format, write down which parts were processed, which were dropped entirely, and how the result was checked. That note becomes part of the de-identification method a buyer will ask to see.
Illustrative: a 3PL learns the limits of redaction#
Illustrative: a fictional third-party logistics provider keeps customer service tickets in Zendesk and exception emails in a shared mailbox. It enables the helpdesk's redaction feature and masks names, emails and phone numbers in its ticket export before discussing a license.
A sampled review finds consignee street addresses in delivery notes, driver first names in exception descriptions, and dock appointment numbers that link back to the warehouse system. The company switches to a de-identification method: consistent tokens for customer and driver names, addresses generalized to city and state, appointment numbers removed, and exception threads with one-off incidents suppressed. The reviewed sample passes, and the method is written up before delivery.
How SourceX approaches de-identification#
SourceX treats de-identification as the Preparation step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Personal and confidential details are removed with a method suited to each record type, the output is sampled and reviewed, and the supplier approves the result before anything is released.
The privacy record in the SourceX Evidence Packet describes the techniques applied and the review performed, so the buyer receives a stated standard rather than a list of masked fields.
Frequently asked questions
Is redaction enough for AI training data?
For structured fields with no narrative content, careful redaction can be enough. For tickets, emails, chat logs and notes it usually is not, because context identifies people. Buyers of narrative records generally expect a de-identification method with a sampled review.
Should replaced names be blank, tags or fake names?
Each choice trades usefulness against risk. Blank removal breaks the flow of a conversation, generic tags such as CUSTOMER keep roles clear, and consistent tokens keep track of who said what across a thread. Realistic fake names read naturally but can be mistaken for real people, so label them in documentation.
Does de-identification have to be irreversible?
In practice, the aim is that records cannot reasonably be linked back. If a mapping key exists, the data is pseudonymized rather than de-identified, and many laws still treat it as personal data. Delete or isolate keys when de-identification is the goal.
Who decides whether data is de-identified?
The company releasing the data decides, guided by counsel and by the standard written into the license. Some buyers ask for specific techniques or review steps. The decision should rest on a written method and review results, not on the setting of a tool.
Can de-identified records be re-identified later?
It is possible if new data becomes available or the method was weak, which is why licenses commonly prohibit re-identification and onward sharing. Reviewing quasi-identifiers and suppressing rare records lowers the chance, and the contract assigns responsibility if a recipient tries.
Sources
- The Data Provenance Standards' Privacy Enhancing Tools code list includes data anonymization, encryption, masking, minimization, redaction, differential privacy, k-anonymity, pseudonymization and tokenization, among others. Source
- Google's Sensitive Data Protection API supports de-identification transforms including deterministic encryption and date shifting by a random number of days. Source
- Post-CPRA, Cal. Civ. Code 1798.140(m) treats information as deidentified only if the business takes reasonable measures to ensure it cannot be associated with a consumer or household, publicly commits not to reidentify it, and contractually obligates recipients to comply. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.