Privacy and preparation
Why helpdesk redaction settings don't clean up old tickets
By SourceX Editorial · Updated
Short answer
Helpdesk redaction settings usually apply only to messages that arrive after they are switched on, so old tickets keep every phone number, card number and signature they always held. To clean history, export a copy, detect personal details across bodies, notes and attachments, replace them consistently in the copy, then verify and record the result.
Key takeaways
- Redaction toggles and PII settings generally act when a message is stored or indexed, not on years of existing tickets.
- Internal notes, attachments, side conversations, call transcripts and integration copies often sit outside any redaction setting.
- Cleaning an exported copy is usually safer than bulk redacting inside the live helpdesk, which cannot be undone.
- A backfill has four steps: export, detect, transform in the copy, then verify and record.
Why don't redaction settings reach old tickets?#
Helpdesk redaction settings don't reach old tickets because they are built as filters on new events: an incoming email, a chat message, a call transcript or a record sent to a search index. Tickets stored before the setting was switched on are never passed through the filter again.
Vendors also scope these features narrowly. A setting may target card numbers only, apply to one channel, skip internal notes or depend on the plan tier, and defaults differ between workspaces. Helpdesks such as Zendesk, Freshdesk and Intercom each document their own options, so read the documentation for your plan rather than assuming a toggle covers history.
The practical consequence: a company that switched on redaction recently may still hold years of tickets with phone numbers in signatures, card numbers typed by customers, shipping addresses and passwords pasted during troubleshooting.
Where does personal data hide in a helpdesk archive?#
Personal data in a helpdesk archive hides well beyond the public ticket body. Map every location before planning a cleanup, because each one exports differently and fails differently.
Custom fields deserve a second look. Admins often add fields for order numbers, serial numbers or account references, and older intake forms sometimes asked customers for a date of birth or a verification code. Those fields keep whatever was collected, regardless of settings added later.
| Location | What it often contains | Reached by a new redaction setting? |
|---|---|---|
| Public replies | Names, signatures, phone numbers, order details | Only messages after the change |
| Internal notes | Account lookups, escalation context, pasted logs | Often not; check the scope |
| Attachments | Screenshots, invoices, exported spreadsheets | Rarely |
| Call recordings and transcripts | Voices, spoken card or account details | Depends on the voice product and its settings |
| Side conversations and email copies | Forwarded threads with outside parties | Varies by feature |
| User and organization profiles | Contact fields, custom fields, profile notes | Usually not |
| Satisfaction survey comments | Free-text feedback that names agents and customers | Usually not |
| Integration copies | Ticket text synced to a CRM, warehouse or search tool | No; each copy is separate |
Should you bulk redact tickets inside the live helpdesk?#
Bulk redaction inside the live helpdesk is usually the wrong first move for a data licensing project. In-place redaction is generally permanent, so it removes context agents still rely on, and it can conflict with retention schedules, legal holds or audit obligations.
Cleaning an exported copy keeps the operating system intact and produces a version you can rerun when the rules improve. Reserve in-place redaction for its proper job: removing data you should not be holding at all, such as full card numbers, under your own retention and security policies.
| Approach | Best for | Main risk |
|---|---|---|
| Redact inside the live helpdesk | Removing data you should never have stored | Irreversible; can break holds and agent context |
| Clean an exported copy | Preparing a dataset for analysis or licensing | Gaps if fields or attachments are missed in export |
| Delete old tickets | Enforcing an agreed retention schedule | Loses history permanently, including its value |
A four-step backfill plan for historical tickets#
A backfill plan cleans history in a copy, in four steps that can be repeated as rules improve. Each step produces something a reviewer can check.
Export routes shape Step 1, so confirm them before promising a date. Zendesk's account data exports must be switched on by Zendesk support and are not offered on Team plans, where the REST API is the route. Zendesk also archives tickets 120 days after they close; archived tickets drop out of views but stay reachable by search and API, so a view-based export can miss them. Intercom's S3 data export offers a historical export of two years of conversations, so older history may need another route. Check, too, whether any ticket deletion schedule is still removing archived tickets while you plan.
- Step 1, export: pull tickets, comments, internal notes, users, organizations, custom fields and attachment metadata with their IDs, through the API or a full export, and note the date range, any gaps in access and the date each redaction setting was switched on, since tickets older than that were never filtered.
- Step 2, detect: run pattern rules for card numbers, phone numbers, emails and credentials, a named-entity model for names and addresses, and a deny list of customer names taken from your CRM.
- Step 3, transform: replace each finding in the copy with one consistent placeholder per person or account, remove credentials entirely, and keep product names, error codes and outcomes.
- Step 4, verify and record: review a sample by hand, rescan the transformed copy, log what changed and keep the configuration so the run can be repeated.
Credentials found in old tickets are a security task first#
Credentials found in old tickets need rotation at the source system, not just masking in the copy. A password, API key or remote access code pasted into a ticket years ago may still work, and anyone with helpdesk access has been able to read it since.
Send credential findings from Step 2 to the security owner as soon as they appear, rather than waiting for the backfill to finish. Record that each one was rotated or confirmed dead, and keep that list separate from the dataset.
How do you know the backfill worked?#
The backfill worked when a fresh scan of the cleaned copy finds no remaining contact details or credentials, and a human review of a sample confirms that threads still make sense. Both checks are needed; a clean scan only proves that the patterns you wrote stopped matching.
Draw the review sample from across the archive's history, not just recent tickets. Older tickets often came through a previous helpdesk or a different intake form, and they fail in their own ways: legacy signatures, scanned attachments, or customer numbers in a format your rules do not recognize.
- Search the cleaned copy for a set of known customer and agent names taken from the CRM and the staff roster.
- Search for card number, phone and email patterns across bodies, notes and custom fields, not only ticket text.
- Open threads from each year of history and confirm the conversation still reads in order.
- Confirm every attachment is either excluded or reviewed, with none left in an unknown state.
- Compare record counts before and after transformation so no tickets were silently dropped.
Illustrative: a property management software vendor finds its history untouched#
Illustrative: a fictional vendor of property management software switches on its helpdesk's PII redaction option when it connects an AI assistant to the ticket index. Leadership assumes the archive is now clean and asks IT to prepare a ticket sample for a licensing review.
The export shows otherwise. Older tickets still contain landlord and tenant phone numbers in email signatures, bank details in pasted rent ledgers and gate codes in maintenance requests. Internal notes hold account lookups, and attachments include lease scans.
IT leaves the live helpdesk alone, runs the four-step backfill on the exported copy, excludes attachments from the first package and rotates the shared passwords found in old notes. The privacy lead approves the cleaned sample, and the company records the date the setting was enabled so later reviewers know which tickets were never filtered.
How SourceX handles historical ticket archives#
SourceX starts with metadata: which helpdesk, how many years of tickets, which channels and which settings were in force when. No ticket leaves the company during that fit check. Cleaning happens later, in the Preparation step of the SourceX five-step transaction, on a copy the supplier controls.
The SourceX Evidence Packet records the backfill in its privacy record: export scope, detection rules, transformation choices, review results and the date any helpdesk setting changed. The supplier approves that record before release.
Frequently asked questions
If we turn on redaction today, will last year's export be clean?
No. An export reflects what is stored, and settings switched on today usually affect only new messages. Treat any export of tickets created before the change as unredacted until you have cleaned the copy yourself and checked it.
Does deleting old tickets solve the problem?
Deletion removes the risk and the value together, and it may conflict with retention or legal hold obligations. If tickets are past your retention schedule, deletion may be the right call. If they are kept for business reasons, cleaning an exported copy is the more useful route.
Do internal notes need the same treatment as public replies?
Yes, and often more. Internal notes tend to hold account lookups, escalation details and pasted logs, which carry identifiers and sometimes credentials. They also hold the reasoning behind a resolution, so replace identifiers rather than dropping the notes.
Can we reuse the helpdesk's own redaction rules for the backfill?
They are a reasonable starting point, but they were written for live traffic and a narrow set of patterns. A backfill also needs rules for legacy formats, customer names from your CRM and attachments, plus a human review that the built-in feature does not provide.
What about ticket copies in our CRM or data warehouse?
Each synced copy is a separate store with its own history, and a helpdesk setting does not reach it. List every integration that received ticket text, then decide whether each copy is in scope for the dataset or should be excluded.
Sources
- Zendesk account data exports are not on by default and must be enabled by Zendesk support; they are not available on Team plans, where the REST API is the export route. Source
- Zendesk archives tickets 120 days after they reach Closed status; archived tickets remain reachable by search, direct link and API but do not appear in views. Source
- Zendesk ticket deletion schedules permanently delete archived tickets that match their criteria. Source
- Intercom's S3 Data Export offers a historical export covering two years of conversations data. Source
Related resources
- IndustryBPO & contact centers data
- QuestionDo AI labs buy medical data?
- QuestionDo AI labs buy video of people working?
- InsightHow do I de-identify customer support transcripts for AI training?
- InsightHow do I de-identify IT service tickets for AI training?
- InsightPurpose limitation: can records collected for one purpose be licensed for AI?
See if your company qualifies
A short company assessment. No data uploads are needed.