Skip to content

Privacy and preparation

What not to redact: keeping product names, error codes and outcomes

By SourceX Editorial · Updated

Short answer

Do not redact product names, version numbers, error codes, reproduction steps or outcomes: they are rarely personal data and carry most of the value in support and engineering records. Remove names, contact details, account identifiers and credentials instead. The rule: redact who was involved and how to reach them, not what broke and how it was fixed.

Key takeaways

  • Product names, version numbers, error codes and resolution notes are rarely personal data and carry most of a record's value.
  • Names, emails, phone numbers, account IDs, IP addresses and credentials should be removed or replaced in every field, including logs and signatures.
  • Replace identities with consistent placeholders instead of deleting them, so a thread still shows who said what.
  • Allow lists for product terms and custom patterns for error codes stop detection tools from erasing technical detail.
  • Review a sample of redacted records for lost meaning, not only for missed personal data.

What does over-redaction destroy in support and engineering records?#

Over-redaction destroys the technical chain that explains a problem: which product failed, in which version, with which error, and what fixed it. A ticket that reads 'customer reported [REDACTED] on [REDACTED] after upgrading to [REDACTED]' tells a model nothing about how software support works.

The damage usually comes from detection tools tuned for recall. Named-entity models flag capitalized module names as organizations or places, number patterns catch error codes and build numbers as account numbers, and blanket rules delete every URL, including links to public documentation. Each choice looks cautious on its own; together they turn a resolved incident into noise.

AI developers license operational records to learn how work gets done. For a software company, that work is diagnosis and resolution, so the technical nouns are the payload. Protecting people and protecting value are compatible goals once the redaction plan separates who was involved from what happened.

Keep or remove: a field-by-field list#

A keep and remove list assigns every field type a default treatment before any tool runs. Writing it down forces the team to settle edge cases once, instead of letting a model's confidence threshold decide them ticket by ticket.

The list below fits most helpdesk, issue tracker and engineering chat exports. Adjust it for your product, then hand it to whoever configures the detection tool.

Keep or remove: a field-by-field list
Field typeDefault treatmentReason
Product, module and feature namesKeepDescribes what the work was about, not who did it
Version, build and release numbersKeepLinks the issue to a code change and a fix
Error codes and exception namesKeepThe core diagnostic signal for support and coding models
Steps to reproduce and workaroundsKeep, scrub embedded identifiersThe reasoning lives here, but identifiers sometimes hide inside
Resolution, root cause and outcomeKeepShows how the problem ended, the most useful part of the record
TimestampsKeep, or shift consistentlyOrder and elapsed time matter more than the calendar date
Customer and agent namesReplace with consistent placeholdersIdentity is not needed, but roles in the thread are
Emails, phone numbers, postal addressesRemoveContact details add no diagnostic value
Account IDs, tenant subdomains, IP addressesReplace with tokensCan point to a specific customer or person
Passwords, API keys, session tokensRemove, and rotate at the sourceA security exposure before it is a privacy question

Which details look risky but are usually safe to keep?#

Product terminology, plan tiers, browser and operating system versions, integration names and generic job titles usually look riskier to a detection tool than they are. A note that a customer on an enterprise plan hit a sync error between your app and NetSuite after a browser update names no person, contact detail or account.

The reverse also happens: some technical strings quietly carry identities. Check these before marking a field type safe.

  • File paths that include a user folder, such as a home directory named after an employee.
  • Hostnames and subdomains built from a customer's company name.
  • Log lines and stack traces that print an email address, session token or customer ID.
  • URLs with query strings that carry account numbers or password reset tokens.
  • Screenshots and attachments that show a name, avatar or inbox.
  • Email signatures pasted into ticket bodies, with titles, direct lines and office addresses.

How do you stop a detection tool from erasing technical detail?#

A detection tool erases far less technical detail once you give it an allow list and custom patterns built from your own product, loaded before the first full run rather than after the first complaint.

Then test on a small labeled sample. Track misses (personal details left in) and over-hits (technical terms removed) separately, because they are fixed differently: misses call for new patterns or a stronger model, over-hits call for allow list entries. Repeat until both are acceptable to the person approving release.

Most PII detection tools, from open-source SDKs to cloud services, support allow lists, deny lists and custom recognizers, though names and syntax differ. Keep the configuration file with the release record so the same rules can be rerun on the next export. Build the allow list and patterns from sources you already control:

  • Product, module and feature names from your documentation, admin console and codebase.
  • Names of the third-party products you integrate with, such as accounting, ERP or identity tools.
  • Your error code and exception formats, written as patterns so a code like a prefix plus four digits is never read as an account number.
  • Version and build number formats, including release train names.
  • The domain of your public documentation site, so links to help articles survive while customer subdomains do not.
  • Common job titles and role names used in the thread, such as agent, tier 2 or customer admin.

Why consistent placeholders keep a thread readable#

Consistent placeholders keep a thread readable because each participant keeps one label from the first message to the last. If the customer is CUSTOMER_1 at intake and CUSTOMER_1 after the escalation, a reader can still follow who asked, who answered and who approved the fix.

Deleting names outright, or issuing a new tag every time a name appears, breaks that structure. Keep role information where it matters, such as agent, engineer or customer admin, and keep the mapping table out of the delivery entirely, under the company's own control.

Illustrative: a warehouse software vendor fixes an over-redacted sample#

Illustrative: a fictional vendor of warehouse management software prepares a sample of Zendesk tickets, the linked Jira issues and the related Slack threads from its engineering channel. The first pass uses a general-purpose detection model at its strictest setting.

The review finds that the product's module names, several of which are also common first names, were replaced as people. Error codes in the product's own format were removed as account numbers, and every resolution note that quoted a release number was gutted. Customer names, meanwhile, survived inside log excerpts pasted into tickets.

The CTO loads the module glossary as an allow list, writes a custom pattern for the error code format, adds a rule for log excerpts and reruns the sample. The second version keeps the diagnosis and the fix in every reviewed ticket and has no customer names in the logs, so the CEO approves it for a buyer's review.

What to check before approving a redacted dataset#

Approval should rest on two reviews of the same sample: one looking for personal details that remain, and one looking for meaning that was lost. Most teams only run the first, which is how over-redacted datasets get released.

Use the signs below as a reviewer's guide. A sample that shows several over-redaction signs needs allow list work before anyone debates whether it is safe.

What to check before approving a redacted dataset
CheckUnder-redaction signOver-redaction sign
NamesReal names in signatures, logs or greetingsProduct modules replaced as people
NumbersPhone or account numbers in free textError codes and build numbers removed
LinksCustomer subdomains or tokens in URLsPublic documentation links deleted
OutcomeCustomer details quoted in the resolutionResolution notes empty or unreadable
ThreadPersonal email addresses in quoted repliesSpeakers impossible to tell apart

How SourceX approaches redaction scope#

SourceX treats redaction scope as part of Preparation in the SourceX five-step transaction, once the Rights step has settled what may be licensed at all. The aim is to lower the privacy burden, which the SourceX Enterprise Data Value Framework counts against net value, without stripping the domain expertise and AI utility that raise it.

The keep and remove list, the tool configuration and the review results go into the privacy record of the SourceX Evidence Packet. The supplier approves that record before anything is released, so the company can see exactly what was removed and what was kept.

Frequently asked questions

Should we keep our customers' company names?

Usually not. A customer's company name can identify the customer and may be confidential under your contract with them. Replace it with a consistent token such as ORG_1, and keep useful attributes like industry or plan tier as separate fields if your contracts and privacy review allow it.

Are IP addresses and device IDs personal data?

They can be. Many privacy laws treat online identifiers that can be linked to a person or household as personal information. In support and engineering records they rarely add diagnostic value, so replacing them with tokens is the safer default. Keep a note of whether an address was internal or external if that matters to the fix.

Does keeping product names expose trade secrets?

Product names alone rarely do, but internal architecture details, unreleased feature names and security findings might. That is a confidentiality question, separate from privacy. Review it with the product owner and decide which internal code names, roadmap items or vulnerability details should be generalized or excluded.

What about screenshots and attachments?

Treat them as a separate stream. Screenshots often show names, avatars, inboxes or customer records even when the ticket text is clean. Many sellers exclude images from a first package, or include only those that pass an image-level review, rather than relying on text redaction.

How can we tell if redaction went too far?

Read a sample of redacted tickets and ask whether an engineer who never saw the originals could still tell what broke, in which version, and how it was fixed. If the answer is often no, the rules are removing technical content and need allow list entries or narrower patterns.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify