Skip to content

Software companies

Redacting customer company names and deal details from B2B records

By SourceX Editorial · Updated

Short answer

To redact customer company names and deal details from B2B records, build an entity list from your CRM and billing system, replace each customer with one consistent pseudonym across every system, and catch indirect identifiers such as email domains, subdomains, URLs and deal names. Remove or band contract values and discounts, and shift dates rather than leaving them exact.

Key takeaways

  • General-purpose PII detectors focus on people and contact details, so customer companies need an entity list you build from your own systems.
  • One customer should map to one pseudonym across support, CRM, engineering and chat records, or the links that give the data value are lost.
  • Customer names hide in domains, subdomains, tenant slugs, channel names, file names and Jira components as often as in plain text.
  • Deal values, discounts and custom terms are removed or banded; outcomes such as won, lost or churned are kept.
  • The mapping between real names and pseudonyms stays with the company and is never part of the delivered data.

Why do company names need a different method?#

Customer company names need their own method because standard redaction tools are built mainly to find people, contact details and account numbers. A detector can find an email address, but it cannot know that a short abbreviation in a Jira title is your largest customer, or that a subdomain in a log line identifies a tenant.

The obligation is different too. Customer names are usually protected by confidentiality and publicity clauses in customer contracts rather than by privacy law alone, so the test is whether a reader could tell which company a record is about. Automated tools still help, but Presidio's documentation, for example, warns that automated detection cannot guarantee it finds all sensitive information.

Step 1: build the entity list#

The entity list is a table of every customer, prospect and partner that could appear in the records, with every way each one is written. Building it from systems of record is faster and more complete than trying to spot names in text.

Give each entity one row with its canonical name, variants, domains, IDs and the pseudonym it will receive. Ask account managers and support leads to add nicknames; engineers and agents often use shorthand that never appears in the CRM.

  • CRM accounts from Salesforce or HubSpot, including parent and child accounts, former names and lost prospects.
  • Billing customers from Stripe, NetSuite or your billing system.
  • Product tenants: workspace names, subdomains, tenant slugs and tenant IDs.
  • Help desk organizations from Zendesk, Intercom or similar tools.
  • Email domains of every contact, including subsidiaries and regional domains.
  • Partners, resellers and integration vendors whose names appear in deal and support records.
  • Known variants: abbreviations, misspellings, ticker-style short names and internal nicknames.

Step 2: assign consistent pseudonyms#

Consistent pseudonyms keep records linked: the same customer becomes the same token in its support tickets, CRM notes, Jira issues and Slack threads, so a reader can still follow one account's story. Generic tags such as a plain customer placeholder destroy that linkage and much of the value with it.

Normalize variants before replacing them, so that every spelling of one customer maps to the same token. Coarse attributes such as industry or size band can be added to the pseudonym if they help the buyer, but only if the combination cannot single out a customer.

Step 2: assign consistent pseudonyms
MethodHow it worksKeeps linkageWatch out for
Lookup tableA company-held table maps each name to a token such as a customer letter-number codeYesThe table is sensitive and never ships
Keyed deterministic encryption or hashingThe same input always yields the same token under a secret key; Google's Sensitive Data Protection offers deterministic encryption for thisYesVariants must be normalized first or they get different tokens
Generic tagsEvery customer becomes the same placeholderNoLoses the ability to follow one account
Invented company namesEach customer gets a realistic fictional nameYesAn invented name may match a real company; check each one

Step 3: catch domains, URLs and IDs#

Indirect identifiers reveal a customer as clearly as the name itself, and they sit in places text search for names will miss. Build patterns from the entity list's domains, slugs and IDs, and run them across every record type in scope.

Replace email domains with a pseudonym domain under a reserved documentation domain, so the shape of an address survives without pointing anywhere real. Apply the same token used for the company name, so a domain and a name for one customer still match after redaction.

  • Email addresses at customer domains, including subsidiaries.
  • Tenant URLs and subdomains, such as a customer slug in front of your product's domain or inside an account path.
  • Webhook URLs, API endpoints and storage bucket names that embed a tenant slug.
  • Slack Connect and shared channel names that carry the customer's name.
  • File names on attachments and documents, such as statements of work named after the customer.
  • Jira components, labels and Git branch names created for a specific account.
  • IP allowlists, SSO identity provider names and customer-specific configuration keys.

Deal details: remove, band or keep#

Deal details are treated by type: values and terms that a competitor or the customer would consider confidential come out, while the decisions and outcomes that make the record useful stay in. Decide the treatment once per field and apply it everywhere.

Deal details travel further than the CRM. Look for them in quote and order form attachments, deal-desk channels in Slack, renewal discussions in support tickets where customers mention what they pay, and engineering issues that cite a contractual service level. Apply the same field rules wherever the detail appears.

Deal details: remove, band or keep
DetailTreatment
Contract value, annual recurring revenue, price per seatRemove, or band into broad ranges you define
Discounts, credits and concessionsRemove
Signature, renewal and churn datesShift consistently per customer, or reduce to the quarter
Opportunity and deal namesReplace with the customer pseudonym and the deal stage
Custom contract terms and service levelsRemove the specifics; keep a note that a custom term existed
Contact names and titlesPseudonymize the person and keep the role
Outcome and reason: won, lost, renewed, churnedKeep; this is the signal that matters

Where does redaction usually fail?#

Redaction usually fails through combinations rather than names. A record that keeps a customer's industry, region, size band and a distinctive product configuration can point to one company even with its name removed, especially in a vertical market with a small customer base. Check that each pseudonym's remaining attributes still describe several plausible companies.

Exact dates are the second route. A go-live or churn date that matches a press release or a published case study identifies the customer, which is why dates are shifted per customer rather than left exact.

The third route is free text written by people who knew the account well: an agent mentioning the customer's annual conference, its founder by first name or its headquarters city. Patterns rarely catch these, so a reviewer who knows the accounts should read a sample from every system in scope.

Illustrative: a supply chain software vendor redacts linked records#

Illustrative: a fictional supply chain visibility software company plans to license support-to-fix histories that join Zendesk tickets, Jira issues and Salesforce notes. The COO's team builds an entity list from Salesforce accounts, the product's tenant table and Zendesk organizations, and assigns pseudonyms from a company-held lookup table.

Pattern searches catch tenant subdomains in pasted log lines and customer names in Slack Connect channel names. Deal values are removed, renewal dates are shifted per customer, and outcomes are kept. A sample review finds that engineers refer to one customer by a nickname in Jira titles, so the variant is added and the run repeated. The delivered records still let a reader follow one account from complaint to fix, without naming it.

How SourceX approaches customer names#

SourceX handles customer name redaction in the Preparation step of the SourceX five-step transaction, with the supplier approving the method. The pseudonym mapping stays with the supplier; SourceX does not need it and the buyer never receives it.

The SourceX Evidence Packet's privacy record notes the entity sources, the pseudonym method, the patterns used and the sample review results, so the buyer can rely on the redaction without seeing what was redacted.

Frequently asked questions

Is a customer company name personal data?

Usually not by itself, but it can be. A sole trader or a small firm named after its owner may identify a person, and company names often appear next to individual contacts. Either way, customer contracts usually treat customer identity as confidential, so redact it regardless of how privacy law classifies it.

Can we keep customers that appear publicly on our website?

Treat them like everyone else. Permission to show a logo in marketing does not usually extend to licensing that customer's records, and keeping some names while removing others makes the remaining pseudonyms easier to guess. Consistent treatment is simpler to explain.

Do prospects and lost deals need redacting too?

Yes. A prospect's name in a sales note or a lost deal in the CRM can be just as confidential as a customer's, and sometimes more sensitive, because the prospect may have shared plans during evaluation. Include prospects, lost opportunities and churned customers in the entity list.

Should we redact our own company and product names?

Usually not. Your product names, feature names and error codes are part of what makes the records useful, and they identify the supplier, not the customer. Redact them only if the license is meant to be anonymous as to supplier, which is uncommon.

How do we check that the redaction worked?

Search the finished copy for every name, variant, domain and ID on the entity list, then have someone who knows the accounts read a sample of records from each system. Add any missed variants to the list and rerun until both checks come back clean.

What happens to new customers after the export?

Nothing, if the export has a fixed end date. The entity list only needs to cover customers who appear in the records being licensed. If the license includes later updates, refresh the entity list from the CRM before each new delivery.

Sources

  • Google's Sensitive Data Protection API supports de-identification transforms including deterministic encryption (CryptoDeterministicConfig) and date shifting by a random number of days. Source
  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify