Skip to content

Privacy and preparation

Spoken numbers and spelled-out emails: why call transcript redaction misses them

By SourceX Editorial · Updated

Short answer

Call transcript redaction misses personal details because speech-to-text writes them in forms detectors were not built for: digits as words, emails spelled letter by letter, and card or account numbers split across speaker turns. The fix is to normalize spoken numbers, rejoin spelled letters and scan across neighboring turns before any detector runs, then review samples by hand.

Key takeaways

  • Run detection on normalized text: convert number words, 'oh' for zero and 'double' patterns into digits before pattern matching.
  • Rejoin single letters, phonetic spellings and the words 'at' and 'dot' so spelled-out emails become detectable strings.
  • Scan a window of neighboring speaker turns, because callers and agents often split one number between them.
  • Redacting the transcript does not redact the audio, the call summary or the CRM note created from the same call.
  • Automated detectors need a human sample review; even well-built open-source tools say they cannot guarantee full detection.

Why does call transcript redaction miss spoken numbers?#

Call transcript redaction misses spoken numbers because most detectors look for digit patterns, and speech-to-text engines often write numbers as words. A card number read aloud may arrive as a run of number words, broken by pauses, filler words and the agent saying 'okay' halfway through.

Pattern rules for phone, card and account numbers expect a continuous string of digits with familiar separators. Checksum tests, which confirm a card number before flagging it, cannot run on a sequence of words. Named-entity models help with names and places but are weak on long numbers spoken in fragments.

The result is a transcript that passes an automated scan yet still holds the details billing and service calls exist to collect: account numbers, dates of birth, street addresses and callback numbers.

Which leak patterns are specific to call transcripts?#

The leak patterns specific to call transcripts come from how people speak and how speech engines transcribe them. Each pattern below needs its own fix, and a phone-heavy company will usually find several in one archive.

Review the list against a handful of real calls from your own archive before choosing tools. Every speech engine has its own habits with numbers, letters and punctuation, and a change of engine or settings partway through your history changes the patterns too.

Which leak patterns are specific to call transcripts?
Leak patternHow it appears in the transcriptWhy detectors miss itFix
Digits as wordsCard or account number written as a run of number wordsPatterns and checksums expect numeralsNormalize number words to digits before detection
Zero as 'oh' and repeated digits'Oh' for zero, 'double seven', 'triple two'Normalizers built for prose skip these formsAdd speech rules for oh, double and triple
Spelled-out emailSingle letters separated by spaces, then 'at' and 'dot com'Email patterns need an @ sign and no spacesCollapse letter runs and map 'at' and 'dot' to symbols
Phonetic spelling'B as in boy' or radio alphabet words for each letterEach word looks harmless on its ownMap phonetic words back to letters, then rescan
Split across turnsCaller gives part of a number, agent repeats it, caller finishes itEach turn holds too few digits to matchScan a sliding window of neighboring turns
Read-back confirmationAgent repeats the full address or number to confirm itThe second copy is formatted differently from the firstRedact every copy once one copy is found
Mis-transcribed namesA surname transcribed as a common word or a placeName models see an ordinary nounMatch against the name on the linked customer record
Spoken dates'March third, eighty-four' in answer to an identity questionDate patterns expect numeric formatsNormalize spoken dates and flag answers to verification prompts

How do you fix digit and letter detection before scanning?#

Digit and letter detection improves most when you build a separate scanning copy of each transcript before any detector runs. The scanning copy is never delivered; it exists so patterns, checksums and models see personal details in forms they recognize, and every hit is mapped back to the original words.

Keep word-level timestamps and turn identifiers through each step. Without them, a hit found in the scanning copy cannot be traced to the exact span in the original transcript or in the audio.

  • Normalize spoken numbers: convert number words, 'oh', 'double' and 'triple' into digits, and drop fillers such as 'um' and 'okay' that sit inside a number.
  • Rejoin spelled letters: merge runs of single letters and phonetic words into one token, and turn 'at' and 'dot' into symbols when they sit between letter runs.
  • Normalize spoken dates and street numbers so they reach the same patterns as typed forms.
  • Build a turn window: join each turn with the turns just before and after it, so split numbers appear as one string.
  • Run detectors on the scanning copy: patterns with checksums for card numbers, models for names and places, and context cues such as 'card', 'date of birth' and 'last four'.
  • Match known values: search for the name, phone, email and address held on the customer record linked to the call, including near spellings.
  • Map every hit back to the original transcript span and redact it there with a typed placeholder, such as card number or email.

Where else does the same call leak details?#

The same call leaks details in several places besides the transcript text: the audio file, the automatic call summary, the agent's wrap-up note in the CRM or help desk, and the call metadata.

Audio is the hardest. Redacting a transcript changes nothing in the recording, so a dataset that includes audio needs each redacted span silenced or replaced with a tone, using the word timestamps kept from the scanning copy. Many companies decide to license text only, which also removes the voice itself as an identifier.

Summaries and wrap-up notes are written by a different system or by the agent, so they need their own pass. Metadata such as caller ID, the dialed extension, the agent's name and the exact call time can identify a person even when the words are clean.

Speaker separation matters too. When the speech engine attributes the caller's words to the agent, rules that scan only the customer channel miss them, so scan both sides of every call.

How far should automated detectors be trusted?#

Automated detectors should be trusted as a first pass, not a final answer. Presidio, an open-source SDK for finding and masking personal data that began as a Microsoft project and is now community-governed under the Data Privacy Stack organization, mixes named-entity recognition with pattern matching, rule logic and context-aware checksums, yet its documentation cautions that 'there is no guarantee that Presidio will find all sensitive information' and recommends layering further systems and protections on top.

That warning applies to every detector, commercial or open source, and more strongly to speech. A detector tuned on typed text has rarely seen a number spoken as words with 'oh' and 'double' mixed in. Plan a human review of sampled calls drawn across call types, agents, years and speech engine versions.

Treat each reviewer finding as a rule gap, not a one-off correction. Add the pattern, rerun the whole archive and sample again until reviews stop finding new leak types.

Which numbers should stay in the transcript?#

Numbers that describe the work should stay in the transcript, because they carry much of a call's value to an AI developer. Error codes, model and part numbers, quantities, appointment windows and ticket references show how problems are diagnosed and resolved, and blanking every digit erases them.

The rule is to redact by context and type rather than by format. A digit string after 'my card number is' goes; the same length of string after 'the model number on the unit is' stays, unless it is a serial number tied to one customer's equipment.

Which numbers should stay in the transcript?
Content heard on the callDefaultReason
Card, bank and account numbers, card security codesRedactFinancial identifiers, often read in fragments; PCI DSS says card verification codes are not kept after authorization
Dates of birth and identity answersRedactUsed to verify one specific person
Phone numbers and spelled emailsRedactDirect contact identifiers
Street addresses, gate and alarm codesRedactLocate a home or grant physical access
Equipment serial numbersRedact or tokenizeCan link back to one customer's account
Error codes, model and part numbersKeepDescribe the problem and the fix
Quantities, appointment windows, ticket referencesKeep; tokenize internal IDsOperational detail with little identifying power

Illustrative: a home services call center finds what its first pass missed#

Illustrative: a fictional plumbing and HVAC company records calls in its cloud phone system and books jobs in ServiceTitan. Before licensing transcripts of booking and troubleshooting calls, its operations team ran a standard PII detector over the text and saw few hits.

A sample review told a different story. Callers read card numbers for deposits as number words, spelled email addresses for invoices letter by letter, and gave gate and lockbox codes using 'oh' for zero. Agents read back every address. The detector had caught almost none of it.

The COO approved a second pass with number normalization, letter rejoining, a window of neighboring turns and matching against the address and phone on each linked customer record. Gate and alarm codes became their own redaction type. The team licensed text only, dropped call summaries that copied customer notes, and kept model numbers and error codes. A fresh sample review found the earlier leak patterns closed before the dataset went to approval.

How SourceX handles call transcripts#

SourceX treats call transcripts as a Preparation task inside the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Before anything leaves the supplier, the redaction method, the leak patterns tested and what the sample reviewers found are filed in the SourceX Evidence Packet's privacy record.

The supplier approves the prepared dataset and its scope, including whether audio is included at all. Nothing is shared during the initial assessment; the fit check asks which phone system and speech engine produced the transcripts and which call types they cover, not for the calls themselves.

Frequently asked questions

Does pausing call recording during card payments solve the problem?

Pausing the recording while a caller reads a card number helps, but only for that moment. Agents forget to pause, callers repeat numbers later in the call, and the pause rarely covers dates of birth, addresses or spelled emails. Card security codes need particular care: PCI DSS Requirement 3.3.1 says sensitive authentication data, including card verification codes, is not retained after authorization, even if encrypted. Treat pause-and-resume as one control and still scan every transcript and any summary made from it.

Should we license audio or only transcripts?

Text-only release is simpler to prepare and removes the voice, which can identify a speaker on its own. Audio adds tone, pacing and background context that some developers want, but every redacted span must also be silenced in the recording and checked by ear. Decide based on buyer demand and the review effort you can support.

Does re-transcribing old calls with a newer speech engine help?

It can. Newer engines may format numbers and emails as digits and symbols that standard detectors catch more easily. A new transcript also brings new errors, so run the same normalization, turn windows and sample review on it, and record which engine and settings produced the delivered text.

Who should do the sample review?

People who know how your calls run, such as senior agents or supervisors who recognize your verification scripts, product names and local street names, working from a written checklist. Pair them with whoever owns the redaction rules so each miss becomes a rule change. Reviewers should work in a restricted environment because they see unredacted text.

Do chat and email transcripts have the same problem?

Partly. Typed channels rarely contain number words, but customers still split numbers across messages, paste screenshots and space out email addresses to get past filters. Turn windows and matching against the linked customer record apply to chat and email as well; number normalization matters mainly for speech.

Sources

  • Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization that combines named-entity recognition, regular expressions, rule-based logic and checksums with context. Source
  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee that it will find all sensitive information, and that additional systems and protections should be employed. Source
  • Presidio moved from a Microsoft-owned project to an independent, community-governed open-source project under the GitHub organization Data Privacy Stack, recorded in release 2.2.363 dated 2026-06-28, and remains MIT-licensed. Source
  • PCI DSS v4.0 Requirement 3.3.1 states that sensitive authentication data is not retained after authorization, even if encrypted; sub-requirement 3.3.1.2 covers card verification codes. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify