Skip to content

Privacy, de-identification and sensitive data

Redacting spoken PII from call recordings: transcript alignment, audio masking and card data

Quick answer

To redact PII from call recordings for AI training, detect identifiers in a word-timestamped transcript, map each detected span to audio time with padding, replace that audio with silence or a tone, replace the transcript text with typed tags or surrogates, and verify both by listening to a sample. Payment card numbers and security codes must be gone from audio and text. Content redaction does not remove the speaker's voice, which remains a biometric identifier unless it is separately anonymized.

By SourceX Editorial · Updated

This page is the buyer's acceptance spec: what to require in the redaction method, the delivered files and the QA evidence. For the supplier-side walkthrough, see how to de-identify sales call recordings. For the wider privacy picture, start at the privacy and de-identification hub.

Why call audio needs two redactions, not one

Call recordings carry identifiers in two parallel layers, so a redaction that touches only the transcript leaves every identifier audible in the WAV file. A transcript-only pass is a common defect in contact-center datasets: the JSON looks clean, but the agent still reads back "4-1-1-1, 1-1-1-1" at 00:03:12. Any buyer who trains ASR, voice agents or speech-to-speech models consumes the audio, so the audio is the record that matters.

The two layers also fail differently. Transcripts fail on detection (the model or regex misses a span), while audio fails on alignment (the span was found but the mute landed 300 ms late). Your acceptance criteria need a measurement for each.

Detect on the transcript, using more than regex

Detection should run on the transcript, because text-based detectors are generally more mature and easier to measure on identifiers than audio classifiers. The practical stack is a pattern layer for structured values (card numbers with Luhn checks, SSNs, phone numbers, emails), a named-entity model for names and addresses, and increasingly an LLM pass for context-dependent spans. Research on LLM-based redaction finds that LLMs catch PII that pattern matching misses, while introducing their own error modes that need measuring [2]. Open-source tools such as Presidio are a reasonable baseline, but the project itself warns that ML detection cannot guarantee complete coverage [4].

Spoken data defeats text patterns in specific ways:

  • Spelled and chunked digits. "Four one one one, one one one one" or "forty-one eleven" must be normalized before a Luhn check can fire. ASR inverse text normalization (ITN) settings change whether digits appear as words or numerals.
  • Corrections and repeats. "My zip is 9-0-2... sorry, 9-0-1-2-1" puts identifiers in two places.
  • Spelled names and emails. "That's S as in Sam, M-I-T-H" and "j dot smith at..." evade NER models trained on written text.
  • ASR errors. A misrecognized digit breaks the Luhn check, so a 16-digit string that fails Luhn is still a candidate, not a pass.
  • Cross-turn context. The agent asks "and the last four?" and the customer answers "two two nine one" in the next turn; detection must read the dialogue, not single utterances.

For a comparison of detector families and their leakage risk, see LLM-based PII redaction vs NER and regex.

Map transcript spans to audio with forced alignment

Audio redaction is only as accurate as the word timestamps behind it, so require forced alignment rather than the coarse segment times most ASR APIs return. A forced aligner (for example a CTC-based or Montreal Forced Aligner pass) gives start and end times per word; each detected span becomes an interval from its first word's start to its last word's end.

Then pad the interval. Alignment is least reliable on exactly the content you are redacting: fast digit strings, spelled letters and overlapping speech on 8 kHz telephony audio. A pre-pad and post-pad of a few hundred milliseconds, set by measurement on the supplier's own data rather than by habit, catches most boundary drift. Overlap is where mono recordings fail: when the agent and customer talk at once, muting the customer's card number also mutes the agent. Dual-channel recordings let you mute only the speaking channel; the trade-offs are covered in dual-channel stereo call recordings for ASR and diarization and 8 kHz telephony versus wideband audio.

Choose the audio replacement for the model you are training

The replacement signal is a modeling decision, and the right choice depends on the downstream task. Silence is clean but can teach a voice-activity detector that a turn ended. A 1 kHz tone ("bleep") is unambiguous for QA listeners but is an out-of-distribution sound an ASR model may learn to transcribe as noise tokens. Low-level pink or room noise matched to the call's noise floor keeps acoustic continuity for ASR and turn-taking models.

Whatever the choice, keep the duration. Shortening audio by cutting spans breaks every timestamp downstream (diarization RTTM files, sentiment labels, hold markers) and teaches models unrealistic turn timing. Keep a marker in the transcript at the same position, such as [CREDIT_CARD] or a surrogate value, so summarization and QA models still learn that a card number was given. The trade-off between tags and realistic surrogates is covered in masking vs surrogate replacement.

Payment card data: what PCI expects from recordings

Card data is the one category where the requirement comes from an industry standard rather than a privacy statute. The PCI Security Standards Council's information supplement on telephone-based payments addresses call recordings that capture card details, and it reflects the PCI DSS rule that sensitive authentication data, such as the card verification code, must not be retained after authorization, including in recordings [1]. The supplement is informational and does not override PCI DSS itself [1].

For a buyer, the practical consequences are:

  • Security codes and PINs must be absent from audio and transcripts, with no exceptions for "internal" training use.
  • Primary account numbers should be removed entirely; a training corpus has no legitimate need for a real PAN, even truncated.
  • Upstream controls change the risk. Calls taken with DTMF masking or pause-and-resume recording may never have captured card digits, but partial pauses (agent resumes recording mid-number) are a known failure. Ask which control was in place and for which date range.
  • Expiry dates and cardholder names travel with PANs and should be redacted in the same pass.

What survives redaction: the voice itself

Removing every spoken identifier does not make a recording anonymous, because the voice is an identifier in its own right. The VoicePrivacy Challenge treats speaker identity as a separate attribute from spoken content and evaluates anonymization by whether an attacker's speaker-verification system can still link anonymized speech to the original speaker, alongside ASR and emotion-recognition utility measures [3]. HIPAA lists voice prints among the biometric identifiers [5], and under GDPR Recital 26 data is anonymous only if a person is no longer identifiable by means reasonably likely to be used [7].

This creates a real choice. Voice anonymization (voice conversion or pitch and formant transformation) reduces linkability but degrades exactly the acoustic signal ASR, emotion and voice-agent models need [3]. Many buyers training speech models accept unconverted voices, treat the corpus as pseudonymized personal data rather than anonymous, and control it through license terms, access limits and a ban on speaker re-identification. Health-related calls, such as provider-to-payer phone calls, need a HIPAA method on top: Safe Harbor or Expert Determination [6]. Because voice prints are a listed identifier, unconverted audio typically pushes health calls toward Expert Determination; see Safe Harbor vs Expert Determination.

Acceptance spec for redacted call recordings

A buyer's acceptance spec should name each identifier class, the required treatment in both layers and the evidence the supplier must deliver. The table below is a starting template; adjust classes to your jurisdiction and use case.

Illustrative example: invented to show structure; it does not describe an available dataset.

Identifier classTranscript treatmentAudio treatmentEvidence required
Card PAN, expiry, CVV/CVCTyped tag [CREDIT_CARD], [CVV]Masked full span plus paddingZero Luhn-valid 13 to 19 digit strings in transcripts; listening sample of all flagged calls
SSN, account and member numbersTyped tag or format-preserving surrogateMaskedRegex and LLM recall report on a labeled sample
Names (customer, agent, third parties)Surrogate names, consistent per callMaskedPer-class precision and recall
Addresses, ZIP, date of birthTag or generalized valueMaskedSame as above
Email, phone, spelled stringsTagMaskedSpelled-out cases included in the test set
VoiceNot applicableUnchanged or anonymized, as licensedStated explicitly in the license and data card
Silence, holds, tonesPreserve [HOLD], [DTMF] markersPreserve durationTotal duration per file unchanged

Pair the table with a per-span redaction manifest so you can audit any mute against its transcript source. JSON Lines works well for this, one record per redacted span:

Illustrative example: invented to show structure; it does not describe an available dataset.

{"call_id":"c_000184","channel":"customer","span_class":"CREDIT_CARD","text_tag":"[CREDIT_CARD]","word_start_s":192.41,"word_end_s":199.87,"pad_pre_s":0.25,"pad_post_s":0.35,"audio_fill":"matched_noise","detector":"luhn_regex+llm","aligner":"ctc_forced","qa_listened":true}

QA: how to prove the audio is clean

Verification must include listening, because text metrics cannot detect a mute that missed by half a second. Ask for a stratified listening sample weighted toward high-risk segments (any call with a payment intent, any span flagged as spelled characters, overlap regions) plus a random sample, and require the fail count and how failures were fixed. A missed fragment such as the final two digits of a card number is still a fail.

Re-run the transcript detector on a fresh ASR pass over the redacted audio. If a new ASR transcript of the redacted WAV recovers digits or names, the audio mask missed. Then apply a formal sampling plan to the delivered set; residual PII audit sampling covers sample sizes and acceptance thresholds, and speech dataset manifest packaging covers how the manifest, segment timing and redaction files should ship together.

How SourceX handles call recording requests

SourceX sources operational datasets, including support and sales call histories, from US companies on request; nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phone and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Health records require HIPAA de-identification by Safe Harbor or Expert Determination. You can review the contact-center call recording category and call recordings overview, then describe your spec on the SourceX buyer page.

Get redacted call recordings scoped to your spec

SourceX sources contact-center and sales call recordings from US companies on request, with every release approved by the supplying company and personal details removed or replaced before delivery under a written license. Describe the recordings, channels and redaction requirements you need on the SourceX buyer page.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Frequently asked questions

Is muting the audio enough if the transcript is redacted?

No. Muting and transcript redaction are both required, and they must agree. A transcript tag with no matching audio mask means the identifier is still audible, and an audio mask with no transcript tag misaligns text and speech for ASR training.

Should redaction run before or after diarization?

After. Detection benefits from speaker labels (a digit string from the customer after "card number?" is high risk), and dual-channel or diarized audio lets you mute one speaker without muting overlapping speech from the other.

Does a redacted call recording count as anonymous data?

Usually not. With the voice intact, the recording remains linkable to a person by speaker verification [3], so treat it as pseudonymized personal data and control it through license terms and access limits.

Sources

  1. PCI Security Standards Council, "Information Supplement: Protecting Telephone-Based Payment Card Data, v3.0" (2018). https://listings.pcisecuritystandards.org/documents/Protecting_Telephone_Based_Payment_Card_Data_v3-0_nov_2018.pdf
  2. arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
  3. arXiv, "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677
  4. Data Privacy Stack (GitHub Pages), "Presidio - Data Protection and De-identification SDK" (2026). https://data-privacy-stack.github.io/presidio
  5. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  6. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  7. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation), Recital 26" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data