Privacy and preparation
Do you need to redact the audio, or only the transcript?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
You need to redact the audio only if you deliver the audio. For most licensing deals the safer default is redacted, checked transcripts. If a buyer genuinely needs recordings, mute every span the transcript redacted, treat the speaker's voice itself as an identifier, and review recording consent and biometric rules with counsel before any audio leaves the company.
Key takeaways
- Default to transcripts; deliver audio only when the buyer's use requires sound.
- Every redaction in the transcript needs a matching muted span in the audio.
- A voice can identify a person even when no name is spoken.
- Recording disclosures and biometric rules are reviewed with counsel before audio is licensed.
- Caller ID, file names and call detail records need their own pass.
The short answer: transcripts by default#
The default for licensing call recordings is to deliver redacted transcripts and keep the audio in house. A transcript carries much of what buyers use from business calls: the problem described, the questions asked, the resolution and the outcome. Audio adds tone, pacing, accents and background sound, and it adds the speaker's voice, which is far harder to de-identify than text.
Audio belongs in scope only when the buyer's use depends on sound, such as speech recognition, turn-taking, voice agents or call-quality models. Ask what the buyer will do with recordings, and write that permitted use into the license so the scope cannot drift later.
Transcript only or audio too: a decision table#
The buyer's stated use decides the format. When the use is vague, start narrow, because a dataset can grow in a later delivery but released audio cannot be recalled.
| Buyer's need | Deliver | Why |
|---|---|---|
| Conversation flow, intents and resolutions | Transcripts only | The text carries the content |
| Summaries or agent-assist training | Transcripts with speaker turns | Turn labels keep structure without the voice |
| Speech recognition on industry vocabulary | Audio with muted spans, paired with transcripts | The model needs sound aligned to text |
| Voice agents or tone and emotion modeling | Audio, with voice treatment agreed in writing | Voice characteristics are central to the use |
| Unclear, or everything you have | Transcripts first; audio later if justified | Scope can widen, but audio cannot be pulled back |
Getting the transcript right comes first#
Transcript redaction comes first in either path, because muted audio is only as good as the transcript that located the spans. A transcript-only delivery also lives or dies on that pass, so it deserves the same care you would give a ticket or email export.
Call transcripts carry identifiers that ticket scanners rarely see. Agents introduce themselves by first name at the start of every call, speaker labels may hold agent names or extensions, and callers read out account numbers, addresses and order references in fragments. Replace speaker labels with consistent role tags such as Agent and Caller, and apply the same person placeholders used elsewhere in the dataset.
Speech-to-text errors cut both ways. A misheard surname may slip past detection, while a misheard product name can be redacted by mistake. Review a sample of the worst-quality calls, not only clean ones, since noisy lines and heavy crosstalk are where both kinds of error cluster.
If audio is in scope, how do you redact it?#
Audio redaction works from a time-aligned transcript. Speech-to-text output with word-level timestamps shows where each name, address, account number or card number was spoken, so those spans can be muted in the recording. Without timestamps, redaction becomes manual listening, which is slow and easy to get wrong.
Spelled-out names and numbers read in pieces are the classic misses. Transcription may render them as single letters or scattered digits that a PII detector does not recognize, so reviewers should listen for them specifically.
- Transcribe with word-level timestamps and speaker labels.
- Run PII detection on the transcript, then review a human sample.
- Map each finding to its time span, with a small pad on each side so partial syllables do not survive.
- Replace each span with silence or a tone, and use one choice across the whole dataset.
- Re-transcribe the redacted audio and confirm the removed words no longer appear.
- Listen to a sample, focusing on spelled names, numbers read digit by digit and crosstalk.
The voice is an identifier too#
The voice itself identifies a person, even when every name has been muted. Coworkers, customers and family members recognize voices, and voiceprint technology can match them at scale. Where recordings are used to create voiceprints or similar templates, biometric privacy laws may apply, so counsel should assess them before audio is scoped.
The Illinois Biometric Information Privacy Act is the best-known example. It lets a prevailing party recover liquidated damages of $1,000 per negligent violation or $5,000 per intentional or reckless violation, and in Rosenbach v. Six Flags (2019) the Illinois Supreme Court held that a person need not show actual injury beyond the violation to sue. A 2024 amendment limits recovery to one violation per person when the same identifier is collected repeatedly by the same method, but the exposure remains significant.
Each technical option changes the data. Pitch shifting is weak and often reversible. Voice conversion, which re-synthesizes speech in a different voice, protects identity better but alters the acoustic detail some buyers want. Excluding certain speakers, such as employees who were never told about this use, is sometimes the simplest route. Agree the approach with the buyer and record it in the license.
Consent, disclosures and the rest of the file#
Recording disclosures were usually written for quality assurance and staff training, not for licensing to an AI developer. Some states require consent from every party to a recording, and the wording of the original disclosure matters. Pull the IVR announcement text, agent call scripts, employee recording acknowledgments, the phone or transcription vendor's terms and any customer contracts that mention recordings, and have counsel compare them with the planned use.
Payment details need their own check. PCI DSS Requirement 3.3.1 says sensitive authentication data, which includes card verification codes, is not retained after authorization, so calls where a customer read out a security code are a payment-security problem before they are a licensing one. Exclude them rather than relying on muting.
Audio files also carry material a transcript pass does not catch. Give each item below its own check.
- Caller ID, phone numbers and customer names in file names, folder paths and call detail records.
- Voicemail greetings that state a person's name or direct number.
- Keypad tones entered during payment, which can reveal card digits.
- Hold music and recorded messages, which may be third-party content.
- Background voices, such as a television, a child or a coworker on another call.
Illustrative: an industrial distributor's inside sales calls#
Illustrative: a fictional industrial distributor records inside sales and customer service calls in its cloud phone system and stores transcripts alongside them. A developer building a parts-ordering assistant asked for the call history.
Counsel reviewed the recording disclosure and the buyer's stated use. The assistant needed conversation flow and product vocabulary, not voices, so the company licensed redacted transcripts with speaker turns and kept every recording in house. Customer names, account numbers and spoken card details were removed, and call detail records were reduced to date, duration band and outcome. A later request for audio was parked until the buyer could explain why sound was needed.
How SourceX scopes recordings#
SourceX scopes recordings by settling the transcript-or-audio question first, before any recording is prepared. Rights covers recording disclosures, employee notices and any biometric questions, and Preparation covers transcript redaction and, only where the supplier agrees, muted audio and voice treatment.
The permitted use in the SourceX Evidence Packet states whether audio is included and what it may be used for. The privacy record holds the padding rule, the silence or tone choice and the re-transcription results, so the buyer can confirm what was removed and the supplier can answer later questions without replaying recordings. The supplier approves both before Delivery.
Frequently asked questions
Is a tone better than silence for muted spans?
Either works if it is applied consistently. A tone marks clearly where something was removed, which helps reviewers and buyers. Silence is less jarring but can be mistaken for a pause in speech. Whichever you choose, state it in the dataset card so the buyer's team knows how to treat those spans.
Do vendor-generated transcripts change the rights picture?
They can. If a phone system or transcription vendor produced the transcripts, check its terms for limits on how outputs may be used and whether the vendor kept any rights. That review belongs in the rights step, before a buyer sees a sample.
Can we delete the audio after licensing transcripts?
Retention is a separate decision governed by your own policies, contracts and any legal holds. Licensing transcripts neither requires you to keep the audio nor to delete it. If you plan to delete, confirm that no dispute, regulatory request or agreement needs the recordings first.
What about recorded video meetings?
Video adds faces, screens and shared documents on top of voices. The same rule applies with more force: deliver transcripts unless the buyer needs the picture, and treat video as its own scope decision with its own review.
Are employee voices treated differently from customer voices?
They can be. Employees may have received different notices, and employment rules differ from consumer rules. A common approach is to include agent speech in transcripts but exclude employee audio unless employees were told about the use and counsel agrees.
Should the delivered transcripts be corrected by hand?
Not usually. Buyers often want transcripts as the system produced them, errors included, or prefer to run their own transcription on audio. Hand correction is costly and can introduce new identifiers. Fix only what redaction requires, and note the transcription source and any corrections in the dataset documentation.
Sources
- BIPA (740 ILCS 14/20) allows liquidated damages of $1,000 per negligent violation or $5,000 per intentional or reckless violation; Illinois SB 2979, signed August 2, 2024, limits recovery to a single violation per person when the same biometric identifier is collected repeatedly by the same method. Source
- In Rosenbach v. Six Flags Entertainment Corp., decided January 25, 2019, the Illinois Supreme Court held that a person need not allege actual injury beyond a BIPA violation to be an aggrieved party entitled to sue. Source
- PCI DSS v4.0 Requirement 3.3.1 states that sensitive authentication data, including card verification codes, is not retained after authorization, even if encrypted. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.