Skip to content

Speech and audio data

Dual-Channel Call Recordings for ASR and Diarization: Why Stereo Audio Matters

Quick answer

Dual-channel (stereo) call recordings put the agent and the customer on separate audio channels, so every word arrives already attributed to a speaker. That gives you near-free diarization references, clean per-speaker ASR targets, and intact overlaps and backchannels for turn-taking models. A mono mix forces you to buy diarization labels or run a model to recover who spoke when, and overlapped speech is lost for good. When you source call audio, specify channel layout, channel mapping and how separation was produced, not just "stereo."

By SourceX Editorial · Updated

This page covers one spec decision. For the full requirements template, see how to specify contact-center audio data; for the wider category, the speech and audio data buyer's guide.

What channel separation gives an ASR or diarization model

Channel separation turns speaker attribution from a modeling problem into metadata. In a native two-leg recording, channel 0 is one party and channel 1 is the other, so a voice-activity detector run per channel yields a who-spoke-when reference without a human labeler. Diarization challenge organizers have used the same idea, building reference RTTM files by combining per-speaker close-talking annotations rather than labeling a single mixed track [1].

Three model families benefit directly:

  • ASR fine-tuning. Each channel can be segmented and transcribed as single-speaker audio, which avoids training on overlapped segments where the "correct" transcript is ambiguous. You can also weight agent and customer speech separately, which matters because customer audio carries more device, network and accent variability.
  • Speaker diarization. Per-channel activity gives you reference segments for training and for a held-out test set scored with diarization error rate (DER). See evaluating diarization on your own audio for collars and overlap scoring.
  • Full-duplex and turn-taking models. Separate tracks preserve overlaps, backchannels ("mm-hm," "right") and laughter that dual-track datasets were built to capture [2]. Full-duplex dialogue systems such as Moshi model the user's and the system's speech as parallel streams instead of chaining VAD, ASR, a text model and TTS [3], so training data in that shape needs each speaker's audio kept apart.

If your roadmap includes voice agents trained on real conversational audio, mono archives are a weak foundation: you cannot reconstruct an interruption you never recorded separately.

Mono vs stereo call audio: what you lose in a mixdown

A mono mixdown destroys information you cannot buy back later. When two legs are summed, overlapped speech becomes a single waveform, and separation models only approximate the originals. The cost shows up in three places.

PropertyNative dual-channelMono mix
Speaker attributionFrom channel index plus mapping metadataNeeds human diarization labels or a diarization model
Overlap and interruptionsPreserved on both tracksSummed; transcripts usually drop or garble one speaker
BackchannelsAudible and timestamped per speakerOften masked by the louder party
Per-speaker redactionMask customer PII without touching agent audioMasking removes both speakers in the window
Diarization reference costLow (per-channel VAD plus spot checks)High (full manual RTTM labeling)
Fit for turn-taking modelsDirectPoor

When diarization is partial and audio is mono, the labeling bill moves to you. Ask for the exact hours covered by speaker labels and transcripts, not the headline total.

Redaction is a frequently missed benefit. Spoken card numbers and account details usually come from the customer leg, so per-channel masking can remove them while keeping agent speech intact. The trade-offs of tones, silence and transcript tags are covered in how audio redaction affects speech model training and redacting spoken PII from call recordings.

How telephony platforms actually produce stereo files

Stereo is a recording setting, not a property of the call, and platform defaults vary. As of October 2026, Twilio's Voice API takes a RecordingChannels parameter of mono or dual and defaults to mono; in dual mode a two-party call keeps each leg on its own channel [4]. The download request has its own RequestedChannels parameter, so a supplier can hold dual-channel originals and still export downmixed files [4].

Conferences behave differently. On Twilio, a dual-channel conference recording puts the first participant on the first channel and mixes everyone else onto the second [4]. A warm transfer or a supervisor barge therefore lands a third voice on the "customer" or "agent" channel, and channel index no longer equals speaker.

Channel order is not standardized either. Twilio Flex documents the customer on the left and the agent on the right, and notes that recording can start on the customer leg, so queue music and hold time may be captured [5]. Amazon Connect can record customer and system audio during IVR and customer, agent or both during agent interactions, with up to two recordings per contact [6]. Two 2018 AWS announcements describe opposite left/right assignments for Connect recordings [7][8]. The practical lesson: never infer mapping from the platform name, require it as metadata per file.

Failure modes to check before you accept a stereo dataset

Most stereo problems are invisible in a file listing and obvious in a 30-second listen. Build these checks into sample review:

  1. Reconstructed, not native, separation. Some suppliers run source separation on mono archives and ship the result as "stereo." Ask whether channels came from the recording system or from post-processing, and look for separation artifacts such as musical noise and bleed during overlaps.
  2. Crosstalk and echo leakage. Agent headsets and softphone echo cancellation are imperfect, so the agent's voice often appears faintly on the customer channel. Measure energy on the "silent" channel while the other party speaks; leakage creates false speech segments in per-channel VAD.
  3. Retention-policy downmixes. Some archives are summed to mono after a retention window to save storage, so the same supplier may hold a stereo recent window and a mono back catalog. Ask for the split by date range.
  4. Mixed sample rates and codecs. Telephony legs are usually narrowband 8 kHz, sometimes stored as G.711 µ-law or compressed to MP3 or Opus. Each leg can differ if one side is a WebRTC softphone. See telephony vs wideband audio and audio file specs for speech datasets.
  5. IVR, bot and hold segments. The customer channel may contain IVR prompts, a voicebot or hold music before an agent joins. These need segment labels, or your "customer speech" contains synthetic audio.
  6. Transfers and conferenced third parties. Interpreters, supervisors and transferred agents share a channel with someone else. Require events with timestamps so you can exclude or relabel those spans.
  7. Channel swaps across a corpus. A platform migration or configuration change can flip left/right halfway through the date range. Spot-check mapping per source system and per month, not once.

Artifact: channel-layout specification for a call-audio request

Put channel layout in writing before any sample review. The table and record below are a starting point for a request or a data-delivery schedule; adjust fields to your platform mix.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldRequired value or question
Channel count2 per file; mono files listed separately with hours
Separation originnative_recording or post_separation (name the tool if the latter)
Channel mapPer file: ch0 and ch1 role (agent, customer, ivr, bot, mixed)
Source systemRecording platform and configuration, with date ranges
Sample rate and codec per channele.g. 8 kHz, 16-bit PCM WAV; disclose any transcoding
Leg eventsTransfers, conference joins, holds, IVR-to-agent handoff with timestamps
Leakage checkMethod and threshold used to flag crosstalk on sample files
Labels coverageHours with transcripts, hours with diarization, hours with neither
RedactionWhich channel was masked, method, and replacement (silence, tone, noise)
Consent basisNotice or consent mechanism per recording population

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "call_id": "c-000184",
  "file": "audio/c-000184.wav",
  "channels": 2,
  "sample_rate_hz": 8000,
  "encoding": "pcm_s16le",
  "separation_origin": "native_recording",
  "channel_map": {"ch0": "customer", "ch1": "agent"},
  "segments": [
    {"start": 0.0, "end": 41.2, "ch0_role": "ivr_and_customer", "ch1_role": "silent"},
    {"start": 41.2, "end": 388.6, "ch0_role": "customer", "ch1_role": "agent"},
    {"start": 212.4, "end": 255.0, "event": "supervisor_joined", "ch1_role": "mixed"}
  ],
  "transcript": "transcripts/c-000184.json",
  "rttm": "rttm/c-000184.rttm",
  "redaction": {"channel": "ch0", "method": "tone", "spans": 3},
  "leakage_db_ch1_into_ch0": -32.5
}

Pair this with a manifest format from packaging speech datasets and the label conventions in speaker diarization training data and RTTM specs.

Stereo files do not change the consent analysis, but they make the parties easier to isolate, which cuts both ways. Under California Penal Code section 632, recording a confidential communication requires the consent of all parties [9], so the customer leg needs a documented notice or consent mechanism just as the agent leg does. Separate channels also make a single speaker's voice easier to extract as a clean, enrollable sample, so treat the customer channel as biometric-adjacent data in your privacy review.

Ask suppliers how the recording notice was delivered (IVR prompt, terms, agent script) and whether it covers the hold and queue audio captured before an agent joins. For the full checklist, see call-recording consent for AI training and the state pages such as Washington call-recording rules.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

When mono call audio is still worth licensing

Mono is acceptable when your target is ASR robustness rather than attribution, and when the price reflects the labeling you will add. Large mono back catalogs can still help acoustic adaptation to narrowband telephony, especially combined with a smaller stereo set used for diarization references and turn-taking. Decide the ratio from your evaluation plan: if DER on overlapped speech or interruption handling is a release criterion, the stereo portion must be large enough to test it. How many hours of audio you need to fine-tune ASR helps size each portion.

The same reasoning applies outside contact centers: sales call recordings and multimodal meeting recordings often come from platforms that record per-participant tracks, which are worth requesting over a composite mix.

How SourceX handles requests for channel-separated call audio

SourceX sources operational datasets, including support and sales histories, from US companies on request; nothing is held in stock and a request does not guarantee a match. You describe the data you need, including channel layout, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and personal details are removed or replaced before delivery, with the method recorded and a sample checked. You can describe your call-audio requirements to SourceX, and the call center audio datasets page explains the broader category. For terminology, see diarization.

Source dual-channel call recordings for your ASR and diarization work

SourceX serves AI teams wherever they are based and manages the process from finding a supplier through assessing data and permissions, agreeing a license and ongoing purchases. Nothing is contracted until a supplier agrees, and terms are agreed per deal. Share your channel-layout spec at sourcex.si/buyers.

Sources

  1. arXiv, "Summary of the DISPLACE Challenge 2023: DIarization of SPeaker and LAnguage in Conversational Environments" (2023). https://arxiv.org/pdf/2311.12564
  2. arXiv, "Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis" (2025). https://arxiv.org/pdf/2509.04093
  3. Kyutai (arXiv), "Moshi: a speech-text foundation model for real-time dialogue" (2024). https://arxiv.org/html/2410.00037v1
  4. Twilio, "Recording resource (Voice API)". https://www.twilio.com/docs/voice/api/recording
  5. Twilio, "Enabling Dual-Channel Recordings (Flex Insights)". https://static1.twilio.com/docs/flex/developer/insights/enable-dual-channel-recordings
  6. Amazon Web Services, "When, what, and where for contact recordings in Amazon Connect". https://docs.aws.amazon.com/connect/latest/adminguide/about-recording-behavior.html
  7. Amazon Web Services, "Coming soon: Amazon Transcribe to Identify Speakers Based on Channels" (2018). https://aws.amazon.com/es/about-aws/whats-new/2018/07/coming-soon-amazon-transcribe-to-identify-speakers-based-on-channels/
  8. Amazon Web Services, "Amazon Transcribe Can Now Identify and Label Transcripts Based on Audio Channels" (2018). https://aws.amazon.com/about-aws/whats-new/2018/08/amazon_transcribe_can_now_identify_and_label_transcripts_based_on_audio_channels/
  9. California Legislature, "California Penal Code section 632 (eavesdropping on or recording confidential communications)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data