Speech and audio data
Full-Duplex Conversation Data for Speech-to-Speech Models: Overlaps, Backchannels and Turn-Taking
Quick answer
A full-duplex conversational speech dataset gives each speaker a separate, sample-aligned audio track of a real, unscripted two-party conversation, plus time-stamped labels for turn starts and ends, overlaps, backchannels, interruptions and non-verbal vocalizations. Speech-to-speech models learn to listen while speaking most directly from audio where both sides stay audible during overlap. Open full-duplex sets are small, about 15 hours in one release [1], so production-scale training usually depends on licensed or commissioned two-party recordings with consent that covers AI use.
By SourceX Editorial · Updated
Why full-duplex models need one track per speaker
Full-duplex models need per-speaker tracks because they model the user stream and the agent stream in parallel, and a mixed-down mono file destroys that structure. A cascaded voice agent (voice activity detection, ASR, a text LLM, then TTS) only needs clean, turn-segmented speech, but a model that listens while it speaks must learn what each party produces at every frame. That includes the frames where both are talking, which is where backchannels, barge-in and overlapping turn transitions live.
In a mono mix, a backchannel such as "mm-hm" under the other speaker's sentence is either inaudible or inseparable. Source separation can recover some of it, but separation artifacts become training signal, and the model learns the separator's errors. The authors of the open full-duplex release made the same point in their own terms: dual-track recording is what makes overlaps, backchannels and laughter learnable, and fine-tuning on that data improved naturalness [1].
Practical consequences for a buyer:
- Native stereo or per-leg capture, not post-hoc separation. Call platforms that record each leg separately (agent and customer channels) are the closest real-world analog; see dual-channel call recordings for ASR and diarization.
- Sample-level alignment between tracks. A 40 ms drift between legs misplaces every overlap onset and teaches the wrong response latency.
- Low cross-talk bleed. Headset or telephone legs bleed less than two room microphones; ask for a measured bleed figure per session rather than an assurance.
File-level details such as sample rate, bit depth and codec are covered in audio file specs for speech datasets.
Which conversational events to label, and how precisely
Label the events the model must produce or react to: turn boundaries, overlaps, backchannels, interruptions (barge-in) and non-verbal vocalizations, each with start and end timestamps on the correct track. Word-level transcripts alone are not enough, because a transcript flattens simultaneous speech into a sequence and loses the timing that defines turn-taking behavior.
A workable taxonomy distinguishes events by function, not only by acoustics:
- Turn start and turn end per speaker, with the gap or overlap to the previous turn (floor transfer offset). Negative offsets are overlapping transitions; positive ones are gaps.
- Backchannel: short listener tokens ("yeah", "right", "mm-hm") that do not claim the floor. These are what a full-duplex agent should emit while the user keeps talking.
- Interruption or barge-in: an overlap where the second speaker claims the floor and the first yields. Mark whether the interrupted speaker stopped, and how many milliseconds later. This is the core signal for interruption handling.
- Non-competitive overlap: collaborative completions or simultaneous starts that resolve without a floor change.
- Non-verbal vocalizations: laughter, breaths, sighs, filled pauses ("uh", "um"), coughs. These carry timing cues and appear in the open full-duplex release as labeled phenomena [1].
- Silence and hold: long pauses where the floor stays with one speaker, which teach the model not to jump in.
Transcription convention matters as much as event labels. Clean-read transcripts delete fillers and false starts, which are exactly the cues turn-taking models use; specify verbatim conventions as described in verbatim vs clean transcription standards.
Why scripted and role-played dialogue underperforms
Scripted dialogue underperforms because actors reading or improvising to a brief rarely overlap, rarely backchannel naturally and leave unnaturally clean gaps between turns. The CASPER authors note that scripted dialogue dominates existing conversational speech data, which limits natural turn-taking examples, and they collected unscripted conversation instead: about 200 recorded hours, 158 hours of actual speech, with 102 hours in the first public release [2].
Role-play collections also skew toward polite alternation. A model trained on them tends to wait for silence before responding, which reads as sluggish in live use, and it has few examples of how a user actually interrupts a long answer. If you commission recordings, design the task so overlap happens: open-ended topics, real stakes, and instructions that permit interruption. The broader trade-off is covered in spontaneous conversational speech data.
Synthetic two-speaker audio generated by TTS has the same limitation in a sharper form: the overlap statistics are whatever the generator was told to produce. See real vs synthetic speech data for where augmentation helps and where it does not.
Where full-duplex data at scale comes from
At scale, full-duplex training data comes from three places: open research releases, commissioned two-party recordings, and licensed operational recordings where each party was captured on its own leg. Open releases are useful for pilots and evaluation, but at roughly 15 hours in the cited release [1] they are far below what pretraining or substantial fine-tuning of a speech-to-speech model typically uses.
| Source | Track separation | Naturalness of overlap | Main risk |
|---|---|---|---|
| Open full-duplex research sets | Dual-track by design | High, but small volume | License terms; limited languages and domains |
| Commissioned studio or remote sessions | Controlled per-speaker capture | Depends on task design | Cost per hour; performed rather than real behavior |
| Business calls recorded per leg (support, sales, scheduling) | Native agent and customer channels | Real, task-driven | Consent scope; telephony bandwidth; PII in audio |
| Mixed-down meeting or podcast audio | Usually none | Real | Separation artifacts; rights usually unclear |
Two-party business calls are a strong candidate because many contact-center platforms store agent and customer legs separately, and the conversations are goal-directed, with genuine interruptions and backchannels. The constraints are real: much of this audio is 8 kHz narrowband (see telephony vs wideband audio), and the original recording notice may not cover AI training. For the commercial category view, see SourceX's call center audio datasets and voice agent training data pages.
Consent, voiceprints and redaction in two-party audio
Two-party recordings need consent from both parties that covers the original capture and the AI training use, and voice itself can be regulated biometric data. California Penal Code section 632(a) prohibits recording a confidential communication without the consent of all parties [4], so a "this call may be recorded" notice is how many businesses obtain consent for capture, not evidence that training use was disclosed; section 632.7 separately covers calls involving cellular or cordless phones. Texas defines a voiceprint as a biometric identifier and requires notice and consent before capture for a commercial purpose [5].
Ask the supplier for the exact notice language played to callers, the date range it applied to, and how the customer leg was handled where the notice changed. Redaction also interacts with full-duplex training: beeps or silence inserted over names and account numbers cut through overlaps and backchannels on one track while the other track keeps talking. Specify how redaction is applied per track and how it is marked in the labels; the failure modes are described in audio redaction artifacts in speech model training, and voice-level de-identification in speaker anonymization for speech datasets.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How to evaluate turn-taking without hiding overlap
Evaluate overlap regions explicitly rather than excluding them. Diarization scoring has long used a no-score collar around speaker boundaries, which conveniently removes the hardest frames; the MISP 2022 audio-visual diarization challenge set no collar and scored overlapping speech [3]. For full-duplex systems, the boundary and overlap frames are the behavior you are buying the data to learn.
Hold out sessions, not segments, so the same speakers and conversations never appear on both sides of the split. Useful held-out measures include floor transfer offset distribution versus human reference, backchannel rate and placement, barge-in yield latency (time from user onset to agent stopping), and false take-over rate during user pauses. Audit labels on the test set as well: label errors are common even in widely used benchmarks, including audio ones, with an estimated average error rate of at least 3.3% across the audited test sets [6]. Broader quality and contamination checks are covered in the training data quality guide.
A request specification for full-duplex conversation data
A good request specifies the capture setup, event labels, timing tolerance, transcript convention, consent scope and delivery package in measurable terms. The template below is a starting point to adapt; packaging conventions for manifests and timing files are covered in packaging speech datasets, and diarization label formats in speaker diarization training data.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: full_duplex_conversation_audio
conversation:
parties: 2
style: unscripted, task-driven (support, scheduling, sales)
languages: [en-US]
min_session_minutes: 3
capture:
tracks: one per speaker, native (no source separation)
alignment: sample-aligned across tracks; max drift 10 ms per session
sample_rate: 16 kHz or higher preferred; 8 kHz telephony accepted and flagged
format: WAV/FLAC, PCM 16-bit, one file per track
bleed: measured cross-talk level reported per session
labels (per track, start/end in ms):
- turn_start / turn_end with floor_transfer_offset_ms
- backchannel (token text)
- interruption {claimed_floor: true|false, yield_latency_ms}
- overlap_noncompetitive
- nonverbal {laugh, breath, filler, cough}
- redaction {method: tone|silence|replace, reason_code}
transcript: verbatim, word-level timestamps, fillers and false starts kept
rights:
consent: all parties; notice text and dates supplied; AI training use covered
biometric: voiceprint handling stated per state of capture
splits: by session and by speaker; no speaker in more than one split
qa: double-annotated sample with inter-annotator agreement on event onsets
Before scaling, run a pilot on a few hours: check per-track alignment by cross-correlating shared audio events, re-label a random sample of overlaps, and confirm the floor transfer offset distribution looks like natural conversation rather than role-play.
How SourceX approaches full-duplex conversation requests
SourceX sources operational datasets from US companies and manages the commercial process, including licensing and ongoing purchases; support and sales histories are among the kinds of data in scope. Datasets are sourced on request rather than held in stock, so a request for per-leg call audio with consent covering AI use does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and no method is perfect. You can describe the data you need on the SourceX buyers page, and the wider speech and audio data guide and AI data hub cover adjacent specs.
Sourcing full-duplex conversational speech data
Describe the tracks, labels, consent scope and volume you need, and SourceX looks for US businesses that hold matching recordings; every release is approved by the supplying company and nothing is contracted until a supplier agrees. Each dataset is delivered under a license defining records, uses, term and delivery. Describe the full-duplex conversation data you need.
Sources
- arXiv, "Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis" (2025). https://arxiv.org/pdf/2509.04093
- arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
- arXiv, "The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition" (2023). https://arxiv.org/pdf/2303.06326
- California Legislative Information, "California Penal Code section 632 (eavesdropping on or recording confidential communications)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.