Skip to content

Real conversational audio for voice agents and speech models

Voice agents and speech models need real two-party conversations recorded on the channels they will run on — phone and VoIP audio with accents, line noise, interruptions and hold time — with accurate transcripts and a record of what each call achieved. SourceX sources contact center and sales call recordings with transcripts, dispositions and QA scores, plus related support tickets, from established businesses, with recording consent and licensing rights reviewed and spoken personal data masked before delivery.

  1. Contact center call recordings and transcripts

    The core training and test audio for telephony agents: inbound service calls, ideally with each speaker on a separate channel, linked to transcripts, dispositions and QA scores that show how calls are opened, verified, resolved or transferred.

  2. Sales call recordings and transcripts

    Discovery and demo calls add a different conversational job — questioning, handling objections, agreeing next steps — often as wideband meeting audio with several participants, and with deal outcomes to judge each conversation against.

  3. Customer support ticket datasets

    Tickets are not audio, but they hold the task layer behind the calls: intents, policies, resolution steps and edge cases. Use them to build intent taxonomies and test scenarios for the agent's reasoning, separately from its recognition and turn-taking.

Why this data is hard to get

Open speech data has no task and no outcome

Public speech datasets are dominated by read text, talks, podcasts and short commands spoken to devices, and their accent mix rarely matches a given customer base. Few capture two people working through a real task on a phone line, and fewer still record whether the caller's problem was solved.

Simulated noise is not a bad line

Adding noise and codec effects to clean audio helps, but it cannot reproduce how people change their speech on a poor connection — speaking louder, repeating themselves, spelling words out, turning away to talk to someone else — or how jitter and dropped packets disrupt turn timing.

Permission to record is not permission to reuse

A recording announcement addresses whether a call may be recorded. Licensing the recording to train another company's model is a separate question, for the caller and for the employee on the other end, so each source's notices, jurisdictions and employee agreements are reviewed before audio is offered.

Anonymizing a voice damages the training signal

Voice conversion can disguise a speaker, but it also alters the acoustic detail a recognizer or speech-to-speech model is meant to learn. Speech training usually keeps the original voices, and some jurisdictions treat voiceprints as biometric data, so protection comes from rights review and license limits on speaker identification and voice cloning.

What a voice agent learns from each layer of a call

A voice agent learns three different things from a real call — hearing speech on the line, timing its turns, and completing the caller's task — and each needs different data. The acoustic layer covers accents, codecs, noise, and spelled-out alphanumerics such as postal codes and policy numbers; it needs audio with accurate transcripts. The conversational layer covers when to speak, how to handle barge-in, backchannels such as "mm-hm", long silences while a caller looks for a card, and repair when something was misheard; it needs per-speaker channels or precise timestamps, because the gaps and overlaps are the labels. The task layer covers getting the job done under the business's rules — verifying identity, finding the order, knowing when to transfer — and needs dispositions, QA scores and the linked tickets or knowledge articles.

Public data rarely covers more than the first layer. Contact center archives cover all three, which is why they are the place to start.

Platform transcripts are weak labels, not ground truth

Automatic transcripts from contact center platforms are useful for text tasks, but they are not ground truth for training or scoring a recognizer. Their errors cluster where a voice agent most needs accuracy — proper nouns, digit strings, code-switching and overlapping speech — so fine-tuning on them teaches the old system's mistakes, and scoring against them hides yours. Ask for a human-verified subset wherever recognition accuracy is the target, above all for evaluation sets. Agree normalization conventions for numbers, fillers, false starts and masked spans before verification starts, or batches transcribed at different times will not be comparable.

Scope from the deployment, not the dataset

A speech data request works best when it describes where the agent will run: narrowband telephony, VoIP or wideband meeting audio; the languages and accent regions it must serve; and the call reasons it will handle. Then list the labels you need — speaker channels, timestamps, dispositions, QA scores, transfer events. Ask for audio at its original sample rate and codec: a narrowband call carries nothing above 4 kHz, and upsampling it to 16 kHz does not restore the missing band.

Name every intended use, because uses clear differently. Voice generation is the sensitive case: a model that can reproduce a real customer's or employee's voice creates likeness and impersonation risks that a recognizer does not, so a data owner may permit recognition and dialogue training while excluding synthesis. Stating uses up front avoids a dataset that clears for one model and not the next.

On the sample, listen for clipped or leaky masking around spoken numbers, confirm that both channels are present and correctly labeled, and spot-check transcripts against the audio.

What good data looks like

  • Each speaker on a separate channel, or reliable diarization, with the original sample rate and codec documented for every file.
  • Transcripts with segment- or word-level timestamps, each marked as human-verified or machine-generated.
  • Call events aligned to the audio, such as holds, transfers, IVR steps and hang-ups, plus dispositions and QA scores where they exist.
  • Spoken identifiers masked in both audio and transcript, with each masked span typed and timestamped so it can be excluded from training loss.
  • Recording disclosures and consent basis documented per source and jurisdiction, with model training confirmed as a permitted use.
  • Coverage reported by language, accent region, channel type and call reason, so it can be compared with your production traffic.

Questions buyers ask

Can I license real call center recordings to train a voice agent?

Yes, provided the business that owns the recordings agrees and its recording disclosures and customer contracts permit that use. Calls handled by an outsourced contact center usually belong to the brand it serves, which then has to authorize licensing. SourceX checks ownership and rights before a dataset is offered. Whether matching audio can be found depends on which businesses hold it and agree to license it.

Is dual-channel or mono audio better for speech model training?

Dual-channel recordings, with each party on its own channel, are usually more useful. They give clean per-speaker audio for recognition, exact overlap and barge-in timing for turn-taking models, and diarization ground truth without extra labeling. Mono mixed audio still trains recognition but needs diarization and loses overlap detail. Some archives keep only mixed mono, so check channel layout, sample rate and codec on the sample.

Does masking spoken numbers hurt a voice agent's training?

It can, because account numbers, card numbers and dates of birth are exactly the strings a voice agent has to capture, and masking removes them from real audio. Spoken identifiers are located through transcript timestamps, silenced or toned out, and replaced with placeholders in the transcript. Marking each masked span with its type keeps the turn structure usable, and alphanumeric recognition can be trained separately on consented read speech or synthetic audio.

Can I license transcripts without the audio?

Yes. Transcripts with speaker turns and timestamps carry no voiceprint, so they are often easier to clear, and they suit dialogue policy, summarization, intent and QA models. Recognition, speech-to-speech and turn-taking models need the audio. Some teams license transcripts broadly and audio only for calls chosen for language, accent or channel coverage.

Can real calls be used to evaluate a voice agent?

Yes, and they are hard to replace for it. Held-out calls with human-verified transcripts measure word error rate on your target accents and channels, and calls with known dispositions become scenarios for a simulated caller, testing whether the agent reaches the same resolution. Keep evaluation calls out of every training run.

Tell us what you are building

Describe the model or agent, the tasks it must handle, and the volume, format and permitted use you need. SourceX will match it to partner data.

Updated 3 October 2026.

See if you qualify