Skip to content

Contact center call recordings and transcripts for AI training

A contact center call recording dataset is a collection of recorded customer service and support calls, usually with speaker-labeled, time-aligned transcripts and call metadata such as queue, holds, transfers, handle time, disposition and QA score. SourceX sources these archives from companies running platforms such as Genesys Cloud, NICE CXone, Five9 or Amazon Connect, reviews recording consent and ownership, and removes spoken card numbers and other personal data from audio and transcripts before delivery.

Dataset manifest

Sourced to your spec
What it is
Recorded service and support calls with diarized transcripts, dispositions and QA scores
Typical systems
Genesys Cloud, NICE CXone, Five9, Amazon Connect, Talkdesk, Verint
Typical history
Varies by partner; audio retention policies set how far back recordings go
Modality
Telephony audio, time-aligned transcripts and structured call metadata
Delivery formats
Agreed per order; FLAC or WAV audio with JSON transcripts is common
Preparation
Names, account details and spoken card numbers removed from audio and transcript
Licensing
Permitted use agreed per license; BPO-held calls need the client brand's authorization
Availability
Depends on retained audio and adequate recording notices; not guaranteed

What a delivery contains

Fields vary by source system and are fixed per order. A typical delivery includes:

FieldTypeWhat it holds
call_idstringPseudonymous call ID shared by the audio file, transcript and metadata.
started_attimestampCall start in UTC; date shifting can be agreed where exact dates are sensitive.
directionenumInbound, outbound or scheduled callback.
queuestringThe queue, skill or IVR-selected reason that routed the call to an agent.
languagestringBCP 47 language tag, with code-switching flagged.
audioobjectFile reference, original codec, sample rate, channel count and which channel holds the agent.
segmentsarrayDiarized transcript turns with speaker role, start and end times and confidence.
segments[].wordsarrayWord-level timestamps and confidences, where the transcription engine produced them.
transcript_sourceenumPlatform ASR, third-party ASR or human-corrected, so you know what each transcript is worth as a label.
eventsarrayIVR path, holds, transfers, conferences and recording pauses, timed against the audio.
handle_timeobjectTalk, hold and after-call work seconds, the parts that make up average handle time (AHT).
dispositionobjectWrap-up code, agent note and whether the customer called back within the partner's repeat-contact window.
qaobjectScorecard items, form version, auto-fail flags and how the call was picked for review.
redactionsarrayTime ranges silenced or masked in audio and transcript, with the category of data removed.

Example record

{
  "call_id": "call_9b2e41",
  "started_at": "2024-02-06T15:31:08Z",
  "direction": "inbound",
  "queue": "billing_past_due",
  "language": "en-US",
  "audio": { "file": "audio/call_9b2e41.flac", "source_codec": "g711_ulaw",
             "sample_rate_hz": 8000, "channels": 2,
             "channel_map": { "0": "agent", "1": "customer" }, "duration_s": 182.4 },
  "transcript_source": "platform_asr",
  "segments": [
    { "start": 0.42, "end": 4.10, "speaker": "agent", "conf": 0.94,
      "text": "Thanks for calling, this is [AGENT_NAME]. Who am I speaking with?" },
    { "start": 4.85, "end": 12.30, "speaker": "customer", "conf": 0.81,
      "text": "Hi, it's [CUSTOMER_NAME]. My service got cut off and I, uh, thought autopay..." },
    { "start": 13.02, "end": 19.64, "speaker": "agent", "conf": 0.92,
      "text": "Sorry about that. Before I look, can you confirm your date of birth?" },
    { "start": 20.10, "end": 22.75, "speaker": "customer", "conf": null,
      "text": "[DOB_REDACTED]" },
    { "start": 90.06, "end": 98.51, "speaker": "agent", "conf": 0.90,
      "text": "Thanks for holding. The card on file expired, so this month's autopay failed." },
    { "start": 99.20, "end": 107.35, "speaker": "customer", "conf": 0.88,
      "text": "Okay, can I just pay now? The number is [PAYMENT_CARD_REDACTED]." },
    { "start": 171.33, "end": 178.90, "speaker": "agent", "conf": 0.93,
      "text": "That's gone through, and your service should reconnect within the hour." }
  ],
  "events": [
    { "t": 0.0, "type": "ivr", "path": ["main_menu", "billing", "agent"] },
    { "t": 24.80, "type": "hold", "duration_s": 65.0 }
  ],
  "handle_time": { "talk_s": 117, "hold_s": 65, "after_call_work_s": 44 },
  "disposition": { "code": "payment_taken_service_restored", "repeat_contact_7d": false },
  "qa": { "form": "billing_v4", "selected_by": "random_sample", "verification": 1,
          "empathy": 0.5, "resolution": 1, "auto_fail": false },
  "redactions": [
    { "start": 20.10, "end": 22.75, "category": "date_of_birth", "audio": "silenced" },
    { "start": 103.10, "end": 107.35, "category": "payment_card", "audio": "tone_replaced" }
  ]
}

Synthetic record for illustration. Field names, structure and format are agreed per order.

What AI teams use it for

Fine-tune telephony speech recognition

Narrowband, recompressed audio with crosstalk, hold music and accented speech is harder for ASR than clean wideband recordings. Human-corrected calls let you measure and close that gap in your own domain.

Train and evaluate voice agents

Real calls show turn-taking, interruptions, identity checks and how good agents calm an upset caller. Held-out calls with known outcomes show whether an agent would have solved the same problem.

Automate QA scoring

Calls scored against a versioned rubric teach a model to apply it to every call, not only the sample a QA team has time to review.

Draft summaries and wrap-up codes

Agent notes and dispositions written after each call are direct targets for summarization and auto-disposition, and after-call work time shows what that step costs today.

Predict transfers and repeat calls

Transfer events and callback flags mark the calls the first agent could not resolve, which is the signal routing and escalation models learn from.

Use-case guides: Voice agents and speech models, Customer support agents, Healthcare administration AI

What makes this data valuable

Dual-channel audio

Agent and customer on separate channels give reliable speaker attribution, even when they talk over each other.

Human-corrected transcripts

Even a modest verified subset makes word error rate measurable on real calls.

Resolution evidence

Dispositions, transfers and repeat-contact flags show whether the call actually fixed the problem.

Event timelines

IVR steps, holds, transfers and recording pauses are timestamped against the audio.

Acoustic range

Mobile, landline, speakerphone and in-car audio across accents and languages.

Linked cases

Calls joined to the support case or account history behind them show what the agent could see.

What open speech corpora miss about service calls

Open speech corpora are mostly read speech, audiobooks, podcasts and broadcast audio, and the classic telephone corpora record recruited participants chatting about assigned topics, not customers trying to get something fixed. A real service call differs on every axis. It crosses narrowband telephony and is often recompressed for storage; caller and agent talk over each other; IVR prompts, hold music, speakerphones and car noise sit under the speech. The words that matter most are the hardest to recognize: spelled surnames, plan names, error codes and account numbers read out in chunks.

Behavior matters as much as acoustics. Callers start mid-thought, answer a different question from the one asked, go quiet while they look for a bill and lose patience on a second transfer. Text-to-speech dialogues rarely reproduce that, and any outcome attached to them is whatever the generator decided. A recorded call joined to its disposition, QA score and repeat-contact flag shows what a good resolution sounded like under one company's actual policies, and what a failed one sounded like.

How call archives change over the years

A multi-year archive is rarely one consistent dataset. A platform migration can swap the channel that holds the agent, change the codec, rename queues and replace the disposition list, so each era needs its own checks and a mapping table for changed codes. QA forms are revised too, which is why scores travel with their form version. Calls recorded before pause-and-resume or keypad masking was adopted need much heavier payment redaction than later ones, so sample redaction quality era by era. Retention schedules can also delete audio while transcripts and metadata survive, leaving the oldest years text-only.

Some labels are noisier than they look. Agents pick wrap-up codes under handle-time pressure and may choose the first plausible option, so cross-check dispositions against transfers, callbacks and the transcript before using them as targets. Redacted spans need consistent markers in audio and text, so that a model neither learns to invent card numbers nor treats inserted silence or tones as speech.

What to check before licensing

  • Establish who owns the recordings. A BPO that records calls for several brands on its own platform needs each brand's written authorization, and the outsourcing contract's data-use clause has to allow it.
  • Map recording announcements by jurisdiction and year, and check whether notices or privacy policies disclosed any use beyond internal quality and training.
  • Listen to sampled calls rather than only reading transcripts, to confirm that card numbers, dates of birth and security answers are gone from the audio too.
  • Ask how calls were chosen for QA review. Reviews triggered by complaints or new-hire coaching skew negative and are not a random sample.
  • Profile language, accent, call reason and device mix on the sample against the callers your model will serve.
  • Check employee-side rules as well. Agents are on every recording, and in some countries works-council agreements or employee notices limit secondary use.
  • Decide in the license whether speech synthesis, voice cloning or speaker recognition is permitted, since reproducing a voice raises consent questions that transcription does not.

How licensing works through SourceX

  1. 1

    Define

    Send the domain, modality, volume, format, timeline and permitted use you need.

  2. 2

    Source

    SourceX identifies businesses that hold matching data and are open to licensing it.

  3. 3

    Qualify

    Fit, rights and quality are checked, and you review samples before committing.

  4. 4

    License

    Scope, permitted use, exclusivity, price and obligations are agreed in writing.

  5. 5

    Deliver

    Approved data is prepared, de-identified where required and transferred securely.

Questions buyers ask

Are the transcripts human-verified or machine-generated?

Usually machine-generated. Archived transcripts tend to come from the contact center platform's speech analytics or a third-party ASR engine, so they carry that engine's errors on names, numbers and accents. Some partners also hold human-corrected subsets from QA or analytics work. The transcript_source field marks which is which; spot-check word error rate on a few calls before treating machine transcripts as labels.

Did callers agree to their calls being used for AI training?

Rarely in those words, so the basis for licensing is reviewed before anything is scoped. Some jurisdictions require every party's consent to record a call, others only one party's. The familiar "this call may be recorded for quality and training purposes" notice was usually written with internal use in mind, so qualification looks at the announcements, privacy notices and jurisdictions involved, and calls without an adequate basis are excluded.

Is a caller's voice treated as biometric data?

In some jurisdictions it can be. Several laws treat voiceprints, or voice data processed to identify a person, as biometric data with stricter consent and retention rules. Redacting words from a transcript leaves the voice intact, so raw audio stays identifying. The usual options are transcripts only, or audio with license limits such as no speaker identification and no voice cloning.

How are card numbers spoken during calls handled?

Card numbers, expiry dates and security codes are removed from the audio and the transcript before delivery. PCI DSS prohibits storing card security codes after authorization, and the PCI Security Standards Council has said this applies to audio recordings too. Many contact centers pause recording or take card details by keypad with tone masking, but older calls, missed pauses and agents reading numbers back can still leave card data in an archive.

What audio formats, channels and sample rates should I expect?

Often narrowband telephony audio sampled at 8 kHz; VoIP and WebRTC calls can carry wideband audio at 16 kHz or more. Some platforms record agent and customer on separate channels, while others store one mixed channel that needs diarization. Archives are often kept in compressed codecs, so ask for the original files or lossless conversions of them rather than audio re-encoded to another lossy format.

Can I request specific languages, accents or call reasons?

Yes, as part of your spec, though supply depends on which partners hold matching calls. Language and accent mix follow a partner's customer base and where its agents sit, so a multilingual service desk and a single domestic center produce very different distributions. List languages, locales, call reasons and noise conditions in the request, then check the sample before committing.

Evaluating this data for procurement?

Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.

Request dataset diligence

Tell us what your models need

Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.

Updated 3 October 2026. Own data like this? See how companies license it to AI developers.

See if you qualify