Skip to content

Speech and audio data

How to Specify Contact-Center Audio Data: A Requirements Template for Buyers

Quick answer

A contact-center audio request should state measurable requirements before any supplier conversation. These are total and labeled hours, call types and verticals, languages and accents, date range, distinct agent and caller counts, channel layout, sample rate and codec, transcript standard, diarization and timestamp format, and outcome metadata. It also needs the redaction method, consent and notice evidence by jurisdiction, and the license scope. A spec written this way makes quotes comparable and screens out research-only corpora early.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why a call-audio spec differs from a generic data request

A call-audio spec has to fix acoustic, conversational and legal properties that a generic data request leaves open. The general structure in how to write a data request for suppliers still applies. But two suppliers can both say "1,000 hours of support calls" and deliver very different things. One may ship mono 8 kHz G.711 mixdowns with clean-read transcripts from three agents. The other may ship dual-channel recordings from 400 agents with verbatim, diarized transcripts.

Those differences decide whether the data helps ASR fine-tuning, a voice agent, speech analytics or summarization. They also decide whether it is legal to use at all. The wider sourcing picture is in the speech and audio data buyer's guide. This page covers the requirements document itself.

Volume, call mix and population: the fields that set coverage

Coverage is set by how hours are spread across call types, verticals, speakers and time, not by the headline hour count. Ask for hours broken down by call type: support, sales, collections, scheduling, billing disputes, retention, and inbound versus outbound. Ask for the same breakdown by vertical, such as telecom, insurance, healthcare scheduling, utilities and retail. Collections and healthcare calls bring their own regulatory overlays, so name them explicitly or exclude them.

Put a cap on how much any one speaker can contribute. A corpus where 20 agents produce most of the audio teaches a model those 20 voices. Request distinct agent and caller counts, plus a maximum share of hours per agent. State a date range too, because IVR prompts, product names and script language drift over time. For guidance on volume targets, see how many hours of audio you need to fine-tune ASR.

Channel layout, sample rate and codec: the acoustic contract

The acoustic fields should describe the audio as captured, not as converted. Telephony audio usually arrives at 8 kHz narrowband over G.711 μ-law or A-law, or at 16 kHz over wideband codecs such as G.722 or Opus on VoIP legs. Upsampled 8 kHz audio saved as 16 kHz WAV still contains no content above 4 kHz. So ask for the original capture rate and codec, plus every transcoding step. The trade-offs are covered in telephony versus wideband audio for ASR.

Channel layout matters most for diarization. Dual-channel recordings, with the agent on one channel and the customer on the other, give near-free speaker attribution and show overlap cleanly. Mono mixdowns force you to rely on model-based diarization. Specify which you need, and whether the left and right channel mapping is consistent across the corpus. See dual-channel call recordings for ASR and diarization and audio file specs for speech datasets.

Transcript standard, timestamps and diarization: the label contract

The transcript standard must be stated as a rule set, not just a label like "human transcribed." Vendors describe call-center datasets by transcript type, time-stamped segments and speaker labels [1]. Your spec should define the following:

  • Verbatim or clean. Verbatim keeps disfluencies, false starts and fillers. Clean drops them. Name which you want and how hesitations, numbers, spelled-out letters and code-switching are written.
  • Tags. List the required tags for overlap, crosstalk, unintelligible speech, hold music, IVR prompts and redacted spans.
  • Timing. Say whether you need segment-level or word-level timestamps, and the timing tolerance you will accept.
  • Speaker turns. Request turns in RTTM so you can score them with standard tools [9].
  • Labeled share. State what share of total hours must be transcribed, and whether transcripts come from one pass or from double transcription with adjudication.

Clean or professional transcripts can still serve as training targets, with caveats explained in training ASR on non-verbatim transcripts. Plan acceptance testing before you sign, using transcript accuracy acceptance testing.

Realism criteria: real calls, not role-play

A production-ready call corpus should be specified by what live calls contain. That means holds, transfers, IVR menus, crosstalk, background noise, speakerphone and Bluetooth audio, and dropped or abandoned calls. One vendor argues that models trained mostly on clean audio fail on live calls [2]. That claim is self-reported, but the failure mode is real: scripted role-play underrepresents interruptions and agent screen-reading pauses.

Write realism into the spec as measurable fields. Examples include the share of calls containing a hold or transfer, a minimum per-call overlap ratio, and an SNR distribution. Require a flag on every call that is staged, simulated or synthetic. If you want transfers and warm handoffs, ask whether transferred legs are stitched into one file or delivered separately with a shared call ID.

Outcome metadata turns audio into training signal for analytics, summarization and agent-assist. Useful fields include disposition codes, wrap-up notes written by agents, QA scorecard results, call-driver tags, handle time, transfer and escalation flags, and links to CRM or ticket outcomes such as resolved, refunded, churned or deal stage. Ask how each field was produced. A disposition chosen by an agent at wrap-up is noisier than a ticket status set by a later system.

Agent desktop actions recorded alongside the call are a separate category, covered in contact-center desktop activity paired with call transcripts. If sales outcomes drive the use case, the sales call transcript datasets page covers that variant.

The compliance section should ask for evidence, not reassurance. Call recordings carry spoken names, addresses, account numbers and often card numbers. PCI SSC guidance on telephone-based payment card data addresses how card data in call recordings should be handled [5]. Ask suppliers to describe the redaction method for both audio and transcripts: silence, tone or noise masking; entity tags in transcripts; how spans are aligned between the two; and the measured miss rate on an audited sample. Some marketplace listings say PII is removed before delivery [4]. Your spec should require a written method, not a checkbox. Redaction artifacts also change model behavior, as covered in how audio redaction affects speech model training and redacting spoken PII from call recordings.

Consent evidence must match where the parties were. California requires all-party consent to record confidential communications [6]. So ask for the recording notice script, IVR disclosure audio and the date ranges each applied. You also need evidence that the supplier's agreements and notices permit reuse for AI training, not just quality assurance. The call recording consent checker and call-recording consent for AI training go deeper.

Voice is a biometric issue if anyone extracts speaker embeddings. Illinois BIPA lists a voiceprint as a biometric identifier [7], and Texas law does the same, with notice and consent required before capture for a commercial purpose [8]. If calls come through a BPO, require proof that each end client authorized the reuse. See sourcing AI training data through BPOs. If any calls involve patient health information, require HIPAA de-identification by Safe Harbor or Expert Determination [11].

License scope: state it before the first call

License scope belongs in the first paragraph of the request, because it disqualifies more candidates than any acoustic field. Name the uses you need: commercial model training, fine-tuning, evaluation only, or redistribution of derived models. Also name the term, the user population and whether the data can be used by affiliates or contractors. Public call-center corpora are often research-only. One academic call-center corpus, for example, is released under CC BY-NC 4.0, which rules out commercial training [3]. An audit of dataset hosting sites found that licenses are frequently missing or wrong [10], so ask for the license that governs the specific files, not a platform tag. If you want licensed audio from operating US businesses rather than public corpora, you can submit a buyer request to SourceX; the call center audio datasets page explains that category.

The requirements template

Use the template below as the body of your request. Mark each field as "must," "preferred" or "informational" so suppliers can respond to each line.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample requirementPriority
Intended useCommercial fine-tuning of a streaming ASR model and held-out evaluationMust
Total hours / labeled hours2,000 h audio; at least 600 h with verbatim transcriptsMust
Call typesSupport 50%, billing 20%, scheduling 15%, collections 0%, other 15%Must
VerticalsTelecom, insurance, utilities; exclude healthcare clinical callsPreferred
Languages / accentsUS English; at least 15% Spanish or code-switched callsMust
Date rangeCalls from 2023-01 onwardPreferred
SpeakersAt least 300 distinct agents; no agent over 1% of hoursMust
Channel layoutDual-channel, agent on ch0, customer on ch1, consistent mappingMust
Sample rate / codecNative capture rate and codec reported per file; no undisclosed upsamplingMust
Transcript standardVerbatim, written style guide supplied, word-level timestampsMust
DiarizationRTTM per call; overlap taggedMust
RealismHolds/transfers in at least 20% of calls; staged calls flaggedMust
Outcome metadataDisposition, QA score, call driver, resolution flagPreferred
RedactionMethod for audio and transcript described; audited miss rate reportedMust
Consent evidenceNotice script, IVR disclosure, jurisdictions, AI-training permissionMust
Biometric handlingStatement on voiceprint extraction and BIPA/Texas exposureMust
Evaluation sample5-10 h stratified sample under NDA before commitmentMust
PackagingManifest with call ID, segment times, checksumsMust

Packaging fields are covered in packaging speech datasets. Metadata fields such as device and environment are in speech dataset metadata fields.

The evaluation sample: what to request and how to score it

Ask for a stratified evaluation sample under NDA before committing. Some vendors offer a sample of a call-center dataset for evaluation on request [1]. Specify the strata yourself: call type, vertical, language and channel layout. Otherwise you may receive the cleanest calls.

Score the sample against thresholds you set in advance. Measure WER of the supplied transcripts against your own re-transcription of a subset, and DER of the supplied RTTM [9]. Check the redaction miss rate on a manual listen, the share of upsampled files (spectrum cut off at 4 kHz), and the per-agent hour concentration. Write those thresholds into the request so the full delivery can be accepted on the same terms. For diarization test design, see evaluating diarization on your own audio.

Sending a contact-center audio request to SourceX

SourceX sources operational datasets from US companies on request, including support and sales histories, and nothing is held in stock, so a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery, and the supplying company approves each release. Use the template above to describe the call audio you need.

Sources

  1. Unidata, "Call Center Audio Dataset". https://unidata.pro/datasets/call-center-audio
  2. AIxBlock, "Call Center Audio Dataset: What Makes It Production-Ready". https://aixblock.io/blogs/call-center-audio-dataset-what-makes-it-production-ready
  3. arXiv, "Academic call-center speech corpus (arXiv 2507.02958)" (2025). https://web3.arxiv.org/pdf/2507.02958
  4. Datarade, "AI Training Data: Audio Data, Unique Consumer Sentiment Data (WiserBrand)". https://datarade.ai/data-products/ai-training-data-audio-data-unique-consumer-sentiment-data-wiserbrand-com
  5. PCI Security Standards Council, "Information Supplement: Protecting Telephone-Based Payment Card Data, v3.0" (2018). https://listings.pcisecuritystandards.org/documents/Protecting_Telephone_Based_Payment_Card_Data_v3-0_nov_2018.pdf
  6. California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  7. Illinois General Assembly, "740 ILCS 14/10 (Biometric Information Privacy Act, definitions)". https://www.ilga.gov/documents/legislation/ilcs/documents/074000140k10.htm
  8. Texas Legislature, "Texas Business and Commerce Code Section 503.001". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  9. GitHub (nryant), "dscore: diarization scoring tools". https://www.github.com/nryant/dscore
  10. Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  11. U.S. HHS Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data