Speech and audio data
How to Specify Contact-Center Audio Data: A Requirements Template for Buyers
Quick answer
A contact-center audio request should state measurable requirements before any supplier conversation. These are total and labeled hours, call types and verticals, languages and accents, date range, distinct agent and caller counts, channel layout, sample rate and codec, transcript standard, diarization and timestamp format, and outcome metadata. It also needs the redaction method, consent and notice evidence by jurisdiction, and the license scope. A spec written this way makes quotes comparable and screens out research-only corpora early.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why a call-audio spec differs from a generic data request
A call-audio spec has to fix acoustic, conversational and legal properties that a generic data request leaves open. The general structure in how to write a data request for suppliers still applies. But two suppliers can both say "1,000 hours of support calls" and deliver very different things. One may ship mono 8 kHz G.711 mixdowns with clean-read transcripts from three agents. The other may ship dual-channel recordings from 400 agents with verbatim, diarized transcripts.
Those differences decide whether the data helps ASR fine-tuning, a voice agent, speech analytics or summarization. They also decide whether it is legal to use at all. The wider sourcing picture is in the speech and audio data buyer's guide. This page covers the requirements document itself.
Volume, call mix and population: the fields that set coverage
Coverage is set by how hours are spread across call types, verticals, speakers and time, not by the headline hour count. Ask for hours broken down by call type: support, sales, collections, scheduling, billing disputes, retention, and inbound versus outbound. Ask for the same breakdown by vertical, such as telecom, insurance, healthcare scheduling, utilities and retail. Collections and healthcare calls bring their own regulatory overlays, so name them explicitly or exclude them.
Put a cap on how much any one speaker can contribute. A corpus where 20 agents produce most of the audio teaches a model those 20 voices. Request distinct agent and caller counts, plus a maximum share of hours per agent. State a date range too, because IVR prompts, product names and script language drift over time. For guidance on volume targets, see how many hours of audio you need to fine-tune ASR.
Channel layout, sample rate and codec: the acoustic contract
The acoustic fields should describe the audio as captured, not as converted. Telephony audio usually arrives at 8 kHz narrowband over G.711 μ-law or A-law, or at 16 kHz over wideband codecs such as G.722 or Opus on VoIP legs. Upsampled 8 kHz audio saved as 16 kHz WAV still contains no content above 4 kHz. So ask for the original capture rate and codec, plus every transcoding step. The trade-offs are covered in telephony versus wideband audio for ASR.
Channel layout matters most for diarization. Dual-channel recordings, with the agent on one channel and the customer on the other, give near-free speaker attribution and show overlap cleanly. Mono mixdowns force you to rely on model-based diarization. Specify which you need, and whether the left and right channel mapping is consistent across the corpus. See dual-channel call recordings for ASR and diarization and audio file specs for speech datasets.
Transcript standard, timestamps and diarization: the label contract
The transcript standard must be stated as a rule set, not just a label like "human transcribed." Vendors describe call-center datasets by transcript type, time-stamped segments and speaker labels [1]. Your spec should define the following:
- Verbatim or clean. Verbatim keeps disfluencies, false starts and fillers. Clean drops them. Name which you want and how hesitations, numbers, spelled-out letters and code-switching are written.
- Tags. List the required tags for overlap, crosstalk, unintelligible speech, hold music, IVR prompts and redacted spans.
- Timing. Say whether you need segment-level or word-level timestamps, and the timing tolerance you will accept.
- Speaker turns. Request turns in RTTM so you can score them with standard tools [9].
- Labeled share. State what share of total hours must be transcribed, and whether transcripts come from one pass or from double transcription with adjudication.
Clean or professional transcripts can still serve as training targets, with caveats explained in training ASR on non-verbatim transcripts. Plan acceptance testing before you sign, using transcript accuracy acceptance testing.
Realism criteria: real calls, not role-play
A production-ready call corpus should be specified by what live calls contain. That means holds, transfers, IVR menus, crosstalk, background noise, speakerphone and Bluetooth audio, and dropped or abandoned calls. One vendor argues that models trained mostly on clean audio fail on live calls [2]. That claim is self-reported, but the failure mode is real: scripted role-play underrepresents interruptions and agent screen-reading pauses.
Write realism into the spec as measurable fields. Examples include the share of calls containing a hold or transfer, a minimum per-call overlap ratio, and an SNR distribution. Require a flag on every call that is staged, simulated or synthetic. If you want transfers and warm handoffs, ask whether transferred legs are stitched into one file or delivered separately with a shared call ID.
Outcome metadata: dispositions, QA scores and CRM links
Outcome metadata turns audio into training signal for analytics, summarization and agent-assist. Useful fields include disposition codes, wrap-up notes written by agents, QA scorecard results, call-driver tags, handle time, transfer and escalation flags, and links to CRM or ticket outcomes such as resolved, refunded, churned or deal stage. Ask how each field was produced. A disposition chosen by an agent at wrap-up is noisier than a ticket status set by a later system.
Agent desktop actions recorded alongside the call are a separate category, covered in contact-center desktop activity paired with call transcripts. If sales outcomes drive the use case, the sales call transcript datasets page covers that variant.
Redaction, consent and biometric evidence: the compliance fields
The compliance section should ask for evidence, not reassurance. Call recordings carry spoken names, addresses, account numbers and often card numbers. PCI SSC guidance on telephone-based payment card data addresses how card data in call recordings should be handled [5]. Ask suppliers to describe the redaction method for both audio and transcripts: silence, tone or noise masking; entity tags in transcripts; how spans are aligned between the two; and the measured miss rate on an audited sample. Some marketplace listings say PII is removed before delivery [4]. Your spec should require a written method, not a checkbox. Redaction artifacts also change model behavior, as covered in how audio redaction affects speech model training and redacting spoken PII from call recordings.
Consent evidence must match where the parties were. California requires all-party consent to record confidential communications [6]. So ask for the recording notice script, IVR disclosure audio and the date ranges each applied. You also need evidence that the supplier's agreements and notices permit reuse for AI training, not just quality assurance. The call recording consent checker and call-recording consent for AI training go deeper.
Voice is a biometric issue if anyone extracts speaker embeddings. Illinois BIPA lists a voiceprint as a biometric identifier [7], and Texas law does the same, with notice and consent required before capture for a commercial purpose [8]. If calls come through a BPO, require proof that each end client authorized the reuse. See sourcing AI training data through BPOs. If any calls involve patient health information, require HIPAA de-identification by Safe Harbor or Expert Determination [11].
License scope: state it before the first call
License scope belongs in the first paragraph of the request, because it disqualifies more candidates than any acoustic field. Name the uses you need: commercial model training, fine-tuning, evaluation only, or redistribution of derived models. Also name the term, the user population and whether the data can be used by affiliates or contractors. Public call-center corpora are often research-only. One academic call-center corpus, for example, is released under CC BY-NC 4.0, which rules out commercial training [3]. An audit of dataset hosting sites found that licenses are frequently missing or wrong [10], so ask for the license that governs the specific files, not a platform tag. If you want licensed audio from operating US businesses rather than public corpora, you can submit a buyer request to SourceX; the call center audio datasets page explains that category.
The requirements template
Use the template below as the body of your request. Mark each field as "must," "preferred" or "informational" so suppliers can respond to each line.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example requirement | Priority |
|---|---|---|
| Intended use | Commercial fine-tuning of a streaming ASR model and held-out evaluation | Must |
| Total hours / labeled hours | 2,000 h audio; at least 600 h with verbatim transcripts | Must |
| Call types | Support 50%, billing 20%, scheduling 15%, collections 0%, other 15% | Must |
| Verticals | Telecom, insurance, utilities; exclude healthcare clinical calls | Preferred |
| Languages / accents | US English; at least 15% Spanish or code-switched calls | Must |
| Date range | Calls from 2023-01 onward | Preferred |
| Speakers | At least 300 distinct agents; no agent over 1% of hours | Must |
| Channel layout | Dual-channel, agent on ch0, customer on ch1, consistent mapping | Must |
| Sample rate / codec | Native capture rate and codec reported per file; no undisclosed upsampling | Must |
| Transcript standard | Verbatim, written style guide supplied, word-level timestamps | Must |
| Diarization | RTTM per call; overlap tagged | Must |
| Realism | Holds/transfers in at least 20% of calls; staged calls flagged | Must |
| Outcome metadata | Disposition, QA score, call driver, resolution flag | Preferred |
| Redaction | Method for audio and transcript described; audited miss rate reported | Must |
| Consent evidence | Notice script, IVR disclosure, jurisdictions, AI-training permission | Must |
| Biometric handling | Statement on voiceprint extraction and BIPA/Texas exposure | Must |
| Evaluation sample | 5-10 h stratified sample under NDA before commitment | Must |
| Packaging | Manifest with call ID, segment times, checksums | Must |
Packaging fields are covered in packaging speech datasets. Metadata fields such as device and environment are in speech dataset metadata fields.
The evaluation sample: what to request and how to score it
Ask for a stratified evaluation sample under NDA before committing. Some vendors offer a sample of a call-center dataset for evaluation on request [1]. Specify the strata yourself: call type, vertical, language and channel layout. Otherwise you may receive the cleanest calls.
Score the sample against thresholds you set in advance. Measure WER of the supplied transcripts against your own re-transcription of a subset, and DER of the supplied RTTM [9]. Check the redaction miss rate on a manual listen, the share of upsampled files (spectrum cut off at 4 kHz), and the per-agent hour concentration. Write those thresholds into the request so the full delivery can be accepted on the same terms. For diarization test design, see evaluating diarization on your own audio.
Sending a contact-center audio request to SourceX
SourceX sources operational datasets from US companies on request, including support and sales histories, and nothing is held in stock, so a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery, and the supplying company approves each release. Use the template above to describe the call audio you need.
Sources
- Unidata, "Call Center Audio Dataset". https://unidata.pro/datasets/call-center-audio
- AIxBlock, "Call Center Audio Dataset: What Makes It Production-Ready". https://aixblock.io/blogs/call-center-audio-dataset-what-makes-it-production-ready
- arXiv, "Academic call-center speech corpus (arXiv 2507.02958)" (2025). https://web3.arxiv.org/pdf/2507.02958
- Datarade, "AI Training Data: Audio Data, Unique Consumer Sentiment Data (WiserBrand)". https://datarade.ai/data-products/ai-training-data-audio-data-unique-consumer-sentiment-data-wiserbrand-com
- PCI Security Standards Council, "Information Supplement: Protecting Telephone-Based Payment Card Data, v3.0" (2018). https://listings.pcisecuritystandards.org/documents/Protecting_Telephone_Based_Payment_Card_Data_v3-0_nov_2018.pdf
- California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- Illinois General Assembly, "740 ILCS 14/10 (Biometric Information Privacy Act, definitions)". https://www.ilga.gov/documents/legislation/ilcs/documents/074000140k10.htm
- Texas Legislature, "Texas Business and Commerce Code Section 503.001". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- GitHub (nryant), "dscore: diarization scoring tools". https://www.github.com/nryant/dscore
- Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- U.S. HHS Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.