Speech and audio data
How Speech Data Is Priced: Per Hour, Per Speaker and Add-On Labeling
Quick answer
Speech data is usually quoted per usable audio hour, with per-speaker, per-utterance or per-project pricing used when speaker diversity or collection effort dominates cost. Labeling (transcription, timestamps, diarization, intent or sentiment tags) is priced as an add-on. Most commercial catalogs quote on request, so budget by cost drivers, not list prices: speech realism, language rarity, domain, consent work, exclusivity, transcription standard and delivery format. Normalize every quote to cost per usable hour licensed for your use before comparing.
By SourceX Editorial · Updated
The pricing units vendors actually use
Per audio hour is the default unit, but the hour being priced is rarely the hour you can train on. Catalog pages for call center and conversational speech describe inventory in hours and speakers and bundle or itemize labeling separately [1]. Before comparing, pin down which of these the quote measures.
| Unit | When it fits | What to pin down |
|---|---|---|
| Per raw audio hour | Large conversational or call center corpora | Does the hour include silence, hold music, IVR prompts and non-target channels? |
| Per speech hour (voice activity) | Spontaneous speech where silence is common | Which VAD method and threshold defines "speech"? |
| Per speaker | TTS voices, speaker verification, accent coverage | Minimum minutes per speaker; are repeat speakers counted once? |
| Per utterance or prompt | Wake words, command sets, read prompts | Accepted vs. recorded utterances; retake policy |
| Per project (fixed fee) | Custom collection with a defined spec | Acceptance criteria and what happens on shortfall |
| Per labeled hour (add-on) | Transcription, diarization, tagging | Which standard, which accuracy target, what share of the corpus |
The gap between raw and speech hours matters most for stereo call recordings. A 10-minute two-channel call might contain 10 minutes of audio per channel but far less overlapping customer speech, so a per-hour price can mean very different training value. See audio file specs for sample rate, channels and codecs for how channel layout changes what you are counting.
Why most speech datasets have no list price
Commercial speech catalogs mostly publish scope, not price, and route buyers to a quote. One large call center catalog, describing more than 13,000 hours of customer service calls with time-stamped transcripts, offers a "get in touch" path and a sample request rather than a rate card [1]. Others advertise on-demand recordings tailored by vertical, which implies custom pricing by definition [3].
A visible price can also mislead. A marketplace listing may be free to access while the buyer pays infrastructure costs, and neither fact tells you whether commercial model training is permitted [2]. Open corpora such as Common Voice are released into the public domain, but they are crowdsourced from volunteers and mostly read speech, which is a different product from real support calls [5]. The license audit questions are covered in whether open speech corpora can be used commercially.
Cost drivers that move a speech data quote
Seven variables explain most of the spread between two quotes for "the same" hours; treat them as hypotheses to test against each supplier's answer rather than fixed multipliers.
- Realism. Real operational conversations (support calls, sales calls, field recordings) are scarcer than scripted or read speech and cannot be produced on demand at volume. See spontaneous conversational speech data.
- Language and accent rarity. Low-resource languages, regional dialects and code-switched speech cost more per hour because recruiting speakers or finding holders is harder. See multilingual ASR training data sourcing.
- Domain. Medical dictation, legal, financial and technical support audio carry vocabulary density and compliance overhead that generic conversation does not.
- Consent and rights work. Recordings captured for quality assurance were not necessarily consented for model training. In Texas, a voiceprint is a biometric identifier, and capturing one for a commercial purpose requires notice and consent before capture [4]. Re-papering consent or de-identifying audio is labor that shows up in price.
- Exclusivity. An exclusive license removes the supplier's ability to resell, so expect it to be priced as lost future revenue, often as a separate line.
- Transcription standard. Verbatim transcripts with disfluencies, false starts and non-speech tags cost more than clean read-style text. See verbatim vs clean transcription standards.
- Delivery format and preparation. Resampling, channel splitting, segmentation, redaction tones and metadata assembly are preparation steps; ask whether they are included or billed.
How labeling add-ons are priced
Labeling is usually priced per labeled audio hour, layered on top of the audio fee, and often covers only part of a corpus. One vendor listing offers transcription and diarization as an additional service on a portion of its dataset rather than across all of it [3]. That pattern means a "10,000-hour dataset with transcripts" may contain far fewer transcribed hours.
Common add-on layers include machine transcripts with light review, human clean transcripts, human verbatim transcripts, word-level timestamps, speaker diarization with overlap marking, and semantic tags such as intent, sentiment or call outcome; each is a separate line item whose cost depends on the human effort it requires. Diarization in RTTM with overlap regions takes more annotator time than turn-level speaker labels; see speaker diarization training data and RTTM labels. Ask each supplier to state the quality target (for example, an audited word error rate on a held-out sample) so cheaper labels are not mistaken for equivalent ones.
Normalizing quotes to cost per usable hour
Compare quotes only after converting each one to cost per usable hour for your intended use. Usable means the audio passes your spec, the labels meet your standard, and the license permits your training and deployment use. The method in comparing vendor quotes by cost per usable record applies directly to audio.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Line | Quote A | Quote B |
|---|---|---|
| Quoted audio | 2,000 raw hours, stereo call recordings | 1,200 speech hours, mono, VAD-trimmed |
| Speech share after VAD | 55% of raw (estimate from sample) | Already trimmed |
| Usable speech hours | ~1,100 | 1,200 |
| Transcripts included | 40% of hours, clean style | 100%, verbatim with timestamps |
| Hours lost in your QA (bad audio, wrong language) | 8% (from sample) | 3% (from sample) |
| Effective usable, labeled hours | ~405 | ~1,164 |
| Training use in license | Internal research only | Commercial training and deployment |
| Comparable? | No, until license and transcripts match | Yes |
Even if Quote A has the lower per-hour headline, once transcripts and license scope are matched it may be the more expensive option. Run the same worksheet on every supplier sample before negotiating.
Speech data quote request template
A precise request gets comparable quotes; send the same fields to every supplier.
Illustrative example: invented to show structure; it does not describe an available dataset.
request:
use: "ASR fine-tuning and evaluation for a support voice agent"
speech_type: "real inbound support calls, spontaneous"
languages: ["en-US", "es-US"]
volume: { target_speech_hours: 1000, unit_basis: "speech hours after VAD" }
speakers: { min_unique: 800, max_minutes_per_speaker: 30 }
audio: { sample_rate_hz: 8000, channels: "stereo, agent/customer split", codec: "PCM WAV" }
labels:
transcript: "verbatim, disfluencies tagged"
timestamps: "word-level"
diarization: "RTTM with overlap"
coverage: "100% of delivered hours"
privacy: "names, emails, phone and account numbers removed or replaced; method documented"
rights: "consent basis for model training stated per source"
license_scope: ["commercial training", "deployment", "term", "exclusivity: price as option"]
pricing_requested: ["per speech hour", "per labeled hour by layer", "exclusivity premium", "preparation fees"]
sample: "30-60 minutes representative, with labels"
Ask for pricing broken out by line so you can drop or add layers without a full requote. The broader enterprise pricing picture, beyond audio, is in what drives the price of licensed enterprise data and AI data license pricing structures.
Budget traps specific to speech purchases
Most overruns come from scope that was assumed rather than written down. Watch for these failure modes:
- Channel double counting. Stereo calls billed as two hours per call hour, while only one channel matches your target speaker.
- Partial labeling. Transcripts on a subset, priced as if the whole corpus is labeled [3].
- Redaction artifacts. Beeps or silence inserted over personal details change acoustics and transcripts; see how audio redaction affects speech model training.
- Missing metadata. Speaker, accent, device and environment fields added later as a paid extra; specify them up front using speech dataset metadata fields.
- License mismatch. Free or cheap access that excludes commercial training [2].
- Voice cloning scope. TTS buyers need explicit speaker consent for synthesis; see license terms for speech and voice recordings.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Where SourceX fits in pricing real-world speech data
SourceX sources operational datasets, including support and sales histories and new recordings of hands-on work, from US companies on request, and every release is approved by the supplying company. It does not publish prices; terms are agreed per deal, and nothing is contracted until a supplier agrees. You can describe the speech data you need on the SourceX buyers page, and the speech and audio datasets buyer's guide covers the rest of this cluster.
Get speech data scoped and priced
Describe the speech data you need, not the businesses that might hold it, and SourceX looks for US companies that hold it, assesses data and licensing permissions, then agrees pricing and allowed uses in a license. Personal details are removed or replaced before delivery, and a request does not guarantee a match. Start a speech data request with SourceX.
Sources
- Unidata, "Call Center Audio Dataset". https://unidata.pro/datasets/call-center-audio
- Amazon Web Services, "AWS Marketplace speech data listing". https://aws.amazon.com/marketplace/pp/prodview-yyfwirpya2mp6
- Datarade, "AI Training Data: Audio Data, Unique Consumer Sentiment Data (WiserBrand)". https://datarade.ai/data-products/ai-training-data-audio-data-unique-consumer-sentiment-data-wiserbrand-com
- Texas Legislature, "Texas Business and Commerce Code Section 503.001, Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- Ardila et al. (Mozilla), arXiv, "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.