Skip to content

Speech and audio data

Speech and Audio Datasets for AI: A Buyer's Guide to Sourcing and Licensing

Quick answer

Speech datasets for AI fall into distinct types: contact-center calls, spontaneous conversation, accented and multilingual speech, far-field recordings, dictation, consented TTS voice, and non-speech audio such as machine sounds and music. Each type is reached through open corpora, vendor catalogs, commissioned collection or licensed business recordings. Choose the type from the acoustic conditions your model will meet in deployment, then check every route for commercial license terms, recording consent and voiceprint (biometric) exposure.

By SourceX Editorial · Updated

This hub covers the speech and audio cluster of the SourceX guide to AI data. Licensing mechanics shared by every data type are in the AI training data licensing guide.

Speech and audio data types and the systems that produce them

The system that captured a recording fixes most of what a buyer inherits: sample rate, channel layout, speaker mix, noise and the legal overlay. Specify a dataset by its origin and traits, not only by a label such as "conversational English."

Data typeTypical originTraits a buyer must specify
Contact-center callsCall-recording and quality-management tools in contact-center platformsNarrowband telephony; stereo (agent and customer on separate channels) or mixed mono; hold music, transfers, spoken card and account numbers
Spontaneous and full-duplex conversationMeetings, interviews, commissioned unscripted sessionsOverlap, backchannels and laughter; one track per speaker for turn-taking models
Accented, dialectal and multilingual speechCrowdsourced reading, field collection, multinational call operationsSelf-reported first language, region and code-switching for each speaker
Far-field and noisy speechSmart speakers, vehicles, meeting-room arrays, drive-thru lanesDistance, reverberation, microphone-array geometry, signal-to-noise ratio (SNR) range
Medical and legal dictationDictation systems, digital recorders, telephone dictation linesOne speaker, dense terminology, protected health information in clinical audio
TTS and expressive voiceStudio sessions with contracted voice talentHigh-fidelity single-speaker capture, scripted and expressive prompts, a signed release
Machine and environmental soundMicrophones on pumps, fans, valves and production linesFew fault examples; MIMII's authors built their set because, at the time, no public dataset covered real-factory machines in normal and anomalous states [1]
Music and broadcast audioLabel and publisher catalogs, production libraries, podcast archivesSeparate rights in the composition and the sound recording, which one licensing vendor cites as a reason licensed music is scarce [2]; identifiable on-air voices

Automatic speech recognition (speech-to-text) draws on most of the speech rows. Speech-language models also need paired audio-text instruction and dialogue data, covered in training data for speech-language models.

Four routes to speech data and where each one breaks

Speech data reaches buyers through open corpora, vendor catalogs, commissioned collection or licensed operational recordings, and each route trades license clarity, acoustic realism and consent evidence differently. The routes can be combined: an open corpus can supply breadth while licensed recordings cover the deployment domain.

RouteWhat it gives youLicense patternWhere it breaks
Open corporaFree, documented sets such as Common Voice, a crowdsourced public-domain corpus [3], or The People's Speech, about 30,000 hours of mostly English speech [4]Public domain, CC-BY and share-alike CC-BY-SA [4] or non-commercial: the CallCenterEN corpus is CC BY-NC 4.0 [5]Hosting-site license fields are unreliable: an audit of text datasets found license omissions above 70% and error rates above 50% [6]; mostly read or found speech
Vendor catalogsPackaged hours with transcriptsOften quoted on request, with evaluation samples instead of list prices [7]Self-reported hours; check how much of the audio the transcripts and diarization actually cover
Commissioned collectionSpeakers, prompts, devices and rooms to your specificationCustom terms plus per-speaker releasesElicited speech is not real use; CASPER's authors note that most conversational datasets are scripted [8]
Licensed operational recordingsReal calls, dictation or meetings a business recorded in its workNegotiated license defining records, uses, term and deliveryRecording notices may not mention AI training; spoken personal data and voiceprints need treatment

Synthetic speech supplements these routes rather than replacing them: a 2025 paper on emotional speech-language models notes that even models that take speech as input can fail to perceive the paralinguistic information, such as pitch and speed, that is crucial for understanding emotion [9]. The trade-offs are worked through in off-the-shelf, custom or licensed speech data, real vs synthetic speech data and the open speech corpora license audit.

SourceX works the fourth route for AI teams wherever they are based. It sources operational datasets from US companies, including support and sales histories and new recordings of hands-on work, and manages the licensing agreement and later purchases. Datasets are sourced on request, not held in stock, so a request does not guarantee a match, and SourceX does not source scraped public web content.

Every dataset goes through rights review, which checks that the business may share the records and that required consents are in place. Buyers can describe the recordings they need to SourceX; the voice and audio data licensing page covers this data type specifically.

Acoustic fit matters before hour counts

Match bandwidth, distance, speakers, overlap and transcription style to deployment first; extra hours add value only once those match. Each mismatch below shows up in a sample through a listening test and a manifest check.

  • Bandwidth. Telephone audio sampled at 8 kHz carries content only up to 4 kHz, so a model trained on 16 kHz wideband speech meets calls without the 4-8 kHz cues it learned [10]. Specify native sample rate and capture codec, and require upsampled files to be flagged (telephony vs wideband audio).
  • Distance and room. Far-field training data is often simulated: one study generated about 2,000 hours of reverberant, noisy speech at 0-25 dB SNR with microphone geometry matched to the evaluation device [11]. Simulation helps only when it matches your array, so state the geometry (far-field speech data).
  • Speakers. On the Edinburgh International Accents of English Corpus (EdAcc), the best model its authors tested averaged 19.7% word error rate (WER), against 2.7% on US English clean read speech [12]. Per-speaker accent and first-language labels let you build an accent-stratified evaluation set.
  • Overlap and turn-taking. Full-duplex speech-to-speech models learn interruptions and backchannels from separate per-speaker tracks; one open dual-track release totals 15 hours [13]. Stereo call recordings are the operational equivalent (dual-channel call recordings).
  • Transcription style. Verbatim and clean transcripts score differently, and a multi-reference study argues that standard WER probably overstates the contentful errors of top ASR systems because reference styles vary [14]. Fix one style guide and one text normalizer before comparing suppliers (verbatim vs clean transcription).
  • Real hours. Recorded time is not speech time: CASPER's authors report 200 hours recorded, of which 158 hours are speech [8]. Ask for speech hours after silence and hold-music removal, unique speakers and minutes per speaker (how many hours to fine-tune ASR).

Voice data carries legal layers that text mostly does not: consent to record the conversation, biometric rules for voiceprints and publicity rights in a voice, on top of license and privacy terms. The points below reflect the cited sources as of October 2026; general de-identification law is covered in the de-identified data guide.

  • Recording consent. California Penal Code section 632 prohibits recording a confidential communication without the consent of all parties [15]. Ask which states the calls came from, what notice callers heard and whether it mentioned AI model training (call-recording consent checks; California call-recording rules).
  • Voiceprints. Illinois's Biometric Information Privacy Act (BIPA) lists voiceprints as biometric identifiers, requires a written release before collection and provides liquidated damages of $1,000 per negligent and $5,000 per intentional or reckless violation [16]. A 2024 amendment (SB 2979) counts repeated collection of the same biometric from the same person by the same method as a single violation [16]. Texas Business and Commerce Code section 503.001 also covers voiceprints and requires notice and consent before capture for a commercial purpose [17]. The EU GDPR treats biometric data processed for the purpose of uniquely identifying a person as a special category of personal data under Article 9 [18].
  • Pending voiceprint suits. Class-action lawsuits filed in May 2026 allege that voiceprints were extracted from recorded speech to train AI models [19]. As of October 2026 they are pending, with no merits ruling in the sources reviewed (voiceprints and BIPA risk).
  • Voice likeness. According to a law-firm analysis, Tennessee's ELVIS Act, effective July 1, 2024, protects a person's voice, including a simulation of it, as a property right; the same analysis notes that broad exclusive licenses of a person's publicity rights are generally not permitted [20]. A TTS position paper notes that many systems are trained on internet-collected voices whose consent and licensing are hard to verify [21], so buy only with releases that name training and synthesis (voice talent consent and release terms).
  • Health audio. HIPAA's Safe Harbor method lists biometric identifiers, including voice prints, among the identifiers to remove [22], so ask how the audio itself was treated, not only the transcript. For health records, SourceX requires HIPAA de-identification (Safe Harbor or Expert Determination) before anything is considered for a license.
  • Model documentation. Providers placing general-purpose AI models on the EU market must publish a summary of training content using the AI Office template [23]. Keep source and license records for every audio set you buy.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What a deliverable speech record should contain

A usable speech dataset ships audio plus a manifest that ties every segment to its audio specs, speaker metadata, transcript, labels, redactions and license terms. Review the manifest schema before you negotiate hours.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "segment_id": "rec_000481_s007",
  "recording_id": "rec_000481",
  "audio_file": "audio/rec_000481.flac",
  "start_s": 41.20,
  "end_s": 47.85,
  "sample_rate_hz": 8000,
  "bit_depth": 16,
  "capture_codec": "G.711 mu-law",
  "channels": 2,
  "channel_map": {"0": "agent", "1": "customer"},
  "upsampled": false,
  "speaker_id": "spk_c_2291",
  "speaker_role": "customer",
  "speaker_meta": {"first_language": "es", "region": "US-TX", "source": "self_reported"},
  "language": "en-US",
  "environment": "mobile_handset_in_vehicle",
  "transcript": "yeah I- I was charged twice on the [ACCOUNT_NUMBER] card",
  "transcript_style": "verbatim_guide_v2",
  "word_alignment": "align/rec_000481.ctm",
  "diarization_ref": "rttm/rec_000481.rttm",
  "redactions": [
    {"pii_type": "account_number", "start_s": 45.10, "end_s": 46.30,
     "audio_method": "tone_1khz", "text_tag": "[ACCOUNT_NUMBER]"}
  ],
  "consent_basis": "recording_notice_v3",
  "license_id": "LIC-EXAMPLE-01",
  "permitted_uses": ["asr_training", "internal_evaluation"]
}
  • sample_rate_hz, capture_codec and upsampled separate native telephony audio from resampled audio.
  • diarization_ref points to an RTTM file, the format that scoring tools such as dscore expect for both reference and system output [24] (diarization training data).
  • redactions records how spoken personal data was masked in audio and text, because a tone, a silence or a tag changes what a model learns at that span (audio redaction artifacts).
  • consent_basis and permitted_uses tie each segment to the notice it was recorded under and the license it ships under.

Packaging conventions are in packaging speech datasets, and field requirements in speech dataset metadata fields.

Start here: match your model goal to data and first checks

Your model goal decides which data type to source first, which check to run before anything else and which guide to read next.

If you are buildingSource firstFirst checkRead next
Contact-center ASR or call analyticsReal stereo calls at native telephony rateOrigin states, recording notice, spoken-data redaction methodcontact-center audio requirements template; call center audio datasets
Voice agent or speech-to-speech modelDual-track spontaneous dialogueSeparate tracks and overlap labelsfull-duplex conversation data; voice agent training data
Commercial TTS voiceStudio recordings from contracted talentRelease naming AI training, synthesis and termconsented TTS recordings
Medical ASR or ambient scribeDe-identified dictation or clinician-patient audioDe-identification method for audio and textphysician dictation audio; doctor-patient conversation audio
Multilingual or accent-robust ASRSpeaker-labeled speech per language and accentHours and speakers per language; self-reported labelsmultilingual ASR data; accented English speech
Acoustic anomaly detectionMachine audio tied to maintenance recordsFault labels traceable to work ordersmachine audio labeled with maintenance records
Voice model evaluationHeld-out real audio with task labelsNo speaker, call or source overlap with training dataevaluation data for voice models; LLM evaluation datasets

Mistakes that make purchased speech data unusable

Most failed speech data purchases trace to an acoustic mismatch or a rights gap that a sample, a manifest review and a close reading of the license would have caught.

  • Clean read speech for a phone product. Studio or read audio lacks telephony bandwidth, codec artifacts and crosstalk.
  • Trusting a hosting site's license tag. Audits find hosting-site license fields often missing or wrong [6], so read the license text for each subset; share-alike or non-commercial terms bind downstream use.
  • Redacting the transcript but not the audio. A masked account number in text is still spoken in the waveform, and the voice itself remains. The VoicePrivacy Challenge scores voice anonymization against speaker-verification attackers and measures utility by ASR WER [25]; treat anonymized audio as lower risk, not zero risk.
  • Comparing vendor WER claims as published. Re-score every candidate on your own test set with one normalizer (measuring WER fairly).
  • Splitting train and test by file. Split by speaker and by call or session so the same voices never appear on both sides.
  • Paying for volume before a pilot. Measure the WER change a sample produces on your model first (piloting speech data); per-hour price structures are explained in how speech data is priced.

For datasets sourced through SourceX, personal details such as names, phone numbers and account numbers are removed or replaced before delivery; the method used is recorded for each dataset, and a sample is checked after processing. No de-identification method is perfect.

Describe the speech data your model needs

At the SourceX buyer page you can submit a data request, talk to SourceX or browse the dataset catalog. Describe the audio type, channel layout, languages, labels and permitted uses you need; SourceX looks for US companies that hold matching recordings, checks the data and the supplier's licensing permissions, and manages the license and delivery through private, access-controlled workflows. Nothing is contracted until a supplier agrees.

Share your speech data requirements with SourceX

Guides in this section

Sources

  1. arXiv (Purohit et al.), "MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection" (2019). https://arxiv.org/pdf/1909.09347
  2. Troveo (vendor page), "Licensed Audio Data for AI Training: What Labs Buy and Why". https://www.troveo.ai/resources/licensed-audio-data-for-ai
  3. Ardila et al. (Mozilla), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
  4. arXiv, "The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage" (2021). https://arxiv.org/pdf/2111.09344
  5. arXiv, "Real-World En Call Center Transcripts Dataset with PII Redaction (arXiv:2507.02958)". https://arxiv.org/abs/2507.02958
  6. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  7. Unidata (vendor page), "Call Center Audio Dataset". https://unidata.pro/datasets/call-center-audio
  8. arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
  9. arXiv, "Dual Information Speech Language Models for Emotional Conversations" (2025). https://arxiv.org/pdf/2508.08095
  10. U.S. Patent and Trademark Office, "Feature domain bandwidth extension and spectral rebalance for ASR data augmentation" (US Patent 12148437). https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/12148437
  11. arXiv, "Spatial Attention for Far-field Speech Recognition with Deep Beamforming Neural Networks" (2019). https://arxiv.org/pdf/1911.02115
  12. ar5iv / arXiv, "The Edinburgh International Accents of English Corpus (EdAcc)" (2023). https://ar5iv.labs.arxiv.org/html/2303.18110
  13. arXiv, "Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis" (2025). https://arxiv.org/pdf/2509.04093
  14. arXiv, "Style-agnostic evaluation of ASR using multiple reference transcripts" (2024). https://arxiv.org/pdf/2412.07937
  15. California Legislature, "Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  16. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  17. Texas Legislature, "Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  18. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)". https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  19. Biometric Update, "Tech giants sued under BIPA over voiceprints used to train AI" (2026). https://www.biometricupdate.com/202605/tech-giants-sued-under-bipa-over-voiceprints-used-to-train-ai
  20. Morgan Lewis (law-firm blog), "Rise of text-to-speech AI models, part 1: intellectual property issues" (2024). https://www.morganlewis.com/blogs/sourcingatmorganlewis/2024/07/rise-of-text-to-speech-ai-models-part-1-intellectual-property-issues
  21. arXiv, "Position: Towards Responsible Evaluation for Text-to-Speech" (2025). https://arxiv.org/pdf/2510.06927
  22. eCFR, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  23. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  24. srvk / GitHub, "dscore (GitHub repository)". https://github.com/srvk/dscore
  25. arXiv, "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data