Speech and audio data
Speech and Audio Datasets for AI: A Buyer's Guide to Sourcing and Licensing
Quick answer
Speech datasets for AI fall into distinct types: contact-center calls, spontaneous conversation, accented and multilingual speech, far-field recordings, dictation, consented TTS voice, and non-speech audio such as machine sounds and music. Each type is reached through open corpora, vendor catalogs, commissioned collection or licensed business recordings. Choose the type from the acoustic conditions your model will meet in deployment, then check every route for commercial license terms, recording consent and voiceprint (biometric) exposure.
By SourceX Editorial · Updated
This hub covers the speech and audio cluster of the SourceX guide to AI data. Licensing mechanics shared by every data type are in the AI training data licensing guide.
Speech and audio data types and the systems that produce them
The system that captured a recording fixes most of what a buyer inherits: sample rate, channel layout, speaker mix, noise and the legal overlay. Specify a dataset by its origin and traits, not only by a label such as "conversational English."
| Data type | Typical origin | Traits a buyer must specify |
|---|---|---|
| Contact-center calls | Call-recording and quality-management tools in contact-center platforms | Narrowband telephony; stereo (agent and customer on separate channels) or mixed mono; hold music, transfers, spoken card and account numbers |
| Spontaneous and full-duplex conversation | Meetings, interviews, commissioned unscripted sessions | Overlap, backchannels and laughter; one track per speaker for turn-taking models |
| Accented, dialectal and multilingual speech | Crowdsourced reading, field collection, multinational call operations | Self-reported first language, region and code-switching for each speaker |
| Far-field and noisy speech | Smart speakers, vehicles, meeting-room arrays, drive-thru lanes | Distance, reverberation, microphone-array geometry, signal-to-noise ratio (SNR) range |
| Medical and legal dictation | Dictation systems, digital recorders, telephone dictation lines | One speaker, dense terminology, protected health information in clinical audio |
| TTS and expressive voice | Studio sessions with contracted voice talent | High-fidelity single-speaker capture, scripted and expressive prompts, a signed release |
| Machine and environmental sound | Microphones on pumps, fans, valves and production lines | Few fault examples; MIMII's authors built their set because, at the time, no public dataset covered real-factory machines in normal and anomalous states [1] |
| Music and broadcast audio | Label and publisher catalogs, production libraries, podcast archives | Separate rights in the composition and the sound recording, which one licensing vendor cites as a reason licensed music is scarce [2]; identifiable on-air voices |
Automatic speech recognition (speech-to-text) draws on most of the speech rows. Speech-language models also need paired audio-text instruction and dialogue data, covered in training data for speech-language models.
Four routes to speech data and where each one breaks
Speech data reaches buyers through open corpora, vendor catalogs, commissioned collection or licensed operational recordings, and each route trades license clarity, acoustic realism and consent evidence differently. The routes can be combined: an open corpus can supply breadth while licensed recordings cover the deployment domain.
| Route | What it gives you | License pattern | Where it breaks |
|---|---|---|---|
| Open corpora | Free, documented sets such as Common Voice, a crowdsourced public-domain corpus [3], or The People's Speech, about 30,000 hours of mostly English speech [4] | Public domain, CC-BY and share-alike CC-BY-SA [4] or non-commercial: the CallCenterEN corpus is CC BY-NC 4.0 [5] | Hosting-site license fields are unreliable: an audit of text datasets found license omissions above 70% and error rates above 50% [6]; mostly read or found speech |
| Vendor catalogs | Packaged hours with transcripts | Often quoted on request, with evaluation samples instead of list prices [7] | Self-reported hours; check how much of the audio the transcripts and diarization actually cover |
| Commissioned collection | Speakers, prompts, devices and rooms to your specification | Custom terms plus per-speaker releases | Elicited speech is not real use; CASPER's authors note that most conversational datasets are scripted [8] |
| Licensed operational recordings | Real calls, dictation or meetings a business recorded in its work | Negotiated license defining records, uses, term and delivery | Recording notices may not mention AI training; spoken personal data and voiceprints need treatment |
Synthetic speech supplements these routes rather than replacing them: a 2025 paper on emotional speech-language models notes that even models that take speech as input can fail to perceive the paralinguistic information, such as pitch and speed, that is crucial for understanding emotion [9]. The trade-offs are worked through in off-the-shelf, custom or licensed speech data, real vs synthetic speech data and the open speech corpora license audit.
SourceX works the fourth route for AI teams wherever they are based. It sources operational datasets from US companies, including support and sales histories and new recordings of hands-on work, and manages the licensing agreement and later purchases. Datasets are sourced on request, not held in stock, so a request does not guarantee a match, and SourceX does not source scraped public web content.
Every dataset goes through rights review, which checks that the business may share the records and that required consents are in place. Buyers can describe the recordings they need to SourceX; the voice and audio data licensing page covers this data type specifically.
Acoustic fit matters before hour counts
Match bandwidth, distance, speakers, overlap and transcription style to deployment first; extra hours add value only once those match. Each mismatch below shows up in a sample through a listening test and a manifest check.
- Bandwidth. Telephone audio sampled at 8 kHz carries content only up to 4 kHz, so a model trained on 16 kHz wideband speech meets calls without the 4-8 kHz cues it learned [10]. Specify native sample rate and capture codec, and require upsampled files to be flagged (telephony vs wideband audio).
- Distance and room. Far-field training data is often simulated: one study generated about 2,000 hours of reverberant, noisy speech at 0-25 dB SNR with microphone geometry matched to the evaluation device [11]. Simulation helps only when it matches your array, so state the geometry (far-field speech data).
- Speakers. On the Edinburgh International Accents of English Corpus (EdAcc), the best model its authors tested averaged 19.7% word error rate (WER), against 2.7% on US English clean read speech [12]. Per-speaker accent and first-language labels let you build an accent-stratified evaluation set.
- Overlap and turn-taking. Full-duplex speech-to-speech models learn interruptions and backchannels from separate per-speaker tracks; one open dual-track release totals 15 hours [13]. Stereo call recordings are the operational equivalent (dual-channel call recordings).
- Transcription style. Verbatim and clean transcripts score differently, and a multi-reference study argues that standard WER probably overstates the contentful errors of top ASR systems because reference styles vary [14]. Fix one style guide and one text normalizer before comparing suppliers (verbatim vs clean transcription).
- Real hours. Recorded time is not speech time: CASPER's authors report 200 hours recorded, of which 158 hours are speech [8]. Ask for speech hours after silence and hold-music removal, unique speakers and minutes per speaker (how many hours to fine-tune ASR).
Voice rights: recording consent, voiceprints and likeness
Voice data carries legal layers that text mostly does not: consent to record the conversation, biometric rules for voiceprints and publicity rights in a voice, on top of license and privacy terms. The points below reflect the cited sources as of October 2026; general de-identification law is covered in the de-identified data guide.
- Recording consent. California Penal Code section 632 prohibits recording a confidential communication without the consent of all parties [15]. Ask which states the calls came from, what notice callers heard and whether it mentioned AI model training (call-recording consent checks; California call-recording rules).
- Voiceprints. Illinois's Biometric Information Privacy Act (BIPA) lists voiceprints as biometric identifiers, requires a written release before collection and provides liquidated damages of $1,000 per negligent and $5,000 per intentional or reckless violation [16]. A 2024 amendment (SB 2979) counts repeated collection of the same biometric from the same person by the same method as a single violation [16]. Texas Business and Commerce Code section 503.001 also covers voiceprints and requires notice and consent before capture for a commercial purpose [17]. The EU GDPR treats biometric data processed for the purpose of uniquely identifying a person as a special category of personal data under Article 9 [18].
- Pending voiceprint suits. Class-action lawsuits filed in May 2026 allege that voiceprints were extracted from recorded speech to train AI models [19]. As of October 2026 they are pending, with no merits ruling in the sources reviewed (voiceprints and BIPA risk).
- Voice likeness. According to a law-firm analysis, Tennessee's ELVIS Act, effective July 1, 2024, protects a person's voice, including a simulation of it, as a property right; the same analysis notes that broad exclusive licenses of a person's publicity rights are generally not permitted [20]. A TTS position paper notes that many systems are trained on internet-collected voices whose consent and licensing are hard to verify [21], so buy only with releases that name training and synthesis (voice talent consent and release terms).
- Health audio. HIPAA's Safe Harbor method lists biometric identifiers, including voice prints, among the identifiers to remove [22], so ask how the audio itself was treated, not only the transcript. For health records, SourceX requires HIPAA de-identification (Safe Harbor or Expert Determination) before anything is considered for a license.
- Model documentation. Providers placing general-purpose AI models on the EU market must publish a summary of training content using the AI Office template [23]. Keep source and license records for every audio set you buy.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What a deliverable speech record should contain
A usable speech dataset ships audio plus a manifest that ties every segment to its audio specs, speaker metadata, transcript, labels, redactions and license terms. Review the manifest schema before you negotiate hours.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"segment_id": "rec_000481_s007",
"recording_id": "rec_000481",
"audio_file": "audio/rec_000481.flac",
"start_s": 41.20,
"end_s": 47.85,
"sample_rate_hz": 8000,
"bit_depth": 16,
"capture_codec": "G.711 mu-law",
"channels": 2,
"channel_map": {"0": "agent", "1": "customer"},
"upsampled": false,
"speaker_id": "spk_c_2291",
"speaker_role": "customer",
"speaker_meta": {"first_language": "es", "region": "US-TX", "source": "self_reported"},
"language": "en-US",
"environment": "mobile_handset_in_vehicle",
"transcript": "yeah I- I was charged twice on the [ACCOUNT_NUMBER] card",
"transcript_style": "verbatim_guide_v2",
"word_alignment": "align/rec_000481.ctm",
"diarization_ref": "rttm/rec_000481.rttm",
"redactions": [
{"pii_type": "account_number", "start_s": 45.10, "end_s": 46.30,
"audio_method": "tone_1khz", "text_tag": "[ACCOUNT_NUMBER]"}
],
"consent_basis": "recording_notice_v3",
"license_id": "LIC-EXAMPLE-01",
"permitted_uses": ["asr_training", "internal_evaluation"]
}
sample_rate_hz,capture_codecandupsampledseparate native telephony audio from resampled audio.diarization_refpoints to an RTTM file, the format that scoring tools such as dscore expect for both reference and system output [24] (diarization training data).redactionsrecords how spoken personal data was masked in audio and text, because a tone, a silence or a tag changes what a model learns at that span (audio redaction artifacts).consent_basisandpermitted_usestie each segment to the notice it was recorded under and the license it ships under.
Packaging conventions are in packaging speech datasets, and field requirements in speech dataset metadata fields.
Start here: match your model goal to data and first checks
Your model goal decides which data type to source first, which check to run before anything else and which guide to read next.
| If you are building | Source first | First check | Read next |
|---|---|---|---|
| Contact-center ASR or call analytics | Real stereo calls at native telephony rate | Origin states, recording notice, spoken-data redaction method | contact-center audio requirements template; call center audio datasets |
| Voice agent or speech-to-speech model | Dual-track spontaneous dialogue | Separate tracks and overlap labels | full-duplex conversation data; voice agent training data |
| Commercial TTS voice | Studio recordings from contracted talent | Release naming AI training, synthesis and term | consented TTS recordings |
| Medical ASR or ambient scribe | De-identified dictation or clinician-patient audio | De-identification method for audio and text | physician dictation audio; doctor-patient conversation audio |
| Multilingual or accent-robust ASR | Speaker-labeled speech per language and accent | Hours and speakers per language; self-reported labels | multilingual ASR data; accented English speech |
| Acoustic anomaly detection | Machine audio tied to maintenance records | Fault labels traceable to work orders | machine audio labeled with maintenance records |
| Voice model evaluation | Held-out real audio with task labels | No speaker, call or source overlap with training data | evaluation data for voice models; LLM evaluation datasets |
Mistakes that make purchased speech data unusable
Most failed speech data purchases trace to an acoustic mismatch or a rights gap that a sample, a manifest review and a close reading of the license would have caught.
- Clean read speech for a phone product. Studio or read audio lacks telephony bandwidth, codec artifacts and crosstalk.
- Trusting a hosting site's license tag. Audits find hosting-site license fields often missing or wrong [6], so read the license text for each subset; share-alike or non-commercial terms bind downstream use.
- Redacting the transcript but not the audio. A masked account number in text is still spoken in the waveform, and the voice itself remains. The VoicePrivacy Challenge scores voice anonymization against speaker-verification attackers and measures utility by ASR WER [25]; treat anonymized audio as lower risk, not zero risk.
- Comparing vendor WER claims as published. Re-score every candidate on your own test set with one normalizer (measuring WER fairly).
- Splitting train and test by file. Split by speaker and by call or session so the same voices never appear on both sides.
- Paying for volume before a pilot. Measure the WER change a sample produces on your model first (piloting speech data); per-hour price structures are explained in how speech data is priced.
For datasets sourced through SourceX, personal details such as names, phone numbers and account numbers are removed or replaced before delivery; the method used is recorded for each dataset, and a sample is checked after processing. No de-identification method is perfect.
Describe the speech data your model needs
At the SourceX buyer page you can submit a data request, talk to SourceX or browse the dataset catalog. Describe the audio type, channel layout, languages, labels and permitted uses you need; SourceX looks for US companies that hold matching recordings, checks the data and the supplier's licensing permissions, and manages the license and delivery through private, access-controlled workflows. Nothing is contracted until a supplier agrees.
Guides in this section
- 8 kHz Telephony vs Wideband Audio for ASR TrainingHow the 8 kHz vs 16 kHz bandwidth gap hurts ASR on phone calls, when to upsample, augment or split models, and what metadata to require.
- Accented English Speech Data for ASR: Sourcing GuideHow to source accented and L2 English speech for ASR: accent and L1 metadata, conversational coverage, WER gap evidence, license checks and a request spec.
- Audio Specs for Speech Datasets: Rate, Depth, CodecHow to specify sample rate, bit depth, channels, container and codec for speech dataset deliveries, and what resampling and lossy transcoding lose.
- BIPA Voiceprint Risk in Voice Data for AI TrainingWhen voice recordings for ASR, TTS or speech-LLM training may count as BIPA voiceprints, and the consent, notice and retention evidence to demand.
- Call Center Audio Dataset Requirements: A Buyer's TemplateSpecify call center audio data before contacting suppliers: hours, call types, channel layout, transcripts, labels, consent evidence and license scope.
- Call-Recording Consent for AI Training: Buyer ChecksWhat consent and notice evidence to require before licensing call recordings for ASR or voice AI training: all-party states, IVR scripts, biometrics.
- Doctor-Patient Conversation Audio for Ambient Scribe AIHow to source doctor-patient conversation audio paired with signed notes for ambient scribe training: consent, de-identification, specs and metrics.
- Full-Duplex Conversation Data for Speech-to-Speech ModelsWhat full-duplex speech-to-speech models need from conversation data: per-speaker tracks, overlap, backchannel and barge-in labels, timing and consent.
- How Many Hours of Audio to Fine-Tune ASR: Sizing GuideSize an ASR fine-tuning dataset in hours and speakers: domain adaptation vs from-scratch scale, learning-curve pilots and a sizing table for speech buyers.
- How to Calculate Word Error Rate with Fair NormalizationCalculate word error rate correctly: shared text normalization, reference style, CER and formatted error rate, plus a fair protocol to compare ASR vendors.
- Machine Sound Data for Acoustic Anomaly DetectionHow to specify and source real machine sound recordings, normal and faulty, for acoustic anomaly detection: mic setup, SNR, labels, domain shift, rights.
- Multilingual ASR Data: Hours, Speakers and DialectsPlan multilingual ASR training data per language: hours, unique speakers, dialect spread, channel, transcription QA and commercial license checks.
- Off-the-Shelf vs Custom vs Licensed Speech DataCompare catalog speech datasets, commissioned crowd or studio collection and licensed real recordings on realism, rights, ownership, cost and time to data.
- Physician Dictation Audio for Medical ASR: Sourcing GuideHow to source physician dictation audio for medical ASR: specialties, devices, verbatim vs report transcripts, HIPAA de-identification of audio and text.
- Real vs Synthetic Speech Data for ASR: Limits of TTS AudioWhere TTS-generated and simulated speech helps ASR and wake-word training (negatives, rare words, rooms) and where real recordings are still required.
- Speaker Diarization Datasets: RTTM Labels and Overlap SpecsHow to source speaker diarization training data: RTTM speaker-turn labels, overlap and backchannel rules, channel setup, and a spec to give suppliers.
- Speech Corpus Licenses: Commercial Use Audit for ASR and TTSAudit open speech corpora before commercial ASR or TTS training: CC BY, CC BY-SA, NC and ND terms, speaker rights, provenance and what to replace.
- Speech-Language Model Training Data: Audio-Text PairsWhat speech-language model teams need beyond TTS prompts: real spoken dialogue, audio-text instruction pairs, labels, specs and voice consent checks.
- Spontaneous Speech Datasets: Sourcing Real ConversationHow to source spontaneous conversational speech for ASR, voice agents and speech LLMs: what read and scripted data miss, what to specify and what to check.
- TTS Training Data for Commercial Use: Consented VoicesHow to source consented studio and conversational voice recordings for commercial TTS: audio specs, consent scope, likeness law and a request template.
- Verbatim vs Clean Transcription Standards for ASR DataHow to choose and document a verbatim or clean verbatim transcription standard for ASR data: fillers, false starts, numbers, tags and casing rules.
- Voice Actor Consent for AI Training: Release TermsConsent and release terms voice data must carry for TTS and cloning: purpose, duration, territory, modification, compensation, revocation and records.
- Accent-Stratified ASR Fairness Eval Sets: Design GuideHow to design an accent-stratified ASR fairness evaluation set: strata, speakers per group, reference transcripts, WER gap metrics and reporting.
- Code-Switched Speech Data: Spanish-English Sourcing GuideHow to source Spanish-English code-switched conversational audio for ASR and voice agents: language labels, transcription rules, metrics, license checks.
- Diarization Error Rate: Collars, Overlap and Test SetsHow to compute diarization error rate correctly: collars, overlap scoring, RTTM and UEM files, JER vs DER, and building a test set from your own audio.
- Domain Vocabulary Coverage in ASR Training DataHow to measure and close domain vocabulary gaps in ASR: term lists, entity error rate, number normalization and the right mix of in-domain audio and text.
- Drive-Thru Speech Data for Voice Ordering AI: Buyer GuideHow to source drive-thru and phone ordering audio for voice AI: lane acoustics, POS-linked order labels, menu versions, recording consent and BIPA checks.
- Dual-Channel Call Recordings for ASR and DiarizationWhy agent/customer stereo call audio beats mono mixes for ASR fine-tuning, diarization and turn-taking, and how to specify channel layout from suppliers.
- Far-Field Speech Data: Real Recordings vs Simulated RoomsHow to split far-field ASR training data between real distant-microphone recordings and simulated reverberation, and what metadata to require for each.
- Forced Alignment and Word Timestamps in Speech DatasetsWhat timestamp granularity to require in licensed speech data, how forced alignment is produced, and how to check alignment quality before you train.
- Machine Fault Audio Labels from Maintenance RecordsHow to turn work orders and CMMS failure codes into fault labels for machine audio: time windows, taxonomies, clock alignment and what to ask suppliers.
- Noisy Speech Data for ASR: Noise Types and SNR CoverageHow to specify noisy speech data for robust ASR: a noise-type taxonomy, SNR bands for training and test, real vs added noise, and per-file labels.
- Overlapping Speech Data for Multi-Talker ASR and tcpWERHow to source and evaluate overlapping speech data for multi-talker ASR: per-speaker references, overlap labels, mic metadata, consent review and tcpWER.
- Podcast and Broadcast Audio for AI Training: Rights GuideHow to license podcast and broadcast audio for AI training: publisher, host, guest and music rights, BIPA voiceprint risk, and a rights audit checklist.
- Redacted Audio for ASR Training: Silence, Tones and TagsHow muting, beeps, noise and splicing in redacted call or dictation audio change ASR training, alignment and VAD labels, and what to ask suppliers for.
- Speech Data Pilots: Measuring WER Lift on a SampleRun a speech data pilot before you buy: build an in-domain held-out test set, fine-tune on the sample, and measure normalized WER lift per slice.
- Speech Data Pricing: Per Hour, Per Speaker and Add-OnsHow speech data is priced per audio hour, per speaker and per utterance, which labeling add-ons and cost drivers move quotes, and how to normalize them.
- Speech Dataset Metadata: Speaker, Device and Room FieldsPer-speaker and per-file metadata a licensed speech dataset must carry: speaker ID, L1 and accent, age band, device, channel map and environment.
- Speech Emotion Recognition Data: Acted vs NaturalisticHow to source speech emotion recognition data: acted vs naturalistic corpora, label mapping, annotator agreement, commercial rights and legal limits.
- Speech LLM Evaluation Data: Measuring the Modality GapDesign held-out spoken eval sets for speech LLMs, voice agents and TTS: paired text and speech tasks, real human audio, stage metrics and listener tests.
- Training ASR on Non-Verbatim Transcripts: Align and FilterHow to judge, align and filter audio paired with edited, clean or caption transcripts before using them as ASR training targets, and how to evaluate it.
- Transcript Accuracy Audits for Licensed Speech DataHow to acceptance-test supplier transcripts: re-transcribe a stratified sample, compute normalized WER, and check coverage, timestamps and speaker labels.
- US Dialect Speech Data for ASR: Coverage, Metadata, ConsentHow to source US regional and social dialect speech data for ASR and voice agents: coverage targets, dialect metadata, transcription and consent checks.
Sources
- arXiv (Purohit et al.), "MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection" (2019). https://arxiv.org/pdf/1909.09347
- Troveo (vendor page), "Licensed Audio Data for AI Training: What Labs Buy and Why". https://www.troveo.ai/resources/licensed-audio-data-for-ai
- Ardila et al. (Mozilla), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
- arXiv, "The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage" (2021). https://arxiv.org/pdf/2111.09344
- arXiv, "Real-World En Call Center Transcripts Dataset with PII Redaction (arXiv:2507.02958)". https://arxiv.org/abs/2507.02958
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Unidata (vendor page), "Call Center Audio Dataset". https://unidata.pro/datasets/call-center-audio
- arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
- arXiv, "Dual Information Speech Language Models for Emotional Conversations" (2025). https://arxiv.org/pdf/2508.08095
- U.S. Patent and Trademark Office, "Feature domain bandwidth extension and spectral rebalance for ASR data augmentation" (US Patent 12148437). https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/12148437
- arXiv, "Spatial Attention for Far-field Speech Recognition with Deep Beamforming Neural Networks" (2019). https://arxiv.org/pdf/1911.02115
- ar5iv / arXiv, "The Edinburgh International Accents of English Corpus (EdAcc)" (2023). https://ar5iv.labs.arxiv.org/html/2303.18110
- arXiv, "Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis" (2025). https://arxiv.org/pdf/2509.04093
- arXiv, "Style-agnostic evaluation of ASR using multiple reference transcripts" (2024). https://arxiv.org/pdf/2412.07937
- California Legislature, "Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
- Texas Legislature, "Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)". https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- Biometric Update, "Tech giants sued under BIPA over voiceprints used to train AI" (2026). https://www.biometricupdate.com/202605/tech-giants-sued-under-bipa-over-voiceprints-used-to-train-ai
- Morgan Lewis (law-firm blog), "Rise of text-to-speech AI models, part 1: intellectual property issues" (2024). https://www.morganlewis.com/blogs/sourcingatmorganlewis/2024/07/rise-of-text-to-speech-ai-models-part-1-intellectual-property-issues
- arXiv, "Position: Towards Responsible Evaluation for Text-to-Speech" (2025). https://arxiv.org/pdf/2510.06927
- eCFR, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- srvk / GitHub, "dscore (GitHub repository)". https://github.com/srvk/dscore
- arXiv, "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.