Speech and audio data
TTS Training Data: Consented Studio Voice Recordings for Commercial Speech Synthesis
Quick answer
A TTS dataset suitable for commercial use is one where every voice was recorded with written consent that names synthesis and voice modeling as allowed uses, and where the audio meets synthesis-grade specs, typically 24 kHz or higher, sentence-aligned scripts, consistent mic chain and low room noise. Public read-speech corpora rarely clear both bars at once. Contracted voice talent, documented performer releases and a license that defines records, uses and term are the practical route.
By SourceX Editorial · Updated
Why commercial TTS needs a different dataset than ASR
TTS training data must be cleaner, higher-fidelity and more consistent per speaker than ASR data, because the model reproduces what it hears rather than only transcribing it. An ASR model benefits from noise, accents and channel variety; a synthesis model learns that noise as part of the voice. Audio sampled at 16 kHz is common for recognition, but synthesis corpora are usually built at 24 kHz or above, which is why TTS-oriented releases built from the same source material as ASR corpora are usually re-derived at a higher sample rate and re-segmented rather than reused as-is.
A synthesis-ready derivative usually changes four things: higher sample rate, segmentation at sentence breaks instead of silences, original plus normalized text, and removal of noisy utterances. Those four changes are a useful baseline checklist for any vendor sample. For format detail, see our guide to audio file specs for speech datasets.
Crowdsourced ASR corpora such as Common Voice were designed for recognition across many speakers and devices [7]. They are valuable for ASR and for testing robustness, but uneven microphones and short prompts make them a weak base for a single branded voice.
The three data shapes buyers actually source for TTS
Most commercial TTS programs need one or more of three shapes: single-speaker studio sets, multi-speaker studio sets and purpose-recorded conversational sets. Each has different spec and consent questions.
Single-speaker studio sets feed a flagship or brand voice. The buyer cares about hours from one performer, script coverage (phoneme and diphone balance, numbers, dates, abbreviations, domain terms), session-to-session consistency and a locked mic chain. Drift between sessions, such as a different preamp or a cold, shows up as audible inconsistency in the trained voice.
Multi-speaker studio sets support voice design, zero-shot adaptation and speaker-embedding training. Here the count of distinct performers, balance across age, gender and accent, and per-speaker minimum duration matter more than total hours. Each speaker needs a separate, traceable release.
Conversational sets support expressive and interactive synthesis: backchannels, laughter, hesitations and turn-taking. Research on full-duplex synthesis has turned to recorded dialogue built for the purpose rather than read speech [6]. Found conversational audio is a poor substitute, because disfluencies and overlapping speech make it ill-suited to TTS without heavy cleaning [3]. Our page on spontaneous conversational speech data covers that collection in depth.
Consent is the review item, not a footnote
Documented performer consent is now the first thing legal and responsible-AI reviewers ask about for a TTS corpus. A 2025 position paper notes that many TTS systems train on large internet speech collections where consent and licensing are hard to verify, and that some technical reports describe their data only as "in-house" [1]. That gap is exactly what diligence will probe.
The documented-consent route is contracted voice talent. One TTS vendor states the position directly: record original training data with speaker permission and contracts rather than scraping [2]. Treat that as market practice, not law, but it matches what reviewers expect to see.
Consent for a voice must be specific. Research on voice performers in the AI data economy describes long-tail risks across privacy, reputation, accountability, consent, credit and compensation [4]. A release that covers "recordings for research" or "improving our services" does not clearly cover training a synthetic voice that can say anything. See voice talent consent and release terms and the license terms for speech and voice recordings for clause-level detail, and our consent management glossary entry for how consent records are tracked.
Voice likeness and biometric law that shapes a TTS deal
State law in the US treats a voice as both a likeness right and, in some contexts, a biometric identifier, so a commercial TTS license has to address both. As of October 2026, the main reference points buyers ask counsel about are these.
- Tennessee ELVIS Act (2024). Tennessee extended its statutory publicity right to cover an individual's voice, including simulated voices [8]. A cloned or closely imitated voice is the risk case; have counsel review the current statute text.
- Texas Business and Commerce Code §503.001. It lists a voiceprint as a biometric identifier and requires notice and consent before capture for a commercial purpose [5].
- Illinois BIPA. Voiceprints are also covered [9]; our page on voiceprints and BIPA explains where speaker embeddings create exposure.
The practical consequence: operational audio such as contact-center calls is not a voice-cloning source by default. Agents and customers typically consented to recording for quality or service purposes, not to having a synthetic voice built from them. If you need call audio for ASR or voice-agent understanding, that is a separate use, covered on our call center audio datasets page.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Checking open TTS corpora before you rely on them
Open corpora can seed research, but each needs a license and consent audit before it touches a commercial model. Audiobook-derived corpora are a common example: permissive copyright terms on the text and audio do not mean the volunteer readers agreed to have a synthetic voice built from them, so voice-likeness questions remain even where copyright is clear.
Two failure modes recur. First, non-commercial terms (CC BY-NC) on a popular corpus quietly contaminate a production checkpoint. Second, a derived corpus inherits obligations from its source that the derived release page does not repeat. Our open speech corpora license audit walks through both.
Spec sheet to send with a TTS data request
A precise request returns comparable samples and avoids a second round of questions. Use a sheet like the one below and attach it to any request.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value for a single-speaker brand voice | Why it matters |
|---|---|---|
| Use | Commercial TTS training and fine-tuning; voice design excluded | Scopes consent and license |
| Speakers | 1 primary performer, plus 4 backup performers | Continuity if a performer withdraws |
| Hours | Target per speaker, stated as clean speech after trimming | Raw session hours overstate usable data |
| Sample rate / depth | 48 kHz, 24-bit PCM WAV, mono | Fidelity; resampling down is easy, up is not |
| Mic chain | Same mic, preamp and room for all sessions; chain logged per session | Session drift becomes audible drift |
| Room | Treated booth; room tone captured per session | Noise floor and denoising reference |
| Script | Phonetically balanced core, domain terms, numbers, dates, questions | Coverage of prosody and normalization cases |
| Text | Original and normalized transcripts, sentence-aligned | Front-end training and alignment |
| Styles | Neutral read, plus labeled expressive styles | Expressive control without guessing labels |
| Consent | Signed release naming synthetic voice training, cloning scope, term, withdrawal | Diligence and likeness law |
| Packaging | Per-utterance manifest (file, speaker_id, text, normalized_text, duration, style, session_id) | Reproducible splits; see packaging guide |
For manifest conventions and segment timing, see speech dataset manifest packaging.
Questions to ask a supplier before any audio moves
Ask for documents first, audio second. A short diligence pass catches most problems.
- Show the release template each performer signed, with the clause naming synthetic voice training.
- Does consent cover synthetic voices that resemble the performer, and does it allow the recordings to train multi-speaker or voice-design models?
- How is withdrawal handled, and what happens to trained models if a performer withdraws?
- Were any recordings captured for another purpose (calls, meetings, podcasts) and repurposed?
- What was the mic chain and room per session, and are there session logs?
- What filtering removed noise, clipping and misreads, and what share was rejected?
- Are transcripts verified against audio, and who checked normalization?
- What speaker metadata is held, and is any of it a biometric identifier under Texas or Illinois law [5]?
Where SourceX fits
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases; the kinds of data include new recordings of hands-on work as well as support and sales histories. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and SourceX does not source scraped web content. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, through private, access-controlled workflows after an executed agreement and supplier approval.
For TTS buyers, that means the consent scope above gets assessed before anything is agreed, and nothing is contracted until a supplier agrees. You can describe your voice data need to SourceX alongside the spec sheet, or review our voice and audio data overview and the wider speech and audio data hub.
Request consented voice recordings for commercial TTS
Describe the voices, hours, specs and allowed uses your model needs; SourceX looks for US businesses that hold matching data, assesses data and licensing permissions, and every release is approved by the supplying company. SourceX does not publish prices, and terms are agreed per deal. Start a TTS data request.
Sources
- arXiv, "Position: Towards Responsible Evaluation for Text-to-Speech" (2025). https://arxiv.org/pdf/2510.06927
- ReadSpeaker, "Ethical AI at ReadSpeaker: Best Practices for the AI Voice Industry". https://www.readspeaker.com/blog/ethical-ai
- ISCA Archive, "Adigwe et al., Interspeech 2022 paper on conversational speech for TTS" (2022). https://www.isca-archive.org/interspeech_2022/adigwe22_interspeech.pdf
- arXiv, "PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy" (2025). https://arxiv.org/pdf/2507.16247
- Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- arXiv, "Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis" (2025). https://arxiv.org/pdf/2509.04093
- arXiv (Ardila et al., Mozilla), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
- Tennessee General Assembly, "Ensuring Likeness, Voice, and Image Security (ELVIS) Act of 2024 (Public Chapter 588)". https://publications.tnsosfiles.com/acts/113/pub/pc0588.pdf
- Illinois General Assembly, "Illinois Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004&ChapterID=57
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.