Skip to content

Speech and audio data

Voice Talent Consent and Release Terms for AI Voice Data

Quick answer

Voice actor consent for AI training is only usable if it is written, specific and traceable to each recording. A buyer should see a signed release per speaker that names the purpose (training, cloning, synthetic output), duration, territory, permitted modifications, compensation, revocation handling and any union terms, plus a record linking every audio file to that release. A vague "AI uses" clause can fail under California Labor Code 927 and invites publicity claims elsewhere.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Voice consent carries a person's identity, not just a copyrighted recording, so a clean data license from the recording owner is not enough on its own. The studio or company that owns the WAV files can license copyright in the sound recording, but the speaker's right of publicity in their voice, and in some states biometric rights in a voiceprint, sit with the individual. General contract terms are covered in the speech and voice recording license terms guide; this page focuses on the performer-side release that has to sit underneath it.

Three bodies of law usually meet in one voice dataset. State publicity statutes now name voice explicitly: Tennessee's ELVIS Act, effective July 1, 2024, added voice to its personal-rights statute. California Labor Code 927 governs contracts that grant rights to create digital replicas of performers. Biometric statutes such as Illinois BIPA and Texas Business and Commerce Code 503.001 list voiceprints as biometric identifiers with notice and consent duties [6][7]; see voiceprints and BIPA risk for the biometric side.

Provenance is also a known weak spot in the field. A 2025 position paper on TTS evaluation notes that many systems train on internet-gathered speech where consent and licensing are hard to verify, and that some technical reports describe data only as "in-house" [4]. Buyers who cannot show consent records may also struggle with downstream disclosure duties such as California AB 2013 training-data documentation [8].

The seven terms a voice release must define

A usable voice release defines purpose, duration, territory, modification rights, compensation, exclusivity and revocation in writing, rather than relying on a one-time checkbox [1]. Each term maps to a specific failure mode when it is missing.

  • Purpose. Separate consent for (a) training a general multi-speaker TTS or ASR model, (b) building a voice clone or speaker-adapted model of this person, and (c) generating and distributing synthetic speech in their voice. Consent to (a) does not imply (b) or (c).
  • Duration. A term for training use and a separate term for output use, plus what happens to trained weights at expiry. Silence here invites later disputes.
  • Territory. Worldwide versus named markets. Publicity rights differ by state and country, so territory interacts with which statutes apply.
  • Permitted modifications. Pitch, speed, emotion and style transfer, accent conversion, language translation (speaking languages the talent does not speak), and combination with other voices in a "voice design" blend.
  • Compensation. Session fee versus per-use, per-model or revenue-based payments, and whether new successor models trigger new payment.
  • Exclusivity. Broad exclusive grants of a person's publicity right are generally not permitted, so voice rights are better structured as scoped licenses [2]. If you need a voice nobody else can clone, negotiate a narrow, clearly defined exclusivity term.
  • Revocation and takedown. What the talent can withdraw, how fast synthetic voices are retired, and whether already-trained general models are affected.

How California Labor Code 927 changes release drafting

California Labor Code 927 makes some digital replica provisions unenforceable, so vague digital replica releases are a concrete risk rather than a theoretical one. As summarized from the codified text, a provision fails where it allows a digital replica to replace work the performer would otherwise have done and lacks a reasonably specific description of intended uses, unless the performer was represented by counsel with clearly stated commercial terms or covered by a collective bargaining agreement that expressly addresses digital replicas. The law took effect January 1, 2025 and applies to new performances by a digital replica fixed on or after that date.

For buyers, the practical rule is to reject releases that say only "any and all media now known or hereafter devised, including AI." Ask whether the performer had counsel or union coverage, and get the specific use list in writing. Courts have not, as of October 2026 in the material reviewed here, settled what "reasonably specific" means, so drafting conservatively is cheaper than testing it.

Union coverage is the second question to ask before contracting. Where a performer works under a collective bargaining agreement that addresses digital replicas, that agreement may set minimum terms the release must respect, and the California exception depends on it. Confirm the talent's union status in the intake form, not after recording.

Where voice cloning disputes actually arise

Disputes usually arise when talent learns their voice is being sold or deployed for a use they did not understand at signing. A 2025 New York federal decision on AI voice cloning, summarized by Skadden, shows the exposure shifting toward state publicity and consumer-protection theories rather than copyright in the voice itself [5]. That makes the release, not the copyright chain, the asset a buyer is really purchasing.

Research on voice actors explains why these misunderstandings happen. A 2025 study of voice actors found they struggle to tell legitimate jobs from voice-data harvesting, prefer verifiable clients, and weigh privacy, reputation, credit and compensation risks when deciding what to accept [3]. Releases that were obtained through anonymous gig postings, or that bury AI use inside a general session contract, carry higher dispute risk even if technically signed.

Common failure modes buyers see in diligence:

  • Releases signed for an audiobook or IVR job, later reused to train a cloning model.
  • One master release covering a whole cohort, with no per-speaker signature or date.
  • Consent obtained in English from speakers recorded in other languages.
  • Minors recorded without verified parent or guardian consent.
  • No record of which audio files belong to which release version after re-takes or re-edits.

Consent is only as good as the record that ties it to each file, so a voice data purchase should ship with a per-speaker consent manifest. This complements technical specs such as those in audio file specs for speech datasets and consent tracking practices described under consent management.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "speaker_id": "SPK-0412",
  "release_id": "REL-2026-0412-v2",
  "release_version_hash": "sha256:9f2c...",
  "signed_at": "2026-03-14",
  "signature_method": "e-signature with audit trail",
  "represented_by_counsel": false,
  "union_cba_applies": "yes, agreement named in release",
  "age_at_recording": "18+ verified",
  "language_of_release": "en-US",
  "purposes": {
    "general_tts_training": true,
    "speaker_clone": false,
    "synthetic_output_distribution": false
  },
  "permitted_modifications": ["speed", "pitch"],
  "prohibited_modifications": ["cross_lingual_synthesis", "voice_blending"],
  "territory": "worldwide",
  "training_term_end": "2031-03-14",
  "output_term_end": null,
  "biometric_notice_given": ["IL BIPA written release", "TX 503.001 notice"],
  "compensation_basis": "session fee plus per-model fee",
  "revocation_contact": "talent-rights@example.com",
  "audio_files": ["SPK-0412/sess01/utt_0001.wav", "SPK-0412/sess01/utt_0002.wav"]
}

Ask for the release template text itself, every version used over the collection period, and a mapping from release version to recording date. Spot-check a sample of manifests against signed PDFs before acceptance.

Release term checklist for TTS, cloning and voice design

Use the checklist below to compare what each use case needs before signing a purchase. The stronger the identity link between the output and a single speaker, the narrower and more explicit the release must be.

Illustrative example: invented to show structure; it does not describe an available dataset.

TermMulti-speaker TTS or ASR trainingSingle-speaker voice cloneVoice design (blended or new voices)
Purpose namedTraining onlyClone creation and named output usesTraining plus blending into new voices
Output useNot tied to one speakerMust list products, channels, languagesState that no output should be identifiable as the talent
DurationTraining term, weights at expirySeparate output term and renewalTraining term; blending usually irreversible
ModificationsAugmentation (speed, noise)Emotion, style, cross-lingual listed or excludedMixing with other speakers expressly allowed
CompensationSession fee commonPer-use or ongoing fees commonSession fee plus disclosure of blending
RevocationFuture datasets onlyRetire clone; define takedown stepsUsually future datasets only; disclose this
Biometric noticeRequired where voiceprints capturedRequiredRequired
Union or counsel checkAskStrongly advised where Labor Code 927 appliesAsk

Successor models deserve a line of their own. State whether consent covers retraining, distillation and fine-tuning of later model versions, and whether a revoked speaker must be excluded from the next training run. For broader license structure, see AI training data licensing and AI data license terms explained.

Questions to ask a voice data supplier

The fastest diligence is a short written questionnaire answered before any audio moves. Ask suppliers:

  1. Who recorded the speakers, under what contract, and who signed each release?
  2. Which purposes did each speaker consent to, and is cloning or synthetic output use excluded?
  3. Were any speakers union members or represented by counsel when they signed, and will the release support digital replica performances fixed on or after January 1, 2025?
  4. How were biometric notices handled for Illinois, Texas or other residents [6][7]?
  5. What revocation requests have been received, and how are revoked speakers removed from deliverables?
  6. Can you supply the consent manifest and release versions per file?

If you are commissioning new studio recordings rather than buying existing ones, the TTS training data guide covers script, booth and speaker-selection specs, and the speech and audio cluster hub maps related topics such as call-recording consent.

SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; categories are not inventory and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, and every release is approved by the supplying company. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset. You can describe the voice data and consent scope you need on the SourceX buyers page.

Describe the speakers, purposes and consent terms your voice project requires, and SourceX looks for US businesses that hold matching data and manages the Find, Assess, Agree, Transact and Manage process. Nothing is contracted until a supplier agrees. Start a buyer request at sourcex.si/buyers.

Sources

  1. Speech Actors, "Ethical Considerations in AI Voice Generation and Usage Rights". https://speechactors.com/article/ethical-consideration-ai-voice-generation
  2. Morgan Lewis, "Rise of Text-to-Speech AI Models, Part 1: Intellectual Property Issues" (2024). https://www.morganlewis.com/blogs/sourcingatmorganlewis/2024/07/rise-of-text-to-speech-ai-models-part-1-intellectual-property-issues
  3. arXiv, "PRAC3 (Privacy, Reputation, Accountability, Consent, Credit, Compensation): Long Tailed Risks of Voice Actors in AI Data-Economy" (2025). https://arxiv.org/pdf/2507.16247
  4. arXiv, "Position: Towards Responsible Evaluation for Text-to-Speech" (2025). https://arxiv.org/pdf/2510.06927
  5. Skadden, "New York Court Tackles the Legality of AI Voice Cloning" (2025). https://www.skadden.com/-/media/files/publications/2025/07/new_york_court_tackles_the_legality_of_ai_voice_cloning.pdf
  6. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  7. Texas Legislature, "Texas Business and Commerce Code Section 503.001: Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  8. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data