Skip to content

Speech and audio data

Podcast and Broadcast Audio for AI Training: Rights, Voices and Licensing Paths

Quick answer

Podcast and broadcast audio is valuable for ASR, speech-LLM pre-training, TTS and emotion research because it captures unscripted, multi-speaker conversation. It is also one of the hardest audio types to license, because a single episode can carry four separate rights layers: the publisher or network, the hosts, the guests and the embedded music or clips. Each voice is also potential biometric data. Buy only where a rights holder can grant every layer in writing, and treat "publicly available" as a risk flag, not as permission.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why podcast and broadcast audio is worth the rights work

Podcast and broadcast speech is worth licensing because read-speech corpora do not cover turn-taking, overlap, laughter, disfluency or emotional range. Researchers already use podcast audio as the base for naturalistic emotional speech corpora, which shows its value for affect and prosody modeling [4]. Long-form interviews also give speech-language models multi-turn dialogue with real topical structure, which is hard to get from scripted sources (see training data for speech-language models).

The value comes with technical caveats. Broadcast audio is usually mastered: compressed dynamics, loudness normalization, music under speech and stingers between segments. Distribution files are often lossy MP3 or AAC at 64–128 kbps, so ask whether the publisher holds multitrack WAV session files or the pre-mix stems. Those are better for speaker diarization training and they make music removal far cleaner.

The four rights layers in one episode

Every podcast or broadcast segment should be treated as a bundle of separate rights, and a license that covers only one of them is incomplete. The publisher or network typically owns the copyright in the recording and edit. That ownership does not automatically include the right to license the voices of the people on it for model training.

  • Publisher or network. Owns or controls the episode as a work. Check whether a network only distributes the show while an independent production company or the host owns the masters.
  • Hosts. Talent agreements may grant the network rights in recordings for distribution, but older contracts rarely mention AI training, voice synthesis or derivative voice models. Union or guild terms can also apply to broadcast talent.
  • Guests. Guest release forms usually cover recording, editing and distribution "in all media." Whether that language reaches model training, and especially voice cloning, is a contract question for counsel. Many shows have no signed releases at all for older episodes.
  • Embedded third-party content. Music beds, theme songs, stings, archival clips, ad reads and call-ins. Under US law a sound recording and the underlying composition are separate copyrights, and a show's performance license for broadcast does not usually extend to training.

The rights audit checklist below turns this into a per-episode audit, and buyers who want conversational audio sourced from a company that holds it can describe the requirement to SourceX. For collective arrangements that may cover some catalogue layers at once, see collective licenses for AI training.

Voiceprint and biometric risk in conversational audio

Voices in podcasts and broadcasts are the primary biometric exposure, and in 2026 that exposure moved from theory to litigation. Illinois' Biometric Information Privacy Act lists a voiceprint as a biometric identifier and requires a published retention schedule and a written release before collection [5]. Texas Business and Commerce Code Section 503.001 also covers voiceprints and requires notice and consent before capture for a commercial purpose [7].

In May 2026, class actions filed in the Northern District of Illinois alleged that major technology companies used voiceprints to train AI without BIPA consent [1]. Reported plaintiffs include journalists and podcasters, which is exactly the population a broadcast archive contains [2]. As of October 2026 none of these claims has been decided on the merits, and whether training on public audio amounts to "collecting" a voiceprint remains contested [3]. Law-firm commentary treats the suits as a signal that voice data pipelines need consent documentation, not just copyright licenses [14].

Two details matter for procurement. The 2024 BIPA amendment (SB 2979) makes repeated collection of the same identifier from the same person by the same method a single violation, which limits but does not remove per-person exposure [6]. And the question is not only whether you build speaker embeddings on purpose: diarization, speaker verification heads and voice-conversion models can all derive voice templates. For a deeper treatment, see voiceprints and BIPA in licensed voice datasets and the state-by-state view in biometric data in AI training: BIPA, Texas CUBI and Washington.

Training a TTS or voice-conversion model on a host's voice raises publicity and likeness claims that a copyright license does not resolve. Tennessee's ELVIS Act extended the state's publicity statute to protect an individual's voice against unauthorized AI imitation [8]. Courts in New York have also been asked whether AI voice cloning violates publicity and related rights [9].

The practical consequence is that use scope must be written into the license. ASR training that discards speaker identity, emotion classification, and a TTS model that can reproduce a recognizable host are three different risk profiles. If your intended use includes generation in a speaker's voice, expect to need individual written consent from that speaker, not just the publisher's signature.

Copyright law for training on audio is unsettled, so licensing is the route that does not depend on how fair use resolves. On 29 September 2026 the Third Circuit, on interlocutory appeal, affirmed partial summary judgment for Thomson Reuters, holding that ROSS Intelligence's use of Westlaw headnotes to train a non-generative tool was not fair use [11]. In July 2026 a class settlement in Bartz v. Anthropic received final approval [12]. Neither is a ruling on podcasts, and outcomes in pending cases cannot be predicted.

Disclosure obligations also reach your audio sources. California AB 2013 requires developers of generative AI systems offered to Californians to post documentation about training data, including sources and whether the data contains copyrighted or personal information [13]. In the EU, Article 53 of the AI Act requires general-purpose model providers to keep a copyright compliance policy and respect text-and-data-mining reservations under Article 4(3) of the DSM Directive, duties that have applied since 2 August 2025 [10]. A podcast archive with a clean chain of title is much easier to describe in those disclosures than scraped RSS feeds.

Three ways to acquire podcast and broadcast audio

Most buyers end up choosing among direct archive licenses, commissioned conversational recordings and organization-held operational audio, and each fits a different use. The decision table compares them on the layers above.

Illustrative example: invented to show structure; it does not describe an available dataset.

PathRights layers you must clearVoice/biometric postureBest fitTypical failure mode
License a publisher or network back cataloguePublisher, host contracts, guest releases, music and clipsConsent rarely covers AI training for legacy episodesASR and dialogue pre-training on long-form speechGuest releases silent on AI; music beds unlicensed
Commission new podcast-style recordingsSingle contract per speaker with AI-specific consentWritten BIPA-style release captured at recordingTTS, emotion, speech-LLM instruction dataScripted feel; narrower speaker diversity
Organization-held conversational audio (webinars, internal broadcasts, recorded calls)Company ownership, employee and participant noticesDepends on employer notices and consent recordsDomain ASR, diarization, meeting modelsNotices predate AI use; third-party callers not covered

For general trade-offs among these routes, see off-the-shelf, custom or licensed real-world speech data. Media buyers can also read the media and publishing buyer overview.

Rights audit checklist for a podcast or broadcast archive

A per-episode audit is the practical way to decide which recordings are licensable and which must be excluded or remediated. Ask the rights holder to fill one row per episode and keep it as part of your provenance file (see data provenance for AI training).

Illustrative example: invented to show structure; it does not describe an available dataset.

episode_id: EP-2019-0412
publisher_of_record: "Example Audio LLC"
master_owner: "Example Audio LLC"         # confirm vs. distributor
recording_date: 2019-04-12
duration_s: 3480
format_available: [mp3_128k, wav_48k_24bit_multitrack]
speakers:
  - role: host
    contract_ref: "Talent agreement 2018, s.7 media grant"
    ai_training_consent: false            # needs amendment
    voice_synthesis_consent: false
    residence_state: IL                   # BIPA review required
  - role: guest
    release_on_file: true
    release_mentions_ai: false
third_party_content:
  - type: music_bed
    composition_license: sync_for_podcast_only
    master_license: library_music_subscription
    action: remove_or_mute_segments
  - type: archival_news_clip
    action: exclude_segment
ad_reads_present: true                    # strip dynamic ad insertion
permitted_uses: [asr_training, diarization]
excluded_uses: [voice_cloning, speaker_identification]

Use the checklist to drive five decisions:

  1. Chain of title. Confirm who owns the masters, not just who distributes the RSS feed.
  2. Speaker consents. Map every speaker to a contract or release, and flag Illinois, Texas and Washington residents for biometric review.
  3. Third-party content. Remove or mute music beds, clips and ad reads that are not licensed for training, and log the timestamps you cut.
  4. Use scope. Write permitted and excluded uses (for example, no speaker identification or voice cloning) into the license.
  5. Downstream records. Keep the audit table so you can answer AB 2013 or Article 53 disclosure questions later.

Technical prep that changes the rights picture

Some preprocessing choices reduce risk and some create it, so agree on them before delivery. Music-source separation can strip beds but leaves artifacts that degrade ASR and TTS; segment-level exclusion with logged timestamps is cleaner. Redacting names or account numbers in call-in segments with tones or silence affects model behavior, as covered in how audio redaction affects speech model training.

Transcripts carry their own rights: show notes and professional transcripts may be licensed separately from the audio. Decide whether you need verbatim or clean transcripts before you price them (see verbatim vs clean transcription standards). If you also want the text of a broadcaster's news archive, its licensing questions are covered in news archive text for AI training.

Sourcing podcast and broadcast audio with SourceX

SourceX sources operational datasets from US companies on request, including documents and new recordings of hands-on work, and manages licensing and ongoing purchases; it does not source scraped web content, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, released only with the supplying company's approval, and delivered under a license that defines records, uses, term and delivery. For the wider audio landscape, start with the speech and audio data buyer's guide, then describe the audio you need on the SourceX buyers page.

Frequently asked questions

Is public podcast audio free to use for AI training?

No. Public availability on an RSS feed or streaming platform says nothing about permission to train. Copyright in the episode, publicity rights in voices and biometric statutes such as BIPA all apply independently, and the 2026 voiceprint suits target exactly this assumption [1][3].

Does a guest release form cover model training?

Only if its language reaches that use. Releases drafted for distribution "in all media" were rarely written with training or voice synthesis in mind, so counsel should read each template, and new recordings should use AI-specific consent.

Can we keep music beds if we only train ASR?

Treat unlicensed music as content to remove regardless of the model's purpose. The composition and the sound recording are separate rights, and a podcast sync license usually covers only the show itself.

Sources

  1. Biometric Update, "Tech giants sued under BIPA over voiceprints used to train AI" (2026). https://www.biometricupdate.com/202605/tech-giants-sued-under-bipa-over-voiceprints-used-to-train-ai
  2. WGN-TV, "Illinois journalists, podcasters sue tech companies over using their voice to train AI" (2026). https://wgntv.com/news/illinois/illinois-journalists-podcasters-sue-tech-companies-over-using-their-voice-to-train-ai/amp/
  3. Kaamel, "When does training on public audio become voiceprint collection?" (2026). https://www.kaamel.com/blog/en-public-audio-ai-voiceprint-bipa/
  4. arXiv, "ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics (arXiv:2411.09349)" (2024). https://arxiv.org/pdf/2411.09349
  5. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  6. Illinois General Assembly, "SB 2979 (103rd General Assembly), BIPA amendment, engrossed text" (2024). https://www.ilga.gov/documents/legislation/103/SB/PDF/10300SB2979eng.pdf
  7. Texas Legislature, "Texas Business and Commerce Code Section 503.001, Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  8. Tennessee General Assembly, "HB 2091 (ELVIS Act) bill information" (2024). https://wapp.capitol.tn.gov/apps/Billinfo/default.aspx?BillNumber=HB2091&ga=113
  9. Skadden, Arps, Slate, Meagher & Flom, "New York court tackles the legality of AI voice cloning" (2025). https://www.skadden.com/-/media/files/publications/2025/07/new_york_court_tackles_the_legality_of_ai_voice_cloning.pdf
  10. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  11. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (3d Cir.)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  12. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  13. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  14. Perkins Coie, "New biometrics lawsuits signal potential legal risks for AI" (2026). https://perkinscoie.com/insights/update/new-biometrics-lawsuits-signal-potential-legal-risks-ai

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data