Speech and audio data
Voiceprints and BIPA: Biometric Risk in Licensed Voice Datasets
Quick answer
Voice recordings are not automatically biometric data, but Illinois' BIPA lists "voiceprint" as a biometric identifier without defining it, and 2026 class actions argue that training speech models on recorded voices extracts voiceprints. As of October 2026 those claims are unresolved. A buyer licensing voice data for ASR, TTS or speech-LLM training should therefore require evidence of Illinois speaker exposure, written releases, a published retention and destruction schedule, and a clear record of whether speaker-identifying features were ever derived.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What BIPA actually says about voiceprints
BIPA covers voiceprints by name, but the statute leaves the technical meaning to litigation. Section 10 defines "biometric identifier" as a retina or iris scan, fingerprint, voiceprint, or scan of hand or face geometry, and excludes items such as photographs and written signatures; "biometric information" is information based on an identifier that is used to identify an individual [1]. The Act does not say whether a raw waveform, an MFCC frame or a speaker embedding is the "voiceprint", and commentators identify that gap as the central dispute in the AI training suits [5].
The operative duties sit in Section 15 [1]:
- 15(a): a private entity in possession of biometric identifiers must publish a written retention schedule and destruction guidelines.
- 15(b): before collecting, capturing, purchasing or receiving an identifier, it must inform the person in writing, state the specific purpose and length of term, and obtain a written release.
- 15(c) and 15(d): it may not sell, lease, trade or otherwise profit from the data, and disclosure needs consent or another listed exception.
- 15(e): it must protect the data with a reasonable standard of care.
Section 20 gives "any person aggrieved" a private right of action with liquidated damages per violation, higher for intentional or reckless violations [1]. The 2024 amendment (SB 2979) made repeated collection of the same identifier from the same person by the same method a single violation, and let an electronic signature count toward the written release [2][3]. That narrows per-scan multiplication but does not remove per-person exposure, which is the number that matters for a corpus with thousands of speakers.
Why the 2026 voice suits matter to dataset buyers
The May 2026 suits target model developers, but the duties they invoke also reach entities that buy or receive data. Reporting describes class actions filed in the Northern District of Illinois alleging that voiceprints were extracted from recorded speech to train AI without written consent, written notice or published retention policies [4]. These are allegations, not findings, and none had been decided on the merits as of October 2026.
Section 15(b) applies to an entity that collects, captures, purchases, receives through trade or otherwise obtains identifiers, and 15(c) restricts profiting from them [1]. If a court accepts that a speech corpus contains voiceprints, a licensee is not insulated simply because a supplier did the recording. Earlier BIPA claims against AI transcription tools show plaintiffs already treat speech processing pipelines as voiceprint collection [7].
When does a recording become a voiceprint?
The defensible working answer is that risk rises with how directly a pipeline isolates and stores speaker identity. Commentary on public-audio training distinguishes a recording, which captures speech, from a biometric template, which encodes who is speaking for later matching [6]. Courts have not settled where model training falls on that line [5].
A practical way to map your own exposure is by artifact and use:
| Artifact or step | What it encodes | Relative voiceprint risk |
|---|---|---|
| Raw WAV/FLAC recordings, no speaker labels | Speech content plus incidental voice traits | Contested; plaintiffs argue training extracts voiceprints |
| Transcripts only (text) | Words, no acoustic features | Low under biometric law; still personal data if identifiable |
| Diarization labels (RTTM speaker A/B) | Who spoke when, within a file | Moderate; turns are separated but not linked to a person |
| Persistent speaker IDs across files | Same speaker linked across sessions | Higher; supports cross-session identification |
| Speaker embeddings (x-vectors, ECAPA-TDNN, d-vectors) stored as features | Fixed-length identity vectors | High; closest to a template used for matching |
| Speaker verification or voice cloning objective | Model trained to identify or reproduce a specific voice | High; purpose is identity |
ASR trained on pooled audio sits in the contested middle. TTS and voice-cloning work that conditions on a named speaker, and any speaker-recognition project, sit at the high end and should be priced and documented as biometric collection from the start. For the cross-modality picture, including Texas and Washington, see the biometric data in AI training guide.
Public audio is not consent
Public availability of a podcast, video or broadcast does not, by itself, satisfy a written-release requirement. BIPA's 15(b) is framed around informed written consent before collection, with no public-availability exception in Section 10's exclusions [1], and commentary treats the public-audio argument as unresolved [6].
Texas is different in structure. Business and Commerce Code 503.001 also lists voiceprints and requires notice and consent before capture for a commercial purpose, enforced by the attorney general rather than private suits; the version effective January 1, 2026 adds references to artificial intelligence systems via HB 149 [8][9]. Read the current text for the exact scope of any AI-training carve-out and its limits before relying on it.
Diligence evidence to request from a voice data supplier
The evidence that holds up is speaker-level, dated and tied to the specific recordings you receive. Ask for these items before signature, and make delivery conditional on them.
Illustrative example: invented to show structure; it does not describe an available dataset.
Voiceprint diligence checklist (per dataset)
- Speaker geography: count of speakers with Illinois, Texas or Washington residence or recording location, and how it was determined (enrollment form, IP, site address).
- Release text: the exact consent form shown to speakers, naming biometric collection, the AI-training purpose and the retention term, with version and date.
- Signature evidence: per-speaker signed or e-signed release IDs linked to
speaker_id, with timestamp and form version. - Retention schedule: the supplier's published 15(a) policy and the destruction date that applies to this corpus.
- Derived features: a statement of whether embeddings, speaker IDs or verification scores were ever computed, stored or shipped.
- Chain of receipt: whether the supplier itself purchased the audio, and the upstream release terms if so.
- Withdrawal handling: how a speaker's withdrawal propagates to the licensee and to derived artifacts.
- Third-party voices: treatment of non-consenting voices captured in the background or on the other end of a call.
A manifest field set that makes this auditable might look like:
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"file": "rec_000412.flac",
"speaker_id": "spk_0193",
"speaker_state": "IL",
"release_id": "rel_2026_0193",
"release_form_version": "v3.2-biometric",
"release_signed_at": "2026-03-14T16:02:11Z",
"biometric_purpose_disclosed": "ASR and TTS model training",
"retention_destroy_by": "2029-03-14",
"speaker_embeddings_shipped": false,
"other_party_consent": "not_applicable_single_speaker"
}
Call audio adds a separate wiretap layer, since Illinois is an all-party consent state for recording; that question is covered in call-recording consent checks for AI training and the Illinois call-recording law summary. Talent-recorded corpora are governed by release terms covered in voice talent consent and release terms.
Mitigations that reduce, but do not remove, exposure
Technical controls narrow the biometric surface, yet none converts a voice into non-biometric data with certainty. The VoicePrivacy 2024 Challenge evaluates speaker anonymization against attacker speaker-verification models while measuring ASR and emotion utility, which shows anonymization is a measured trade-off rather than a guarantee [10].
- Exclude or segregate: drop speakers lacking a biometric-specific release, or hold Illinois speakers out until releases are confirmed.
- Do not ship identity features: require transcripts and audio without embeddings or cross-file speaker IDs unless the use case needs them.
- Purpose-bound licenses: restrict speaker verification and voice cloning unless releases cover them.
- Destruction alignment: match your internal deletion of raw audio to the supplier's published schedule.
- Synthetic substitution: consider TTS-generated speech for augmentation; the limits are covered in real vs synthetic speech data.
Remember that names, phone numbers and account numbers spoken in audio are separate personally identifiable information issues, and redacting them from transcripts does not address the voice itself.
How SourceX handles voice data requests
SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, with diligence materials prepared per dataset. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Buyers can describe the voice data they need and the consent evidence they require. More speech sourcing guidance is in the speech and audio data hub and the main AI data buyer guides.
Sourcing voice data with BIPA evidence in hand
If your ASR, TTS or speech-LLM program needs voice data with documented consent and rights review, start by describing the recordings, speakers and allowed uses rather than a supplier. SourceX looks for US businesses that hold the described data, assesses data and licensing permissions, and nothing is contracted until a supplier agrees. Submit a voice data request to SourceX.
Sources
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)" (current). https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
- Illinois General Assembly, "SB 2979 (103rd General Assembly), engrossed text" (2024). https://www.ilga.gov/documents/legislation/103/SB/PDF/10300SB2979eng.pdf
- Davis Wright Tremaine, "Illinois Revises Biometrics Law To Reduce the Prospect of Ruinous Damage Awards" (2024). https://dwt.com/blogs/privacy--security-law-blog/2024/08/illinois-bipa-biometrics-law-amended-for-damages
- Biometric Update, "Tech giants sued under BIPA over voiceprints used to train AI" (2026). https://www.biometricupdate.com/202605/tech-giants-sued-under-bipa-over-voiceprints-used-to-train-ai
- Sigma Law Group, "Voiceprint Class Actions Bring Illinois Biometric Law to AI Training Data" (2026). https://sigmalawgroup.com/blog/2026-08-22-bipa-voiceprint-ai/
- Kaamel, "When Does Public Audio Training Become Voiceprint Collection?" (2026). https://www.kaamel.com/blog/en-public-audio-ai-voiceprint-bipa/
- Lewis Rice, "AI Transcription Tools Give Rise to BIPA Claims". https://www.lewisrice.com/publications/ai-transcription-tools-give-rise-to-bipa-claims
- Texas Legislature, "Texas Business and Commerce Code Section 503.001, Capture or Use of Biometric Identifier" (effective 2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- Texas Legislature, "H.B. No. 149 (89th Legislature), enrolled text" (2025). https://capitol.texas.gov/tlodocs/89R/billtext/html/HB00149F.htm
- VoicePrivacy Challenge organizers (arXiv), "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.