Skip to content

Speech and audio data

Can You Use Open Speech Corpora Commercially? A License Audit for ASR and TTS Teams

Quick answer

Sometimes. CC0 or public-domain releases and CC BY corpora can generally train a commercial ASR or TTS model with attribution. CC BY-SA corpora can too, but counsel must decide what share-alike reaches. Anything marked NC, ND, "research only" or a custom academic EULA usually cannot, even when a commercial vendor hosts it. The corpus license also covers the compilation, not necessarily each speaker's voice or consent. Audit every corpus in the mix against its original license text, then replace what fails.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

How speech corpus licenses sort into five commercial-use buckets

Speech corpora fall into five license patterns, and only the first three are routinely usable for commercial training. The bucket decides whether a corpus stays in the mix, needs a legal review, or must be replaced. For a cross-domain view of the same licenses, see the open data license compatibility matrix.

  • Public domain or CC0. Mozilla describes Common Voice as a crowdsourced, public-domain speech corpus released in numbered versions [7]. Pin the exact release you trained on, because the corpus changes between versions.
  • Permissive (CC BY). Attribution is the main obligation. Augmentation assets often use this pattern; one room-impulse-response database from Brno University of Technology is released under CC BY 4.0 [4].
  • Share-alike (CC BY-SA). The People's Speech, about 30,000 hours of English, is described by its authors as CC-BY-SA for academic and commercial use [1]. Commercial training is intended, but redistributed adaptations of the corpus carry share-alike terms, and whether model weights count as an adaptation is a question for counsel.
  • Non-commercial (CC BY-NC) and research-only. A recent call-center corpus is released under CC BY-NC 4.0 for non-commercial research only [2]. A production ASR model for a paying product is outside that grant.
  • No-derivatives (ND) and custom EULAs. A commercial vendor's speech emotion dataset on Hugging Face is listed as CC BY-NC-ND 4.0 [3]. Who publishes a dataset tells you nothing about its terms.

Why the Hugging Face license tag is not your audit record

The license field on a hub listing is a starting point, not evidence. On Hugging Face, the license sits in YAML metadata at the top of the dataset card README and drives search filters [8]. Anyone uploading a mirror can set or omit it.

The Data Provenance Initiative found license omission above 70% and license error rates above 50% on popular dataset hosting sites [6]. Its audit covered text datasets, but speech mirrors are built the same way: re-uploads, re-segmented shards and "cleaned" forks that drop the original LICENSE file. Trace each corpus back to the releasing institution's page or paper, store a copy of the license text, and record the version or commit hash you downloaded. The broader method is covered in auditing open dataset licenses before commercial training.

A corpus license grants rights in the compilation and transcripts as the releaser held them; it may not settle every speaker's rights. Corpora assembled from internet audio, such as The People's Speech [1], depend on the upstream licenses of each source recording. Crowdsourced corpora depend on the contributor terms each speaker accepted.

Voice can also be regulated as biometric data. Texas defines a biometric identifier to include a voiceprint and requires notice and consent before capturing one for a commercial purpose [9]. If your pipeline extracts speaker embeddings for diarization, voice cloning or speaker verification, ask whether the original consent covered that use. TTS is the sharpest case: a model that reproduces a recognizable voice raises consent questions that a CC BY tag does not answer. For purpose-collected alternatives, see consented studio voice recordings for TTS.

Jurisdiction does not convert a non-commercial corpus into a commercial one

Copyright exceptions for text and data mining rarely rescue an NC corpus used for a commercial model. In the UK, CDPA section 29A permits copies for computational analysis only for non-commercial research [10], and as of October 2026 that limit remains in place. A research team inside a company may experiment under an NC license only if its terms and counsel allow it; the moment those checkpoints feed a product, the analysis changes.

Disclosure obligations add pressure. Providers placing general-purpose AI models on the EU market must publish a training-content summary using the AI Office template dated 24 July 2025 [11]. A corpus you cannot describe by source and license is hard to report. Reviewers increasingly flag vague labels: a 2025 position paper on TTS evaluation notes that some technical reports describe training data only as "in-house" [5].

Speech corpus license audit worksheet

Run one row per corpus, per version, before a training run is approved. The worksheet below shows the fields counsel and the ML lead should both sign off on.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldWhat to recordFail condition
corpus_id and versionRelease tag, download date, commit hash or checksumUnknown version or mirror-only copy
license_as_publishedExact license text from the releaser, stored in your repoLicense taken only from a hub tag
license_bucketPD/CC0, CC BY, CC BY-SA, NC/research-only, ND/customNC, ND or research-only for a commercial model
attribution_textRequired credit string and where it will appearNo plan for attribution
share_alike_scopeCounsel's view on whether weights or outputs are adaptationsUnresolved for CC BY-SA
audio_provenanceCollected, crowdsourced, broadcast, scraped, vendor-recordedScraped with no upstream license trail
speaker_consent_basisContributor terms, release forms, collection noticeNo consent record for voice cloning or speaker ID use
biometric_useEmbeddings, diarization, verification, voice cloningBiometric use not covered by consent
intended_useASR training, TTS, emotion recognition, eval onlyUse outside the license grant
decisionKeep, keep with conditions, replace, eval only if allowedBlank

Store the worksheet with the training config so the record survives staff changes. An eval-only set still needs a row; public benchmark terms are covered in checking eval dataset licenses.

Common failure modes in speech license audits

Most audit failures come from inherited data rather than the corpora a team chose deliberately. Look for these patterns:

  • Derived corpora that inherit the strictest parent. A "combined English ASR set" that mixes a CC BY corpus with an NC call-center corpus [2] is effectively NC.
  • Augmentation assets. Noise, music and impulse-response sets need their own rows; a permissive one [4] is fine, while an unlabeled noise pack is not.
  • Pseudo-labeled audio. Transcripts generated by a third-party model do not clean up the license of the underlying audio.
  • Emotion and paralinguistic sets. Emotion corpora often carry NC or ND terms; one vendor-published set is CC BY-NC-ND [3]; see acted vs naturalistic emotion corpora.
  • Inherited checkpoints. A model fine-tuned on an NC corpus and later used as a teacher can carry the problem forward.

What to replace a failed corpus with

When a corpus fails, replace it with data whose rights trail you can document, matched to the same acoustic role. A failed NC call-center corpus is usually replaced with narrowband conversational audio; a failed acted-emotion set with naturalistic recordings under a negotiated license. Specify sample rate, channels, transcription standard and speaker metadata up front, using the Speech and audio data hub and the research-only to commercial license guide to scope the request.

Three replacement routes exist: ask the original releaser for a commercial license, commission new recordings, or license operational audio such as support calls from a company that holds it. Compare them in off-the-shelf vs custom vs licensed speech data. Whatever the route, get a written license defining records, permitted uses, term and delivery; the AI data license terms guide and dataset licensing glossary entry explain the clauses.

SourceX sources operational datasets, including support and sales histories, from US companies on request, and every dataset is rights-reviewed for ownership and consents before delivery under a license. Buyers can describe the audio they need through the SourceX buyer intake.

Licensed speech data for commercial ASR and TTS

SourceX looks for US businesses that hold the speech or call data you describe and manages the licensing process; nothing is contracted until a supplier agrees, and a request does not guarantee a match. Personal details such as names and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked. Describe your speech data requirements to SourceX.

Frequently asked questions

Can I train a commercial ASR model on a CC BY-NC speech dataset?

Generally no. The non-commercial condition covers the use of the data, and training a model that ships in a paid product is commercial use under most readings; the call-center corpus in [2], for example, is limited to non-commercial research. Ask the releaser for a separate commercial license or replace the corpus.

Is a CC BY-SA speech corpus safe for a closed-weights model?

It is designed for commercial use [1], but share-alike applies to adaptations you share. Whether trained weights are an adaptation is unsettled, so record counsel's position in the audit row before you ship.

Does a commercial vendor's Hugging Face listing mean commercial use is allowed?

No. Vendor-published datasets can carry NC-ND terms [3], and hub tags are often missing or wrong [6]. Read the license text itself.

Sources

  1. arXiv (Galvez et al.), "The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage" (2021). https://arxiv.org/pdf/2111.09344
  2. arXiv, "CallCenterEN: 91706 Real-World English Call Center Transcripts Dataset with PII Redaction" (2025). https://arxiv.org/abs/2507.02958
  3. Hugging Face / TrainingDataPro, "speech-emotion-recognition-dataset (dataset card)". https://www.huggingface.co/datasets/TrainingDataPro/speech-emotion-recognition-dataset
  4. Brno University of Technology, Faculty of Information Technology, "Room impulse response database publication (IEEE, BUT)" (2019). https://www.fit.vut.cz/research/publication-file/c159973/280107/IEEE_Final_Published_08717722-1.pdf
  5. arXiv, "Position: Towards Responsible Evaluation for Text-to-Speech" (2025). https://arxiv.org/pdf/2510.06927
  6. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  7. arXiv (Ardila et al., Mozilla), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
  8. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  9. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier" (2026).(2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  10. UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 - Consolidated (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  11. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data