Speech and audio data
Multilingual ASR Training Data: Sourcing Hours, Speakers and Dialects per Language
Quick answer
Multilingual ASR data should be planned and bought one language at a time, not as a single "multilingual" bundle. For each target language, specify transcribed hours, unique speakers, dialect and accent spread, domain, channel (8 kHz telephony or wideband), and transcription conventions, then confirm that the license permits commercial training. Public corpora such as Common Voice cover read speech in many languages; spontaneous, domain-specific and telephony audio usually has to be licensed or collected per language [4].
By SourceX Editorial · Updated
Why a per-language plan beats a multilingual bundle
A per-language plan works because ASR error drops with data on a curve that is steep at first and flat later, so each language sits at a different point on its own curve [10]. A bundle with a large headline total across dozens of languages can still concentrate most of its hours in a few high-resource languages, and the total says nothing about the license or the hours in your target languages [3]. Averaged word error rate (WER) across languages hides exactly the languages you are expanding into.
Treat each language as its own sourcing line item with its own acceptance test. The hours guide for fine-tuning ASR covers how to size the volume; this page covers what to specify and where the supply tends to come from.
The six fields to specify for each language
Every language request needs six fields: hours, unique speakers, dialect spread, domain, channel, and transcription standard. Hours alone is a weak specification, because 500 hours from 40 speakers trains a model that memorizes voices rather than learning a language.
- Transcribed hours, net of silence. Ask whether hours are file duration or speech duration after voice activity detection. Vendors may count either.
- Unique speakers and hours per speaker. Set a cap, such as no single speaker contributing more than 1% of a language's hours, and require a stable pseudonymous speaker ID so you can split train and test by speaker.
- Dialect and region spread. For Arabic, name the varieties (Egyptian, Levantine, Gulf, Maghrebi) and whether Modern Standard Arabic is in scope. For Spanish, separate Mexican, Caribbean, Rioplatense and Peninsular. For Hindi, decide whether Hinglish code-switching counts or is excluded.
- Domain. Customer support, sales calls, field service or medical dictation each carry different vocabulary; see domain vocabulary coverage in speech data.
- Channel and format. Telephony (8 kHz, often G.711 mu-law, mono or stereo split by speaker) and wideband (16 kHz or higher, PCM WAV or FLAC) are different training conditions. See telephony vs wideband ASR training.
- Transcription standard. Specify script, numeral style, disfluency tags and whether transcripts are verbatim or clean. The verbatim vs clean transcription standards guide lists the tag decisions.
Document all six in a data statement per language, using the University of Washington Version 2 schema fields for language variety, speaker demographics and speech situation [9].
Per-language sourcing plan template
The artifact below is a template a speech team can fill in before contacting any supplier. Each row is one language and channel combination, because telephony Spanish and wideband Spanish are separate procurement problems.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Language (BCP 47) | Target hours | Min unique speakers | Dialect spread required | Domain | Channel | Transcript standard | Likely supply route |
|---|---|---|---|---|---|---|---|
| es-US, es-MX | 1,000 | 800 | Mexican, Caribbean, Central American; 10% code-switched | Support and sales calls | 8 kHz stereo | Clean verbatim, Latin script, numerals as spoken | US business recordings |
| vi-VN, vi-US | 300 | 250 | Northern and Southern | Support calls | 8 kHz mono | Verbatim with tone diacritics | US business recordings plus commissioned collection |
| ar-EG, ar-SA | 400 | 300 per variety | Egyptian, Gulf; MSA excluded | Customer service | 8 kHz | Arabic script, dialect spelling guide attached | In-country collection |
| sw-KE | 150 | 200 | Kenyan Swahili, with English insertions tagged | Field service | Wideband 16 kHz | Verbatim, code-switch tags | Public read speech plus commissioned spontaneous speech |
| el-GR | 60 | 80 | Standard Greek | Medical dictation | Wideband | Clean, drug names normalized | Public corpora plus domain text for LM adaptation |
Use the "Min unique speakers" column as a hard acceptance criterion. Use "Likely supply route" to decide early which languages need in-country collection, since those take the longest to arrange.
Where per-language speech supply actually comes from
Per-language supply comes from four routes: public corpora, consortium catalogs, commercial catalog vendors, and licensed or commissioned recordings. Each route has a different rights profile, and most multilingual programs combine at least two.
Public corpora. Mozilla Common Voice is crowdsourced and released into the public domain, with per-language releases that grew from 29 languages in late 2019 [4]. It is mostly read sentences from volunteers, so it helps with acoustic coverage and speaker diversity but not with spontaneous, domain or telephony speech. Low-resource teams often start here: one Greek medical ASR effort fine-tuned on roughly 49 hours of public crowd-sourced and read speech and adapted the language model with a separate medical text corpus [2].
Consortium catalogs. The Linguistic Data Consortium licenses many conversational telephone corpora, and commercial use runs through for-profit terms, such as the for-profit membership agreement, rather than the research terms many papers cite [6]. Check which agreement covers the specific corpus version you plan to train on.
Commercial catalogs. Vendors list language-specific call-center audio. Marketplace listings can be free to access while still restricting commercial training, so the access price tells you nothing about the license [3].
Licensed operational recordings and commissioned collection. Real customer calls, sales calls and recorded field work from businesses give spontaneous, domain-specific speech that read corpora lack. Languages widely spoken by US customers, such as Spanish, Chinese, Vietnamese and Tagalog, can appear in US business recordings. Languages with few US speakers usually need in-country collection under participant consent; see consent language for commissioned collection and off-the-shelf vs custom vs licensed speech data.
Commercial rights checks for non-English speech
Commercial rights for non-English speech must be checked per corpus, per language and per version, because licenses vary inside a single collection. An audit of more than 1,800 datasets found license information missing for over 70% of entries on popular hosting sites and license errors in over 50% [5].
Some language-specific corpora use custom licenses written to permit AI training, including commercial use, on content whose owners gave permission, as one Hebrew speech corpus does [1]. Read those licenses to the clause, not to the summary badge. The license audit for open speech corpora walks through common traps, such as non-commercial clauses inherited from source audio.
Voice is also biometric data in several US states. Illinois BIPA treats voiceprints as biometric identifiers and requires written notice and a written release before collection [7]. Texas bars capturing a voiceprint for a commercial purpose without prior notice and consent [8]. Ask suppliers how speaker consent covered AI training, and whether speaker-level opt-outs can be honored after delivery.
Rights checklist per language:
- License text names commercial model training, not only "research" or "evaluation".
- Consent records cover AI training for every speaker, including callers in recorded customer calls.
- Translation, transcription and re-recording rights are held by the party licensing them.
- Personal details in transcripts and audio are removed or replaced, with the method recorded. See how audio redaction affects speech model training.
- Jurisdiction-specific limits are identified, for example, the UK s29A text and data analysis exception covers only non-commercial research.
Native-speaker transcription QA is the hidden cost driver
Transcription QA by native speakers is often a major cost line in low-resource languages, because qualified reviewers are scarce. A transcript from a non-native annotator can look clean and still mis-segment words, drop tone marks in Vietnamese or Yoruba, or standardize dialect spellings in ways that hurt recognition.
Specify QA in the request rather than discovering it at acceptance:
- Double-pass review by native speakers of the target variety on a stated sample, such as 5% of hours.
- A target transcript WER between two independent annotators, per language.
- A per-language style guide covering script (for example, Serbian Cyrillic or Latin), numerals, named entities and English loanwords.
- Language identification checks on every segment, since mislabeled language is a common silent error; see language identification QA for multilingual datasets.
Code-switched speech needs its own tagging rules; the code-switched Spanish-English speech data spoke covers them.
Acceptance tests before you train
Accept each language delivery against a held-out test set from speakers who are not in training. Measure WER or character error rate (for Chinese, Japanese and Thai) per dialect and per channel, not as a single average.
Before training, verify the following:
- Speaker IDs are consistent, and no speaker appears in both train and test.
- The hours-per-speaker distribution matches the cap in the request.
- The dialect mix matches the requested spread within an agreed tolerance.
- Sample rate, codec and channel layout match the specification; resampled 8 kHz audio delivered as 16 kHz is a common mismatch. See audio file specs.
- Redaction tags and beeps follow a documented convention that the tokenizer handles.
How SourceX fits a multilingual ASR sourcing plan
SourceX sources operational datasets, including support and sales call histories and new recordings of hands-on work, from US companies and manages the licensing process and ongoing purchases. Data is sourced on request, not held in stock, so a language listed in your plan is a request, not a promise of a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.
SourceX looks for US businesses that hold the data you describe, and each supplying company approves every release. That makes it most relevant for languages spoken by US customers and workers. You can describe your per-language speech requirements to SourceX using the six fields above. For broader context, see the speech and audio data buyer's guide, training data for multilingual enterprise models, whether AI labs buy non-English business data and licensing voice and audio data.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Source multilingual speech data for your ASR languages
Share the languages, hours, speakers, dialects and channels you need, and SourceX will look for US businesses that hold matching recordings. Personal details are removed or replaced before delivery, and nothing is contracted until a supplier agrees. Start a multilingual speech data request.
Sources
- arXiv, "Hebrew speech corpus with a commercial-training license and per-item owner permission (arXiv:2307.08720)" (2023). https://arxiv.org/pdf/2307.08720
- arXiv, "Greek medical ASR fine-tuned on about 49 hours of public crowd-sourced and read speech plus domain text (arXiv:2509.23550)" (2025). https://arxiv.org/pdf/2509.23550
- AWS Marketplace, "Multilingual call-center ASR audio listing". https://aws.amazon.com/marketplace/pp/prodview-yyfwirpya2mp6
- Ardila et al. (Mozilla), arXiv:1912.06670, "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
- Longpre et al., arXiv:2310.16787, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Linguistic Data Consortium, University of Pennsylvania, "LDC For-Profit Membership Agreement". https://Catalog.Ldc.Upenn.Edu/license/ldc-for-profit-membership.pdf
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14)". https://ilga.gov/Legislation/ILCS/Articles?ActID=3004&ChapterID=57&Print=True
- Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/
- Hestness et al., Baidu Research, arXiv:1712.00409, "Deep Learning Scaling is Predictable, Empirically" (2017). https://arxiv.org/abs/1712.00409v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.