Speech and audio data
Accented English Speech Data for ASR: Coverage Beyond US Read Speech
Quick answer
An accented English speech dataset worth buying pairs conversational audio from many first-language (L1) backgrounds with speaker-level accent and L1 metadata, verbatim transcripts and a commercial license. The gap it closes is large: one benchmark found a strong model averaged 19.7% word error rate (WER) on accented conversational English against 2.7% on US clean read speech [1]. Specify accents by deployment share, demand speaker profiles, prefer spontaneous speech and audit licenses before any fine-tuning run.
By SourceX Editorial · Updated
How large is the accent gap in ASR accuracy?
The accent gap is large enough that average WER on a US-centric test set says little about how your product performs for international users. The Edinburgh International Accents of English Corpus (EdAcc) authors report that the best model they evaluated, trained on 680k hours of transcribed data, averaged 19.7% WER on their accented conversational set versus 2.7% on US English clean read speech, with notable drops for Indian, Jamaican and Nigerian English [1]. That is roughly a sevenfold difference, and it compounds two shifts at once: accent and speaking style.
A separate audit of popular ASR services using more than 2,700 speakers from 171 countries found disparities tied to whether English is the speaker's first language [2]. For a buyer, this means "accent" is not one variable. L1 background, country of English acquisition, years of exposure and speaking context all move error rates, and a dataset that collapses them into a single "non-native" flag will hide where your model fails.
Practical consequence: before sourcing, slice your own production WER by user region or locale, and by whatever accent proxy you lawfully hold. That slice table becomes the target distribution for the purchase, and it is the same exercise described in our guide to coverage gap analysis against your deployment distribution.
Which accent and speaker metadata should a dataset carry?
A usable accented English corpus carries metadata at the speaker level, not just at the file or collection level. The AnglistikVoices authors note that existing L2 English corpora mostly lack detailed linguistic speaker profiles and are largely crowd-sourced [3], which is exactly the problem buyers hit when they try to rebalance or evaluate by accent.
Ask for these fields per speaker, keyed by a stable pseudonymous speaker ID:
- L1 (first language) as an ISO 639-3 code, plus any additional languages spoken at home.
- Accent or variety label (for example Indian English, Nigerian English, Singapore English, Scottish English), with the labeling method recorded.
- Country of English acquisition and age of onset, which separate a speaker raised in Lagos from one who learned English as an adult elsewhere.
- Self-rated or assessed proficiency for L2 speakers, using a named scale such as CEFR bands.
- Age band and gender as collected, only if consented and needed for bias audits.
- Recording context: device class, channel (telephony 8 kHz or wideband), environment and speaking style (read, prompted, spontaneous).
Document the whole set with a data statement; the University of Washington Tech Policy Lab's Version 2 schema has fields for speaker demographics and language variety that map well onto speech corpora [8].
Self-reported vs expert-assessed accent labels
Self-reported and expert-assessed accent labels answer different questions, so decide which you need before you write the request. Self-reported labels (the approach used by crowdsourced corpora such as Mozilla Common Voice, which collects volunteer contributions [4]) are cheap and scale, but speakers describe their accents inconsistently: "Indian English," "South Indian," and "Tamil" may all name overlapping groups.
Expert-assessed labels, where a trained phonetician or rater listens and assigns a variety and strength rating, are slower and costlier but give you consistent strata for evaluation. A common working pattern (a hypothesis to test on your data) is to accept self-reported labels for training data and require assessed labels, or at least a normalized taxonomy with a reconciliation pass, for the held-out evaluation slice. Whatever you choose, ask the supplier to record which method produced each label so you can filter later.
Why conversational accented speech matters more than read speech
Conversational accented speech is closer to what production voice agents and call transcription see, so it usually earns a higher priority than read prompts. Read speech flattens the features that hurt ASR most: disfluencies, reduced forms, code-switched words, rapid turn-taking and prosody that differs from US norms. The EdAcc numbers above mix accent shift with a move from read to conversational speech [1], and the CASPER authors built 200 hours of recorded casual English conversation precisely because spontaneous speech is scarce relative to scripted material [5].
Read accented speech still has a role: it gives controlled phonetic coverage and is easy to align for pronunciation work. For a production ASR fine-tune, though, a reasonable buyer stance is to require that most hours be spontaneous or task-based dialogue, and to treat read prompts as a supplement. Our page on spontaneous conversational speech data covers how to specify turn structure and disfluency transcription; pair it with your verbatim vs clean transcription standard so filled pauses and false starts are labeled consistently across accents.
Operational recordings are one realistic source of spontaneous accented English: customer support lines, sales calls and service desks serving international customers. If that is your target domain, see call center audio datasets and the notes on training ASR on 8 kHz telephony audio.
Writing the accented speech request spec
A good request specifies accents as a weighted distribution tied to your users, not a list of every accent you can name. Start from the WER slice table, weight accents by traffic share and error severity, then set a floor of speakers (not just hours) per accent so a few prolific speakers cannot dominate a stratum.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example requirement | Why it matters |
|---|---|---|
| Target accents and weights | Indian English 30%, Nigerian English 15%, Filipino English 15%, Mexican Spanish L1 15%, UK regional 10%, other 15% | Matches deployment traffic and error severity |
| Minimum distinct speakers per accent | 150, no speaker over 1% of stratum hours | Prevents speaker overfitting and leaky splits |
| Speaking style mix | 70% spontaneous dialogue, 20% task-based, 10% read | Production speech is mostly unscripted |
| Channel | 60% telephony (8 kHz, mono), 40% wideband (16 kHz+, WAV/FLAC) | Matches the inference audio path |
| Speaker metadata | L1 (ISO 639-3), variety label and method, acquisition country, proficiency band | Enables stratified training and fairness evaluation |
| Transcription | Verbatim, disfluency tags, non-English words tagged with language | Keeps code-switching and fillers learnable |
| Evaluation split | Speaker-disjoint, assessed accent labels, held out before delivery | Reliable per-accent WER |
| Privacy | PII in audio and transcript removed or replaced; method documented | Personal details commonly occur in conversational audio |
| License | Commercial training and evaluation, named records, term, delivery method | Avoids share-alike or non-commercial surprises |
Two common failure modes: buying hours without a speaker floor (a stratum of 200 hours from 12 speakers generalizes poorly), and leaving the evaluation split to random utterance sampling, which lets the same voice appear in training and test. For volume planning, see how many hours of audio you need to fine-tune ASR.
Open accented corpora: license and coverage checks
Open accented English corpora are useful for benchmarking but need a license audit before any commercial training use. Some carry share-alike terms; the EdAcc paper describes its release terms, which buyers should read against their own use [1]. Common Voice is public domain [4], which simplifies licensing but leaves you with crowdsourced, mostly read-prompt audio and self-reported metadata.
Do not rely on the license field of a hosting site alone. A large audit of more than 1,800 text datasets found license omission rates above 70% and error rates above 50% on popular hosting platforms [9]. Trace each corpus to its original release, record the license version, and check whether derived transcripts or alignments carry different terms. Our license audit for open speech corpora walks through that check for ASR and TTS teams.
Coverage is the second check. Many open L2 corpora are small, read-speech heavy and thin on speaker profiles [3], so they rarely cover the accents in a specific customer base at production volume. Use them to build an early per-accent evaluation, then source licensed conversational audio where the slice table shows the largest gaps.
Consent, voiceprints and privacy in accented speech data
Accented speech data is personal data, and in some US states a voice recording used to identify a speaker can be regulated biometric data. Texas Business and Commerce Code Section 503.001 lists a voiceprint as a biometric identifier and requires notice and consent before capture for a commercial purpose [7]. Ask suppliers how recordings were consented, whether speakers agreed to AI training use, and whether any speaker-identification features were derived.
Accent metadata adds a specific risk: L1, country of origin and accent labels can act as proxies for national origin or ethnicity. Keep them at the speaker-pseudonym level, collect only what your bias audit needs, and restrict access to the metadata table. Voice anonymization is not a solved problem; the VoicePrivacy Challenge evaluates anonymization by both privacy and ASR utility [6], and voice conversion can distort the very accent features you are paying for. Our page on how audio redaction affects speech model training covers silence, tones and transcript tags.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Where licensed real-world accented speech fits
Licensed operational recordings fill the gap between small open L2 corpora and expensive custom collection, especially for spontaneous speech in a real domain. SourceX sources operational datasets from US companies on request, including support and sales histories and new recordings of hands-on work, and manages the licensing and ongoing purchases. Datasets are not held in stock, so buyers describe the accents, channels and metadata they need, and a request does not guarantee a match.
Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. For the wider picture of sourcing options, compare off-the-shelf, custom-collected and licensed speech data, browse the speech and audio data hub, or see how to license voice and audio data. Buyers can also submit a data request once the slice table and spec are ready.
Source accented English speech data for your ASR model
Describe the accents, speaking styles, channels and speaker metadata your ASR model needs, and SourceX looks for US businesses that hold matching recordings. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees on pricing and allowed uses in a license. Start your accented English speech request.
Sources
- arXiv (ar5iv), Sanabria et al., "The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR" (2023). https://ar5iv.labs.arxiv.org/html/2303.18110
- DeepAI (paper listing), "Performance Disparities Between Accents in Automatic Speech Recognition". https://deepai.org/publication/performance-disparities-between-accents-in-automatic-speech-recognition
- Heinrich Heine University Dusseldorf, SLaM Lab, "AnglistikVoices: an L2 English speech dataset for educational and technological advancement in speech research" (2024). https://slam.phil.uni-duesseldorf.de/publication/akhilesh-2024-l2english/
- arXiv, Ardila et al. (Mozilla), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
- arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
- arXiv, VoicePrivacy Challenge organizers, "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677
- Texas Legislature, "Texas Business and Commerce Code Section 503.001, Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/
- arXiv, Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.