Skip to content

Speech and audio data

US Regional and Social Dialect Speech Data for Equitable ASR

Quick answer

US dialect speech data is conversational audio from native speakers of American English varieties, such as African American English, Southern, Appalachian, New York City, Chicano English and Upper Midwest speech, labeled with consented region and self-identified dialect metadata. Buyers need it because ASR error rates depend on speaker background [2], and clean read speech hides the gap [1]. Specify coverage per dialect group, verbatim transcripts and per-group WER reporting before you license anything.

By SourceX Editorial · Updated

Why US dialect coverage is a separate data requirement from accented English

Native dialect variation needs its own coverage targets because it differs from second-language accent in phonology, grammar and vocabulary, and both move word error rate independently. Published audits document ASR disparities that track speaker background, and the usual suspect is training data drawn from a narrow slice of speakers [2]. The EdAcc study shows the scale of the condition effect: US English clean read speech is the easy case, at 2.7% WER for the best evaluated model, while conversational accented speech averaged 19.7% [1]. EdAcc measured international accents, not US dialects, but it shows how far speaking condition alone can move error rates, so test native dialects on conversational audio too.

The practical consequence is that a dataset labeled "American English" can still be overwhelmingly one variety. If your voice agent serves callers in Atlanta, the Mississippi Delta, the Bronx and rural Kentucky, a corpus of read prompts from a narrow speaker pool will not reveal how the model handles them. Treat L2 accents as a separate spoke, covered in our guide to accented English speech data for ASR, and set native-dialect targets on their own axis.

Which US dialect groups and conditions to specify

Define coverage as a matrix of dialect group by speaking condition, because a dialect sample of read sentences tells you little about spontaneous calls. Most existing English speech corpora are read or scripted rather than spontaneous [3], yet casual conversation is where many dialect features surface. The guide to spontaneous conversational speech data explains why scripted collection undercounts disfluency and turn-taking.

Groups buyers commonly name include African American English (urban and rural, which differ), Southern and South Midland, Appalachian, New England (Eastern and Boston), New York City and Mid-Atlantic, Inland North and Great Lakes, Western and California, Chicano and Latino English, Hawaiian Pidgin-influenced English, Gullah-influenced coastal speech, and Native American English varieties. Do not treat this list as a taxonomy; sociolinguists disagree on boundaries, and speakers style-shift by context.

Conditions matter as much as groups. A contact-center voice agent needs 8 kHz narrowband telephony, cross-talk and hold-music edges; see telephony vs wideband audio. A smart-speaker product needs far-field wideband audio. Ask each supplier to state the condition mix per dialect group, not just in aggregate.

Dialect coverage specification (request template)

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldWhat to specifyWhy it matters
Dialect groupsNamed list plus "self-identified other"Prevents one variety dominating "US English"
Minimum distinct speakers per groupA floor set by your eval plan, for example 40+WER varies more across speakers than across utterances
Hours per group per conditionTelephony, wideband, far-fieldCondition and dialect interact
Speaking styleSpontaneous dialogue, task calls, readRead speech understates dialect error
Region granularityState and metro (CBSA), not ZIPUseful signal with lower re-identification risk
Dialect label sourceSelf-identified by speaker, with consentAvoids inferred race or ethnicity labels
Age band and genderSelf-reported, optional, coarse bandsSeparates dialect effects from age and gender effects
Transcription standardVerbatim, dialect forms preservedNormalizing grammar corrupts the reference
Held-out eval splitSpeaker-disjoint, per groupStops speaker leakage inflating scores

What metadata each dialect speech record should carry

Each utterance or call segment should carry speaker-level dialect metadata that the speaker supplied and consented to, not labels an annotator guessed from the audio. Inferring dialect from voice slides into inferring race or ethnicity, which is both unreliable and a sensitive-data problem. Record self-identification, the region where the speaker grew up and where they live now, since the two often differ.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "segment_id": "call_0412_seg_017",
  "audio_path": "audio/call_0412.flac",
  "start_s": 312.48,
  "end_s": 319.92,
  "channel": 1,
  "sample_rate_hz": 8000,
  "speaker_pseudo_id": "spk_9f3c",
  "role": "customer",
  "dialect_self_id": "African American English",
  "region_raised": {"state": "GA", "metro": "Atlanta"},
  "region_current": {"state": "TX", "metro": "Houston"},
  "age_band": "35-49",
  "consent_scope": ["asr_training", "asr_evaluation"],
  "transcript_verbatim": "she be working late most nights so I been calling after six",
  "redaction_spans": [],
  "condition": "telephony_landline"
}

Machine-readable documentation helps reviewers check these fields at scale. The Croissant-RAI vocabulary extends dataset documentation to cover collection, participants and intended use in a form tools can parse [6]. Package manifests and segment timing as described in speech dataset manifest packaging.

How transcription conventions can erase dialect signal

Transcribe dialect speech verbatim with dialect grammar preserved, because a transcriber who "corrects" habitual be, copula absence or multiple negation turns the reference into a different sentence. The model then learns to output standard forms the speaker never said, and your WER measurement rewards it. Write this into the transcription guideline with worked examples per dialect group, and audit a sample per transcriber.

Text normalization before scoring needs the same care. Common English normalizers map spellings and contractions, and they can collapse forms like "finna," "y'all" or "might could" in ways that hide real errors or create false ones. Freeze one normalizer, publish its rules alongside results, and check that it treats dialect lexemes consistently. Our page on verbatim vs clean transcription standards covers tag sets for disfluencies and overlaps.

Transcriber familiarity is a quality variable. Annotators unfamiliar with a variety mishear it, so ask suppliers how transcribers were matched or trained, and request inter-annotator agreement per dialect group rather than one aggregate figure.

Measuring ASR performance by US dialect

Report WER per dialect group with speaker-level confidence intervals, because a few prolific speakers can dominate an utterance-level average. Bootstrap by resampling speakers, not utterances, and report the gap between the best and worst group alongside the mean. Treat this as a repeatable audit, as published accent-disparity studies do [2]: rerun the same per-group eval on every model release.

Keep evaluation speakers disjoint from training speakers and, where possible, from the same recording sessions. Split by condition too, so a telephony regression in one group is not masked by wideband gains elsewhere. Track entity-level errors (names, street names, account phrases) per group, since a voice agent that mishears "Ponchatoula" or "Natchitoches" fails the call even if overall WER looks fine.

Where US dialect speech data comes from, and what each source allows

Public research corpora, crowdsourced collections, commissioned recordings and licensed operational call audio each cover dialects differently and come with very different license terms. Check terms first: large audits of dataset hosting sites found license information missing or wrong for most entries [7].

  • Sociolinguistic research corpora. Interview corpora of African American English and other regional varieties, such as CORAAL from the University of Oregon, are valuable for evaluation research, but many carry non-commercial research licenses, so read each license and component terms before any commercial training. Run any open corpus through a commercial-use license audit for open speech corpora.
  • Crowdsourced read speech. Common Voice is public domain and carries contributor-supplied metadata, but it is mostly read sentences from self-selected volunteers, so dialect balance and spontaneity are limited [4].
  • Commissioned collection. You choose speakers and prompts, and can recruit by self-identified dialect. Consent language must name AI training and evaluation; see consent language for commissioned collection.
  • Licensed operational call audio. Customer-service and sales calls from businesses serving many US regions contain natural dialect variety in real task conditions. Dialect labels are usually absent, so plan for consented self-identification on a subset or region-level proxies only. See our overview of contact-center call recordings.

Dialect speech raises consent questions twice: whether the original recording was lawful and whether dialect labels were collected with permission. A minority of US states require all-party consent to record; California Penal Code 632, for example, bars recording a confidential communication without the consent of all parties [5]. For operational call audio, confirm the disclosure each caller heard, the states involved and whether the recording notice covers use beyond quality assurance.

Dialect labels can correlate with race and ethnicity, so treat them as sensitive. Collect them only by self-report, store them separately from audio where possible, and keep group sizes large enough that a label plus metro does not single out a speaker. Voice is also an identifier in its own right; the guide to speaker anonymization for speech datasets explains what voice-conversion metrics can and cannot promise, and license terms for speech and voice recordings covers consent scope and voice-cloning limits.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

How SourceX approaches US dialect speech requests

SourceX sources operational datasets, including support and sales call histories, from US companies and manages the licensing process, including ongoing purchases. Data is sourced on request rather than held in stock, so a request for multi-region call audio does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, and every release is approved by the supplying company.

Before delivery, personal details such as names, phone numbers and account numbers are removed or replaced, the method is recorded and a sample is checked, though no method is perfect. Buyers describe the dialect coverage and conditions they need on the buyer request page, and SourceX looks for US businesses that hold that data. For broader context, start at the speech and audio data hub or the voice and audio licensing overview.

Source US dialect speech data for your ASR or voice agent

SourceX sources operational speech data, such as support and sales call recordings, from US companies on request, with rights review and supplier approval for every release. Serving AI teams wherever they are based, it manages the license through Find, Assess, Agree, Transact and Manage. Describe the US dialect speech data you need.

Frequently asked questions

Is synthetic TTS speech a substitute for real dialect recordings?

Not for dialect coverage. TTS voices reproduce what their training data contained, so they rarely capture dialect grammar, prosody and lexical choice faithfully. Use synthetic audio for vocabulary or noise augmentation, and real speakers for dialect; see real vs synthetic speech data.

Can we infer dialect from ZIP code instead of asking speakers?

Region is a weak proxy. Speakers move, communities within one metro speak different varieties, and fine-grained location raises re-identification risk. Use state and metro as context and self-identified dialect as the label.

Does text dialect data help speech models?

It helps language-model components and post-ASR understanding, but it does not teach acoustics. Pair audio with the guidance on dialect and regional variety text data.

Sources

  1. Sanabria et al., University of Edinburgh (arXiv:2303.18110), "The Edinburgh International Accents of English Corpus (EdAcc)" (2023). https://ar5iv.labs.arxiv.org/html/2303.18110
  2. DeepAI (paper summary), "Performance Disparities Between Accents in Automatic Speech Recognition". https://deepai.org/publication/performance-disparities-between-accents-in-automatic-speech-recognition
  3. arXiv:2506.00267, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
  4. Ardila et al., Mozilla (arXiv:1912.06670), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
  5. California Legislative Information, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  6. Jain et al., MLCommons (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  7. Longpre et al. (arXiv:2310.16787), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  8. Koenecke et al., "Racial disparities in automated speech recognition" (2020). https://5harad.com/papers/asr-disparities.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data