Skip to content

Speech and audio data

Code-Switched Speech Data: Spanish-English and Other Mixed-Language Conversations

Quick answer

Code-switched speech data is conversational audio where speakers move between languages within a sentence or across turns, transcribed with a consistent mixed-language convention and labeled with language segments as well as speaker turns. For Spanish-English ASR and voice agents, buyers should specify switch density, US Spanish varieties, channel (8 kHz telephony or wideband), a language-diarization reference format, and switch-aware metrics. The public corpora are small and carry restrictive terms, so commercial teams usually license real bilingual business conversations.

By SourceX Editorial · Updated

This guide sits under the speech and audio data hub. It covers training and fine-tuning data. For held-out test sets, see Spanish-English code-switching evaluation sets, and for per-language hour planning see multilingual ASR training data sourcing.

Why monolingual Spanish plus English data does not cover code-switching

Pooling separate Spanish and English corpora rarely teaches a model to handle a switch inside an utterance, because the model never sees the acoustic and lexical transitions at the switch point. A caller who says "necesito cambiar mi billing address porque me mudé" produces coarticulation across the boundary and a Spanish frame around English nouns. Language-ID front ends that pick one language per utterance then route the whole segment to the wrong decoder.

The failure modes are predictable. Embedded English product names get transliterated into Spanish spelling. Short Spanish discourse markers ("o sea", "pues", "este") inside English turns get deleted or misheard as English filler. Intent models downstream of ASR then miss the very terms (account, payment, plan names) that carry the request.

Public Spanish-English resources are thin and mostly academic. The handful of well-known conversational corpora are small, often recorded in a single city with lapel or room microphones rather than phone lines, and distributed under research or share-alike terms that need a license review before any commercial training use. Treat them as a sanity check for your pipeline, not as a training base for a production contact-center model.

What a usable code-switched record contains

A usable record pairs audio with two separate label layers: who is speaking, and which language each stretch of speech is in. The DISPLACE challenge treats speaker diarization and language diarization as distinct tasks with their own reference RTTM files [1], and buyers should demand the same separation. A single "language" field per call or per turn hides exactly the switches you are paying for.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "call_id": "c-000481",
  "audio": {"file": "c-000481.flac", "sample_rate_hz": 8000, "channels": 2, "channel_map": {"0": "agent", "1": "customer"}},
  "speakers": [
    {"id": "spk_cust", "role": "customer", "l1": "es", "variety": "es-MX", "self_reported_dominance": "balanced"},
    {"id": "spk_agent", "role": "agent", "l1": "en", "variety": "en-US", "spanish_proficiency": "professional"}
  ],
  "speaker_rttm": "c-000481.speaker.rttm",
  "language_rttm": "c-000481.lang.rttm",
  "segments": [
    {"start": 12.41, "end": 16.02, "speaker": "spk_cust",
     "text": "necesito cambiar mi billing address porque me mudé",
     "tokens_lang": ["es","es","es","en","en","es","es","es"],
     "switch_type": "intra-sentential"}
  ],
  "redaction": {"method": "tone+tag", "tag_format": "[PII:ADDRESS]"},
  "transcription_convention": "cs-conv-v1.2"
}

Key fields to require: per-token language tags (es, en, mixed, other, named entity), switch type (intra-sentential, inter-sentential, tag switch), a channel map for stereo call audio, and the version of the transcription convention used. Ask for a data statement covering language varieties and speaker demographics; the Data Statements V2 schema is a reasonable template [6].

Transcription conventions for mixed tokens

Write the transcription convention before any labeling starts, because mixed forms are where annotator disagreement concentrates. Academic corpora often use CHAT-style transcripts with word-level language tags, but a contact-center buyer needs rules tuned to its own vocabulary. Inconsistent spelling of the same mixed form across annotators inflates WER and teaches the model noise.

Decisions the convention must settle:

  • Loanwords and established borrowings. Is "email", "parquear" or "el carro" tagged English, Spanish or mixed? Pick a lexicon-based rule and publish the list.
  • Morphologically mixed forms. Hybrids such as "textear" or "chequear" need one canonical spelling and a "mixed" tag.
  • Named entities. Brand, product and place names get a separate tag so they do not count as switches.
  • Accents and diacritics. Require Spanish orthography with diacritics ("mudé", "teléfono"); decide whether scoring is diacritic-insensitive.
  • Disfluencies and fillers. Align with your verbatim policy; see verbatim vs clean transcription standards.
  • Numbers and spelled items. Account digits read in Spanish and confirmed in English must follow one normalization rule.

Normalization matters for scoring too. The Whisper authors built a text normalizer so harmless formatting differences would not count as errors [2]; for code-switched audio you need a normalizer that handles both languages' number words, contractions and diacritics consistently.

Measuring ASR on switched speech

Report error rates separately for monolingual segments, switched segments and the tokens around each switch point, because aggregate WER hides switch failures. A model can improve overall WER while getting worse exactly where callers change language, and that regression stays invisible if switched utterances are a small share of the test set. For mixed-script pairs such as Mandarin-English, teams combine word- and character-level scoring into a mixed error rate; for Spanish-English, word-level scoring split by language is usually enough.

Language diarization deserves its own score. Because DISPLACE scores language diarization against a separate reference from speaker diarization [1], you can measure whether the system finds switch boundaries at all, independently of whether it transcribes the words correctly. Ask suppliers to deliver language RTTM files that make this possible.

A practical evaluation set reports, at minimum: overall WER, WER on Spanish-only segments, WER on English-only segments, WER within switched utterances, and switch-point error at k = 1 or 2. Keep test calls disjoint from training calls by speaker and by agent, not only by recording. The code-switching evaluation sets page covers held-out design in more depth.

Speaker metadata for US Spanish varieties

Capture variety and language dominance per speaker, because US Spanish is not one accent. Mexican, Caribbean (Puerto Rican, Cuban, Dominican) and Central American speakers differ in phonology (for example, aspiration or deletion of syllable-final /s/ in Caribbean varieties) and in which English items they borrow. A corpus drawn from one metro area or one client program can over-represent one variety, so ask for the variety mix per hour of audio, not just per speaker.

Useful speaker fields: self-reported country of family origin, region of the US, generation (first or second), language dominance, and agent versus customer role. Bilingual agents also switch, often to mirror the customer, so agent speech is training signal and not only context. Pair this with accented English speech data if your users' English turns carry Spanish-influenced accents.

Channel, format and redaction choices

Match the training channel to deployment: bilingual contact-center traffic is mostly narrowband telephony, while voice agents in apps capture wideband audio. See telephony vs wideband ASR training for the trade-offs. Ask for stereo with agent and customer on separate channels where the source system recorded them that way; mono call mixes make overlap at switch points much harder to label.

Redaction interacts badly with code-switching. Account numbers and addresses are often where callers switch to English, so tones or silence inserted for PII removal can land precisely on switch points. Require a recorded redaction method and transcript tags such as [PII:ADDRESS], and read how audio redaction affects speech model training before setting the spec.

Confirm that recording and AI-use notices were actually given, and given in a language the caller understood. California Penal Code 632 requires consent of all parties to record a confidential communication [4], and a recording disclosure played only in English to a Spanish-dominant caller is a weak basis for a training license. Texas lists voiceprints as biometric identifiers requiring notice and consent before capture for a commercial purpose [5]; see biometric data rules for AI training.

Read the license, not the price tag. A call-center audio listing on a data marketplace [3] still needs its terms checked for commercial model training, even when it is offered at no charge. Academic corpora are often shared under research-only or share-alike terms, and some hosting listings add explicit restrictions on model training. Our license audit for open speech corpora walks through the clauses to check.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Buyer checklist for a code-switched audio request

Illustrative example: invented to show structure; it does not describe an available dataset.

RequirementWhat to specifyWhy it matters
Switch profileTarget share of utterances with a switch; intra- vs inter-sentential mixLow switch density gives mostly monolingual audio
Label layersSpeaker RTTM and language RTTM, per-token language tagsLanguage diarization needs its own reference [1]
ConventionVersioned guide for loanwords, hybrids, entities, diacriticsPrevents annotator drift on mixed forms
VarietiesSpeaker origin, region, dominance, roleAvoids single-variety skew
Channel8 kHz telephony or wideband; stereo channel mapMatches deployment acoustics
MetricsPer-language WER plus switch-point errorAggregate WER hides switch failures
RedactionMethod, tag format, sample checkPII often sits at switch points
RightsRecording consent, notice language, commercial training licenseFree or academic terms may bar training [3]

For a fuller template, adapt the contact-center audio requirements spec. SourceX's call center audio datasets, voice agent training data and multilingual enterprise model data pages cover adjacent requests, and do AI labs buy non-English business data? covers demand context.

Where SourceX fits in sourcing bilingual conversational audio

SourceX sources operational datasets, including support and sales histories, from US companies on request; nothing is held in stock and a request does not guarantee a match. Buyers describe the data they need, such as Spanish-English support calls with language labels, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. If you serve bilingual US customers, describe your code-switched audio requirements.

Request code-switched conversational audio

Every SourceX dataset is rights-reviewed for ownership and consents, has personal details removed or replaced before delivery with the method recorded, and is delivered under a license defining records, uses, term and delivery. Nothing is contracted until a supplier agrees, and pricing is set per deal. Tell SourceX what bilingual call audio you need.

Sources

  1. arXiv, "Summary of the DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments" (2023). https://arxiv.org/pdf/2311.12564
  2. arXiv (OpenAI), "Robust Speech Recognition via Large-Scale Weak Supervision" (2022). https://arxiv.org/pdf/2212.04356
  3. AWS Marketplace, "Call-center ASR audio dataset listing". https://aws.amazon.com/marketplace/pp/prodview-yyfwirpya2mp6
  4. California Legislative Information, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  5. Texas Legislature, "Texas Business and Commerce Code Section 503.001". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  6. University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data