Skip to content

Data quality, coverage and contamination

Language Identification QA for Multilingual Datasets

Quick answer

Language identification (LID) QA checks that every record in a multilingual dataset is actually in the language its label claims, and that the delivered language mix matches the spec. Do not trust source metadata alone: rerun a modern LID model such as GlotLID at the right granularity, set per-language confidence thresholds, measure false positives that flood small subsets, and confirm each language with a stratified sample read by native reviewers before training or evaluation.

By SourceX Editorial · Updated

This page sits in the data quality, coverage and contamination hub and focuses on one check: whether language tags are true. For the choice between native text, parallel corpora and translation memories, see multilingual training data types compared.

Why language labels in multilingual training data are often wrong

Language labels are often wrong because they come from where the data was found, not from what the text says. A record tagged sw because it came from a Kenyan support queue may be English, Sheng, or code-switched Swahili-English. GlotLID's authors list wrong source metadata as a distinct failure mode alongside model errors [1].

The problem is worst where it matters most. Low-resource languages have the least clean reference text, so both their source labels and the models that check them are weakest, and GlotLID was built specifically because existing LID tools mislabel many of these languages [1].

Labels in metadata are also what downstream tooling filters on. On the Hugging Face Hub, the language field in a dataset card's YAML block drives search and filtering [3], so a wrong tag propagates into every mix you build from it. Treat the declared label as a claim to test, not a fact.

Common LID failure modes to test for

LID errors cluster into a small number of predictable patterns, and each needs its own test. The GlotLID paper names the main ones: wrong metadata, leakage from high-resource languages, confusion between close relatives, macrolanguage ambiguity, and noise [1]. In operational data such as tickets, chat logs and transcripts, a few more show up.

  • High-resource leakage. English, Spanish, French or Russian text gets predicted as a low-resource language, or low-resource text with English loanwords gets pulled into English. The GlotLID README notes higher false-positive rates for high-resource languages [2].
  • Close relatives. Bosnian/Croatian/Serbian, Malay/Indonesian, Hindi/Urdu in romanized form, Galician/Portuguese, Norwegian Bokmål/Nynorsk/Danish. Expect confusion matrices with heavy off-diagonal mass.
  • Macrolanguages and code granularity. zh versus cmn/yue, ar versus arb/arz/ary, ms versus zsm. A dataset may be correctly "Arabic" and still useless for an Egyptian-dialect assistant [1].
  • Script mismatch. Serbian in Latin versus Cyrillic, Hindi in Devanagari versus Latin (Hinglish), Uzbek in Latin versus Cyrillic. Your label scheme should carry script, as in BCP 47 sr-Latn or GlotLID-style srp_Latn.
  • Short and noisy segments. "ok", "thanks", order IDs, URLs, stack traces and templated signatures carry little linguistic signal; the GlotLID README warns against relying on predictions for very short sentences [2].
  • Code-switching. One chat turn mixes two languages; a document-level label hides it and a sentence-level label fragments it.
  • Boilerplate and repetition. A single repeated disclaimer or footer can inflate a language's apparent share; deduplication research found one sentence repeated over 60,000 times in C4 [5].
  • Machine-translated text. Fluent but translated content passes LID and still distorts the mix; see detecting model-generated content in purchased data.

GlotLID vs fastText language ID: choosing a model

Choose an LID model by its label set and its coverage of the languages you actually care about, not by overall accuracy. The original fastText lid.176 model is fast and widely used, but its label set is small and ISO 639-1-style, so many low-resource languages simply have no label and are forced into a near neighbor. GlotLID-M is also built on fastText, keeping its throughput, but the published paper version covered 1,665 languages with ISO 639-3 plus script labels [1]; later releases expanded the label set, so check the version you run [2].

That design difference matters for QA. A model with no label for Twi or Fula will confidently call it something else, and that error looks like clean data. A model with a broad label set can also produce more confusable labels among close relatives, which is why the threshold and audit steps below still apply.

Practical guidance:

  • Run at least two models and log both predictions. Disagreement is a cheap signal for which records to send to human review.
  • Normalize all labels to one scheme (ISO 639-3 plus ISO 15924 script) before comparing models or computing the mix.
  • Pin the model version and record it with the QA report, because label sets and calibration change between releases [2].
  • For speech datasets, run LID on the transcript text and separately check audio-language metadata; see multilingual ASR training data.

Setting a language detection confidence threshold per language

Use per-language confidence thresholds, not one global cutoff. LID confidence is not comparable across languages: the GlotLID README notes higher false-positive rates for high-resource languages [2], and in practice genuine low-resource text often scores lower than confident but wrong high-resource predictions. A single 0.5 or 0.9 cutoff therefore keeps too much leaked English and drops too much real Yoruba.

Set each threshold from a labeled validation slice for that language. Pick the threshold that hits a target precision (for example, 95% of retained records truly in-language) and record the recall you give up. Where you lack validation data, start conservative for languages known to attract false positives and tighten after the native audit.

Add rules the score cannot express. Require a minimum segment length (in characters or tokens) before trusting a prediction, route records below it to a "short/undetermined" bucket (und in ISO 639), and apply script checks that reject a srp_Cyrl label on Latin-only text.

How a small false-positive rate swamps a low-resource subset

A false-positive rate that looks negligible on the whole corpus can make a low-resource subset mostly wrong. Suppose a 10-million-record corpus is 98% English and you expect 20,000 records of a low-resource language. If the LID model mislabels just 0.5% of English records as that language, it adds 49,000 false records, and the "low-resource subset" is now about 70% English.

This is the base-rate problem, and it is why GlotLID's authors stress high-resource leakage [1] and the README flags false positives for high-resource languages [2]. Precision for a small language depends on the error rate of every large language feeding into it. Always compute per-language precision on the predicted set, not accuracy on the whole corpus.

The fix is a combination of per-language thresholds, a second model as a tiebreaker, and native review of the predicted subset itself. Upstream text filtering also changes the balance; see quality filtering for pretraining-scale text.

How to verify the language mix of a dataset: a QA workflow

Verify the language mix by measuring it yourself and reconciling it with the supplier's declared numbers. The workflow below works for delivered samples and full deliveries.

  1. Normalize and segment. Strip HTML, signatures, quoted email history and system templates; decide document-, turn- or sentence-level granularity per use case.
  2. Deduplicate first. Remove exact and near-duplicates before measuring shares, so boilerplate does not inflate a language [5]; see near-duplicate detection with MinHash and LSH.
  3. Run two LID models. Store top-3 labels and scores per record with model versions.
  4. Apply per-language thresholds and length rules. Bucket records as accepted, undetermined, or conflicting.
  5. Compute the measured mix. Report records, tokens and characters per language and script; token share is what matters for training budgets.
  6. Reconcile with metadata. Build a declared-versus-predicted confusion matrix; any cell above your tolerance becomes a finding.
  7. Native-reviewer audit. Draw a stratified sample per language (including the undetermined and conflicting buckets) and have native speakers judge language, variety, script and usability.
  8. Accept or reject per language. Use an explicit sampling plan; see acceptance sampling for dataset deliveries.

The native-reviewer step is a recommendation from practice rather than a published standard: reviewers catch dialect, register and machine-translation issues that LID cannot. Many gross errors, such as wrong script, empty text or obvious English, are visible without fluency, so a quick non-native screen before paid native review saves time.

Per-language LID QA report: an artifact to request or produce

A per-language QA report turns LID findings into a document both sides can sign off on. Ask suppliers for one alongside the sample, or produce it yourself and share it back. Pair it with a data statement describing language variety, speaker population and collection context, as proposed by Bender and Friedman [4].

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample value
dataset_idsupport-chat-sample-v2
lid_modelsglotlid-m (pinned version); fasttext lid.176
label_schemeISO 639-3 + ISO 15924 script
granularityconversation turn
min_length_rule>= 20 characters, else und
languageswh_Latn
declared_records18,400
predicted_records_accepted12,950
threshold_used0.82 (set for 95% precision on validation slice)
top_confusionseng_Latn 3,100; mixed swh/eng 1,700; und 650
model_disagreement_rate9.4%
native_audit_sample400 turns, stratified by queue and month
native_audit_in_language93%
notescode-switching common in billing queue; Sheng flagged as variety
decisionaccept with filter; re-tag mixed turns

Report one block per language and a summary table of declared versus measured shares. A finding like "38% of declared Swahili predicted as English or mixed" is a commercial issue as much as a technical one, because it changes what you are paying for.

What to ask suppliers about language labels before you buy

Ask how every language label was produced before you rely on it. Useful questions:

  • Was the label assigned from source metadata (market, locale setting, queue), by an LID model, or by humans? Which model and version?
  • What code scheme is used, and does it carry script and variety (for example pt-BR versus pt-PT, es-419)?
  • How are short, empty, templated and code-switched segments labeled?
  • Was deduplication done before computing the declared language shares [5]?
  • Can you provide a per-language sample large enough to audit the smallest language you are buying?

Operational data from companies tends to carry language signals tied to customer locale or queue routing, which is useful context but not proof. For how AI teams approach non-English business data generally, see do AI labs buy non-English business data and training data for multilingual enterprise models. If your target is retrieval rather than generation, also read multilingual and cross-language retrieval data.

SourceX sources operational datasets such as support and sales histories from US companies on request, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset; you can describe the language mix you need to SourceX and apply the checks on this page during your own assessment.

Sourcing multilingual operational data with verified language labels

SourceX looks for US businesses that hold the data you describe, including support conversations, sales histories and documents, and manages assessment, licensing and ongoing purchases. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and data is sourced on request, so a request does not guarantee a match. Tell SourceX what multilingual data you need.

Sources

  1. Kargaran et al., arXiv (EMNLP 2023 Findings), "GlotLID: Language Identification for Low-Resource Languages" (2023). https://arxiv.org/pdf/2310.16248
  2. GlotLID project, GitHub, "GlotLID language identification model". https://github.com/AI-Natural-Language-Processing-Lab/GlotLID-language-identification-model
  3. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  4. Bender and Friedman, TACL vol. 6, "Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science" (2018). https://aclanthology.org/Q18-1041/
  5. Lee et al., arXiv (ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data