Skip to content

Text and language data

Tokenizer Fertility and Coverage: Evaluating a Text Corpus for Vocabulary Extension

Quick answer

Tokenizer fertility is the average number of tokens a tokenizer produces per word, and it is the first number to compute on any corpus you plan to use for vocabulary extension. Run your current tokenizer over a stratified sample, report fertility, characters per token, byte-fallback rate and continued-word share per language and script, then compare against a reference language. High fertility shows where extension pays off; the corpus must also be large, clean and licensed enough to train the new embeddings afterward.

By SourceX Editorial · Updated

Why fertility decides whether vocabulary extension is worth buying data for

Fertility matters because every extra token per word raises training and inference cost, adds latency and shrinks effective context for that language. The gap is widest for scripts the base vocabulary barely saw during training. If your base tokenizer splits a Thai or Amharic sentence into several times as many tokens as its English equivalent, a 32k context window holds a fraction of the content.

Extension is the usual fix, and the evidence suggests a modest budget goes a long way. Replacing roughly 5,000 of the least-frequent tokens, about 10% of the vocabulary, improved fertility by 42% for Hungarian and 73% for Thai, with diminishing returns beyond that budget [1]. That result frames the purchase: you are not buying text to rebuild a tokenizer from scratch, you are buying enough representative text to pick the right new merges and then teach the model what they mean.

Fertility is a symptom, not a goal. A corpus that lowers fertility by being dominated by one template (boilerplate legal footers, repeated ticket signatures) produces merges that look efficient on the sample and fail on real traffic. For the token concept itself, see the tokens glossary entry.

Metrics to compute on a candidate corpus before signing

Compute five numbers per language and script slice, using the exact tokenizer you intend to extend. Report them on a held-out sample the supplier did not choose, stratified by source system and document type.

  • Fertility (tokens per word). Tokens divided by whitespace-delimited words. It is meaningful only where words are delimited by spaces.
  • Characters (or UTF-8 bytes) per token. The script-neutral alternative for Thai, Lao, Khmer, Chinese and Japanese, where whitespace does not mark words and per-word fertility is ill-defined.
  • Continued-word share. The fraction of words split into two or more tokens. A high share with moderate fertility points to systematic splitting of common morphemes.
  • Byte-fallback and unknown rate. With SentencePiece byte_fallback=true or byte-level BPE, unseen characters become raw byte tokens rather than <unk>. A high byte-token share means the base vocabulary has no coverage of that script.
  • Parity ratio against a reference. Tokens for a passage in the target language divided by tokens for the same content in English. Measure on parallel text so content is held constant across languages.

Two practical cautions apply. First, count words after Unicode normalization, or decomposed and precomposed forms of the same character will tokenize differently and inflate variance. Second, slice by domain: support transcripts, contracts and engineering logs tokenize very differently, and jargon-heavy slices are covered on the domain vocabulary coverage page.

Script and normalization coverage for non-Latin text

Coverage of non-Latin scripts fails most often because of encoding and normalization defects in the corpus, not because the script is absent. Before measuring, require suppliers to deliver UTF-8 text in Unicode NFC, with consistent whitespace, and with a per-document script or language tag (ISO 15924 script codes and BCP 47 language tags are the common conventions). These are reasonable delivery requirements to put in a data request; they are not a guarantee that any given supplier already meets them.

Check for these failure modes in the sample:

  • Mixed normalization. NFC and NFD forms of Vietnamese diacritics or Devanagari conjuncts mixed in one corpus. SentencePiece applies nmt_nfkc by default, but if you train merges on unnormalized data and serve normalized text, the learned merges miss.
  • Zero-width characters. ZWJ and ZWNJ are meaningful in Persian, Hindi and Malayalam text and should be preserved; stray zero-width spaces from web editors should not.
  • Transliteration and code-switching. Romanized Hindi or Arabizi inside "Hindi" or "Arabic" tagged documents lowers measured fertility while teaching the model nothing about the native script.
  • Rare-character tail. SentencePiece character_coverage decides how much of the character tail becomes vocabulary versus byte fallback. For scripts with large character inventories, inspect the tail rather than accepting a default.

If your coverage gap is regional rather than script-level, the dialect and regional variety page covers how to specify varieties in a request.

How much text you need: tokenizer training versus model adaptation

Training the tokenizer itself needs far less text than adapting the model to use it, so size the purchase for the second step. Merge statistics stabilize on a representative sample long before the model has seen enough text to learn useful embeddings for thousands of new tokens. Treat the tokenizer sample as a subset of the adaptation corpus, not a separate buy.

The adaptation step is where volume matters. The usual recipe is two stages: extend the vocabulary, then run continued pre-training and supervised fine-tuning so the model learns the new tokens. One 2026 study added 30,000 Ge'ez-script subwords to a multilingual model, then ran continued MLM and SFT, and the authors report large QA gains [2]. How the new embedding rows are initialized also matters; a 2026 systematic study compares initialization strategies for exactly this step [3].

Expansion is not free. The new embedding rows start untrained, so quality on tasks such as translation can drop until the model has seen enough target-language text, which is why the adaptation corpus, not the merge sample, sets the purchase size. Continual pre-training research shows that learning-rate re-warming and replay of earlier data help a model absorb a new distribution without retraining from scratch [5], which means you need replay data in the original languages as well as new target-language text. For sizing in tokens and how that drives pricing, use the token count glossary entry and the continued pre-training corpus guide.

A worked evaluation report for a candidate corpus

A useful corpus evaluation is a short, reproducible report you can attach to a data request or diligence file. The schema below shows the fields worth recording for each slice.

Illustrative example: invented to show structure; it does not describe an available dataset.

corpus_eval:
  candidate_id: "cand-thai-support-v1"
  tokenizer: "base-model-tokenizer (byte-level BPE, 32k vocab)"
  sample:
    method: "stratified by source_system and doc_type; buyer-drawn seed"
    documents: 5000
    normalization_checked: "NFC; 0.4% of docs failed and were re-normalized"
  slices:
    - lang: "th"          # BCP 47
      script: "Thai"      # ISO 15924
      chars_per_token: 1.3
      bytes_per_token: 3.6
      byte_fallback_share: 0.11
      parity_vs_en: 3.9   # parallel subset, same content
    - lang: "en"
      script: "Latn"
      fertility_tokens_per_word: 1.3
      continued_word_share: 0.14
  contamination_checks:
    near_duplicate_rate: 0.07       # MinHash, Jaccard >= 0.8
    template_share: 0.18            # repeated signatures/footers
    code_switch_share: 0.05         # romanized text inside 'th'
  extension_trial:
    new_tokens: 3200                # about 10% of vocab replaced
    parity_vs_en_after: 1.8
    held_out_slice: "contracts (not used for merges)"
  license_check:
    tokenization_permitted: "confirm before running this report"
    derivative_restrictions: "none found in draft license"

Read the report as a decision aid. If parity drops sharply after a trial extension on a held-out slice, the corpus supports extension; if it drops only on the slice used to train merges, the corpus is too narrow.

Decision table: extend, replace, or keep the tokenizer

The right choice depends on parity, byte-fallback share and how much clean target text you can license. Use this as a starting point, then validate on your own model.

Illustrative example: invented to show structure; it does not describe an available dataset.

Observation on candidate corpusLikely decisionData you need to buy
Parity under about 1.5 and byte fallback near zeroKeep the tokenizer; spend on contentDomain text for continued pre-training
Parity high, script partly coveredExtend or replace the least-frequent tokensRepresentative target text plus replay data
Byte fallback dominant, script absentAdd a script-specific block of tokensLarge native-script corpus, clean NFC
Low fertility gain on held-out slicesCorpus too narrow; do not extend yetMore source systems and document types
Translation is a core use caseExtend cautiously and test translationLarge target corpus plus parallel text

License terms that can block tokenizer testing

Check the license before you run a single fertility measurement, because some terms restrict tokenization itself. A NoDerivs-style license may prohibit tokenizing the corpus, since tokenization and the resulting artifacts can be treated as derivatives [4]. Dataset metadata is also unreliable: the Data Provenance Initiative audit of 1,800+ text datasets reported license omission above 70% and license error rates above 50% on popular hosting sites [6].

Ask the supplier, in writing, whether evaluation runs, merge statistics and the extended tokenizer file itself are permitted uses, and whether those artifacts survive the end of the license. The pre-training rights guide and non-English corpus licensing page cover the grant language in more depth. For low-resource sources specifically, see low-resource language text corpora.

Where operational text fits in a tokenizer-extension plan

Operational text from businesses, such as support histories, sales conversations, engineering records and documents, fills a gap that web crawls leave: real domain vocabulary, abbreviations and product names in the form your users write them. SourceX sources these kinds of operational datasets from US companies on request; they are not held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and delivery follows an executed license that defines records, uses, term and delivery. If your evaluation shows a domain slice with poor coverage, you can describe the text you need to SourceX and the request goes through Find, Assess, Agree, Transact and Manage, with nothing contracted until a supplier agrees.

For broader context, the text datasets hub covers corpus types and rights, and the training data quality hub covers deduplication and contamination checks that should run alongside fertility measurement.

Sourcing a corpus for tokenizer fertility work

If your fertility and coverage report points to a domain or language slice your current data does not cover, describe the records, languages and allowed uses you need. SourceX looks for US businesses that hold the described data, every release is approved by the supplying company, and terms are agreed per deal. Start a buyer request at sourcex.si/buyers.

Sources

  1. arXiv, "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
  2. arXiv, "Expanding the Lexicon of Ge'ez Based African Languages (arXiv:2607.15209)" (2026). https://arxiv.org/abs/2607.15209
  3. arXiv, "Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension" (2026). https://arxiv.org/pdf/2608.03494
  4. arXiv, "Low-resource language corpus licensing for AI training (arXiv:2606.28867)" (2026). https://arxiv.org/abs/2606.28867
  5. arXiv (Ibrahim et al.), "Simple and Scalable Strategies to Continually Pre-train Large Language Models" (2024). https://arxiv.org/pdf/2403.08763
  6. Nature Machine Intelligence (Longpre et al.), "A large-scale audit of dataset licensing and attribution in AI" (2024). https://www.nature.com/articles/s42256-024-00878-8

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data