Text and language data
Estimating Token Counts for a Text Dataset Before Licensing
Quick answer
To estimate a text dataset's token count, do not multiply a supplier's word or gigabyte figure by a rule of thumb. Instead, have a stratified sample run through your own tokenizer, measure tokens per byte (or per word) for each language and document type, apply deductions for markup, duplicates and filtering, and extrapolate with a confidence interval. Then write the counting basis into the order form so that the delivered count can be audited against the quote.
By SourceX Editorial · Updated
Token count drives budgets, per-token pricing and data-mix planning, but it is not a property of the text alone, because tokens are defined by a vocabulary. It depends on the tokenizer, the language, the format and the cleaning applied. The glossary entry on token count defines the term, and this page covers how to estimate it for a corpus you have not yet licensed. For the wider sourcing context, start at the text and language data hub.
Why supplier size figures cannot be compared directly
Supplier quotes are not comparable because each one uses a different unit measured at a different stage of the pipeline. One multilingual vendor's site, as of October 2026, quotes its collection as 7.4 billion words across 483 language pairs [1]. An SEC filings release is quoted as 43 billion "clean tokens" [2]. Open corpora are described as 8 TB of text [6] or as 154 billion tokens [7].
None of these figures converts to another without assumptions. Words depend on segmentation rules. Bytes depend on encoding and markup. Tokens depend on a vocabulary you may not share, and "clean" depends on what was removed.
The units you will usually be quoted, and what each one hides:
| Quoted unit | What it actually measures | Main hidden variable |
|---|---|---|
| Bytes / GB / TB | Storage on disk, often compressed or in a container format | Compression, HTML/JSON markup, encoding (UTF-8 vs UTF-16) |
| Characters | Unicode code points after decoding | Whitespace, markup, normalization (NFC vs NFKC) |
| Words | Whitespace or tool-based segments | Segmentation for Chinese, Japanese, Thai; hyphenation; numbers |
| Documents / records | Files, rows, tickets, filings | Length distribution, which is usually long-tailed |
| Raw tokens | Output of the supplier's tokenizer | Vocabulary, special tokens, fields included |
| Clean tokens | Tokens after dedup and filtering | Which filters, which dedup scope, what was dropped [2][5] |
How tokenizer and language change the count
The same text can produce very different token counts under different tokenizers, and the gap widens outside English. Fertility, the average number of tokens per word, is a common measure of this. Research on adapting pretrained models to new languages found that replacing about 10% of a vocabulary improved fertility by 42% for Hungarian and 73% for Thai [3]. A corpus quoted in Hungarian or Thai words can therefore map to very different token totals depending on which vocabulary you train with.
Practical consequences for buyers:
- Your tokenizer, not theirs. A supplier's token count from a GPT-2-style BPE, a SentencePiece unigram model or a Llama-family tokenizer does not transfer to yours. Name the exact tokenizer file or library version (for example a
tokenizer.jsonhash or a tiktoken encoding name). - Script matters. CJK text has no whitespace word boundaries, so "words" are a segmenter's output. For these languages, characters or bytes are a more stable base than words.
- Byte fallback. Byte-level tokenizers never produce unknown tokens, but rare scripts can fall back to one token per byte, which multiplies counts.
- Planned vocabulary extension. If you plan to extend your tokenizer, estimate with both the current and the extended vocabulary. The tokenizer fertility and coverage guide covers how to evaluate that.
For multilingual deals, compute fertility per language rather than as a single blended figure. The non-English corpus licensing guide explains why language mix also affects rights review.
The estimation method: stratified sample, measured ratios, extrapolation
The most reliable pre-license estimate comes from tokenizing a stratified sample with your tokenizer and extrapolating ratios stratum by stratum. Treat this as a ratio estimation problem. You know the population total of some cheap unit (bytes, characters or words) and you measure tokens per unit on a sample.
- Get a manifest, not a headline. Ask for per-stratum totals of bytes (uncompressed, decoded), characters, words and document counts. Strata are language × source type × format, for example English support tickets in JSON, German contracts in PDF-extracted text, or Japanese chat logs in CSV.
- Draw the sample inside each stratum. Sample documents at random, but make sure long documents are represented, because token-per-byte ratios drift with length and boilerplate share.
- Tokenize with the target pipeline. Apply your own extraction, normalization (for example NFC), field selection and tokenizer. Count with and without special tokens such as BOS/EOS so that both bases are known.
- Compute the ratio and its uncertainty. For each stratum, use the ratio estimator R = Σtokens / Σbytes over sampled documents. Bootstrap over documents (not over tokens) to get a 90% or 95% interval.
- Extrapolate and sum. Multiply each stratum's R by its population byte total, then sum. Combine stratum variances to get an interval for the corpus total.
- Apply deductions separately. Estimate dedup and filtering losses as their own ratios (next section) instead of folding them into R.
For how many documents to check per stratum, the sample-size guide for dataset error rates applies the same logic. Strata with high variance, such as mixed HTML pages, need more documents than uniform ones like short chat turns. If the supplier cannot run your tokenizer, an alternative is to ask for a small redacted sample under NDA and run it yourself.
Deductions: markup, boilerplate and duplicates
Raw counts overstate usable tokens because markup, boilerplate and repeated text all tokenize. A "clean token" figure is only meaningful when you know which steps produced it [2]. Large public pipelines report their sizes after multi-stage extraction, filtering and deduplication, so their token counts are post-pipeline figures by construction [5].
Where the inflation usually comes from:
- Markup and structure. HTML tags, JSON keys repeated on every record, CSV delimiters, XBRL or XML wrappers. Structured filings such as 10-K reports split into items in JSON [8] tokenize differently depending on whether keys are counted.
- Boilerplate. Email signatures, legal disclaimers, ticket templates, navigation text, page headers and footers in PDF extractions.
- Duplicates. Exact and near-duplicate content is common. In C4, one sentence was repeated over 60,000 times [4]. Operational corpora show the same pattern: quoted replies in email threads, re-opened tickets and versioned documents.
- Redaction placeholders. Replacing names or account numbers with tags such as
[NAME]changes the token count, slightly up or down depending on the tokenizer. - Filtering losses. Language ID thresholds, length filters, quality classifiers and PII filters remove documents. Ask for removal rates per filter.
Measure each deduction on your sample. For example, run MinHash near-dup detection within the sample and against the supplier's stated dedup scope, strip markup to your extraction standard, and record the remaining fraction. Keep exact and near-duplicate removal as separate lines, because a supplier may have done one but not the other.
Worked example: normalizing two offers
Normalizing two offers to your tokenizer and cleaning standard often reverses which one looks larger. In the invented example below, offer A is quoted in words and offer B in gigabytes. Every ratio is an assumption chosen to show the arithmetic, not a reference value.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Step | Offer A (quoted: 2.0B words, English + Thai) | Offer B (quoted: 30 GB JSON, English) |
|---|---|---|
| Strata | English 1.6B words; Thai 0.4B segmented words | English support tickets, 30 GB uncompressed JSON |
| Measured ratio (your tokenizer) | English 1.3 tokens/word; Thai 2.6 tokens/word | 0.21 tokens/byte including JSON keys |
| Raw tokens | 2.08B + 1.04B = 3.12B | 6.3B |
| Markup and field stripping | none needed (plain text): 3.12B | keep body and subject fields only, 55% retained: 3.47B |
| Near-dup and quoted-reply removal | 8% removed: 2.87B | 30% removed (quoted replies): 2.43B |
| Filtering (language ID, length) | 4% removed: 2.76B | 3% removed: 2.35B |
| Usable tokens (90% interval) | about 2.76B (2.6B to 2.9B) | about 2.35B (2.1B to 2.6B) |
Offer B looked twice as large by raw tokens. After stripping and dedup it is smaller than A. If you were to extend your tokenizer for Thai, A's Thai share would shrink further [3], so rerun the estimate with the vocabulary you will actually train on. For valuing the result once it is counted, see estimating a dataset's value before purchase.
Writing the counting basis into the order form
A token count is only enforceable if the order form defines how it is counted. Without a definition, a per-token price or a volume commitment cannot be audited, and disputes become arguments about tokenizers. The per-token pricing guide covers audit mechanics and price comparison. The block below shows the fields to pin down.
Illustrative example: invented to show structure; it does not describe an available dataset.
counting_basis:
tokenizer: "buyer-supplied tokenizer.json, sha256 recorded in schedule"
special_tokens: excluded # BOS/EOS/padding not counted
normalization: "Unicode NFC, no lowercasing"
fields_counted: [subject, body] # JSON keys and metadata excluded
markup: "HTML stripped to visible text per agreed extractor"
dedup:
exact: "document-level SHA-256, across full delivery"
near: "MinHash, Jaccard >= 0.8, across full delivery"
filters_applied_before_count: [language_id, min_length_50_chars]
redaction: "placeholders counted as tokenized"
measured_at: "after dedup and filters, before buyer-side processing"
per_stratum_report: [language, source_type, documents, bytes, tokens]
tolerance: "delivered count within agreed band of quoted estimate"
verification: "buyer recount on full delivery or agreed random sample"
Two further habits help. Record the size figure on the dataset card or manifest with its unit and stage, much as Hugging Face dataset cards record size alongside license and language in YAML metadata [9]. Also, keep counts for each delivery of a recurring feed, so that growth and duplication across deliveries are visible. If you are sourcing operational text from companies rather than open corpora, a SourceX buyer request starts from a description of the data you need, not of the businesses that hold it.
Common estimation failures to check for
Most bad estimates come from a few predictable errors, and each one can be checked on the sample.
- Compressed bytes treated as text bytes. A 30 GB
.jsonl.gzmay decompress to several times that size. Always ask for uncompressed, decoded byte counts. - One blended ratio for a multilingual corpus. A small high-fertility language share can dominate the error.
- Sampling only short documents. Easy-to-share samples skew short and clean, while long PDFs and logs carry more boilerplate.
- Counting metadata as content. Ticket IDs, timestamps and repeated keys inflate counts but rarely add training value.
- Dedup scope mismatch. A supplier dedups within each monthly file, while you dedup across the whole corpus and against data you already hold.
- Ignoring overlap with your existing mix. Text you already have from public crawls adds fewer new tokens than its count suggests. The proprietary text beyond web crawls guide discusses which sources tend to be new.
Sizing a licensed text corpus with SourceX
SourceX sources operational text, such as support and sales histories, engineering records and documents, from US companies on request, so there is no stock catalog to quote from, and a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery, so bring your counting basis to that negotiation. To describe the corpus and token volume you need, start a buyer request at SourceX.
Frequently asked questions
Is there a reliable words-to-tokens ratio?
Only per tokenizer, language and domain. Ratios measured on English prose do not hold for code, tables, chat logs or non-Latin scripts, and vocabulary changes can shift fertility substantially [3]. Measure the ratio on a sample instead of borrowing a published constant.
Should I estimate from bytes or from words?
Bytes (uncompressed, decoded UTF-8) are usually the more stable base, because they need no segmentation rules and suppliers can report them exactly. Use words only where the supplier's segmentation is defined and the languages are whitespace-delimited.
How many tokens do I need?
That depends on the model and the purpose: pre-training, continued pre-training or fine-tuning. The question on how many records AI labs want and the guide to planning licensed pre-training volume cover volume planning. This page covers only how to measure what an offer contains.
Sources
- TAUS, "Data for AI". https://www.taus.net/data-solutions/data-for-ai
- Daft, "SEC EDGAR case study". https://www.daft.ai/blog/sec-edgar-case-study
- arXiv, "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
- arXiv (Lee et al., ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- arXiv, "The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale" (2024). https://arxiv.org/pdf/2406.17557
- arXiv, "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
- arXiv, "The German Commons: 154 Billion Tokens of Openly Licensed Text for German Language Models" (2025). https://arxiv.org/html/2510.13996v1
- arXiv, "EDGAR-CORPUS: Billions of Tokens Make The World Go Round" (2021). https://arxiv.org/pdf/2109.14394
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.