Skip to content

Data licensing for AI training

Per-token pricing for training data: how to count, audit and compare

Quick answer

A per-token price for licensed training text is only meaningful once the contract fixes four things: the tokenizer and version that does the counting, whether the count is taken before or after deduplication, filtering and PII removal, how the buyer verifies it, and what tolerance triggers an adjustment. Without those, two quotes at the same rate can differ in real cost by a large factor. Pin the counting basis in a schedule, reproduce the count on delivery, and normalize every offer to cost per usable deduplicated token before comparing.

By SourceX Editorial · Updated

Why the same corpus produces different token counts

The token count of a corpus is a property of the corpus and the tokenizer together, so an unqualified number such as "40 billion tokens" is not a unit of sale. A byte-level BPE tokenizer with a 32k vocabulary, a 100k-vocabulary tiktoken encoding and a SentencePiece unigram model will each segment the same text differently. The gap widens for code, tables, non-English text and OCR output, where tokenizers trained mostly on English web text fragment words into more pieces.

Research on adapting models to new languages measures this as fertility, the average number of tokens produced per word, and finds it is much higher where the tokenizer covers a language poorly [3]. For a buyer, high fertility means a seller counting with a poorly matched tokenizer reports more tokens, and a per-token price rewards the seller for that mismatch. Openly licensed corpora publish headline counts, such as 154 billion tokens for the German Commons [7], and those figures also depend on the tokenizer the authors used.

Treat the tokenizer as a contract term. See the glossary entries on tokens and token count for definitions, and the guide to estimating token counts before licensing for pre-offer sizing.

Data price per token is not compute price per token

Most published "price per training token" figures are compute billing for fine-tuning, not the price of licensed data. Fine-tuning providers bill per token processed, typically counting training plus validation tokens and varying by model size and method [4]. That count also multiplies with epochs: three passes over a one-billion-token file bill roughly three billion tokens of compute.

A data license price is a one-time or recurring fee for rights to the content, and it should never scale with how many epochs you run. If a seller's draft says "tokens consumed in training" or "tokens processed", strike it and replace it with "unique tokens delivered, counted as defined in Schedule A". Developer forums show how quickly billed counts diverge from a buyer's own estimate once epochs and tokenizer differences enter [5].

Choosing the counting basis: raw, deduplicated, filtered or delivered

The fairest basis for pre-training buyers is usually tokens remaining after exact and near-duplicate removal, quality filtering and PII handling, because those are the tokens that actually train the model. Raw counts overstate value: Lee et al. found long repeated substrings throughout common language-modeling datasets, including one sentence repeated more than 60,000 times in C4, and showed that deduplication reduces verbatim memorization [1].

Each stage removes tokens, and each needs a definition:

  • Raw: everything in the export, including boilerplate, signatures, quoted email chains, templated tickets and log lines.
  • Exact-deduplicated: identical documents or lines removed by hash (SHA-256 over normalized text).
  • Near-deduplicated: documents above a Jaccard similarity threshold removed, typically via MinHash with locality-sensitive hashing [2]. The threshold (for example 0.8 on 5-gram shingles) and the number of permutations belong in the schedule.
  • Filtered: language ID, length floors, quality classifiers, and removal of machine-generated or templated text.
  • PII-processed: names, emails, phone numbers and account numbers removed or replaced. Replacement tokens such as [EMAIL] change the count slightly; decide whether they are billable. See PII redaction for LLM training data.

Operational data from support desks, CRMs and engineering trackers is especially duplicate-heavy because of macros, auto-replies and quoted threads. Pricing it on raw tokens can mean paying for the same canned response thousands of times.

Illustrative token-count schedule for a data license

The counting basis should live in a schedule attached to the license so that both parties can reproduce it. The following is a template you can adapt; it pairs with the data license term sheet and the comparison of pricing structures.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample entryWhy it matters
Counting tokenizercl100k_base via tiktoken, pinned package version and encoding hashRemoves tokenizer ambiguity; version drift changes counts
Text normalizationUnicode NFC; strip HTML; collapse whitespace runs; keep code indentationNormalization alters token counts and dedup matches
Unit of recordOne support ticket thread = one documentDefines what dedup compares
Exact dedupSHA-256 on normalized document textRemoves copies across exports
Near-dedupMinHash, 128 permutations, 5-gram shingles, Jaccard >= 0.8Removes templated and quoted near-copies
FiltersLanguage ID = en; min 50 tokens per document; drop auto-generated noticesDefines "usable"
PII handlingTyped placeholders; placeholders count as tokensPrevents disputes on redacted text
Billable countTokens after all steps above, in the delivered files onlyThe number the price multiplies
ManifestPer-file token count, document count, SHA-256 of each fileLets the buyer reconcile file by file
ToleranceBuyer recount within +/-1.5% of manifest total = acceptedAvoids arguing over rounding
AdjustmentOutside tolerance: invoice re-based on buyer count, or seller recounts with shared scriptConverts disputes into a procedure
Overlap exclusionDocuments with long exact matches in named open corpora excluded from billable countAvoids paying for freely available text

Keep the counting script itself, with its dependency lockfile, as an exhibit. A count that cannot be rerun from the same files and the same code is an estimate, not a delivery measure.

Verifying a delivered count within a tolerance band

Verification means rerunning the agreed pipeline on the delivered files and reconciling against the seller's manifest, not trusting a summary total. A workable procedure:

  1. Check integrity. Confirm every file's SHA-256 against the manifest and that the document count matches.
  2. Recount with the pinned tokenizer. Run the exact tokenizer version from the schedule over normalized text and compare per-file totals; differences concentrated in a few files usually point to encoding or normalization mismatches.
  3. Re-run dedup. Apply the same MinHash parameters [2] and measure residual near-duplicate rate. A residual above an agreed ceiling (for example 2% of documents) is a delivery defect, not a pricing nuance.
  4. Sample for filter quality. Read a stratified sample, such as 400 documents across sources and months, for boilerplate, machine text and missed PII.
  5. Check overlap with open corpora. Exact n-gram indexes such as infini-gram allow fast lookups against trillion-token public corpora [8], which shows whether a slice of the offer duplicates text you could already use under an open license, such as the roughly 8TB Common Pile [6].

Write the tolerance, the residual-duplicate ceiling and the adjustment mechanism into the license, and pair them with usage-reporting terms from the guide to audit rights in AI data licenses. Record the recount output as part of acceptance; it also helps with disclosure duties such as California's AB 2013, which, as of October 2026, requires developers of generative AI systems offered to Californians to post training-data documentation, including the number of data points, with disclosures due from 1 January 2026 [10].

Converting per-token quotes to comparable units

Normalize every offer to cost per usable deduplicated token under one tokenizer before comparing, then translate to per-document cost to sanity-check against record-priced offers. The method: rebase each quote to your pinned tokenizer using a sample, apply the observed survival rate through your dedup and filter pipeline, then divide total price by surviving tokens.

Illustrative example: invented to show structure; it does not describe an available dataset.

Offer AOffer BOffer C
Quoted basisRaw tokens, seller tokenizerDeduplicated tokens, cl100k_basePer document
Quoted volume12.0B tokens7.5B tokens9.0M documents
Rebase factor to your tokenizer (from sample)0.881.00780 tokens per document
Survival after your dedup and filters (sample)52%94%81%
Usable tokens5.49B7.05B5.69B
Total quoted price (index units)100100100
Index cost per billion usable tokens18.214.217.6

Offer A looks largest but loses more than half its volume to duplicates and tokenizer inflation. Offer C, priced per record, is competitive once converted. The same normalization logic applies to other modalities; speech offers are rebased per usable hour, as covered in speech data pricing per hour, and multi-vendor comparisons in comparing training data vendor quotes.

Failure modes that inflate a per-token invoice

Most per-token disputes come from a handful of predictable gaps in the counting definition. Watch for these before signing:

  • Tokenizer drift: the seller upgrades a tokenizer library between sample and delivery.
  • Format tokens: JSON keys, Markdown table pipes, HTML tags or XML wrappers counted as content. Agree whether delivery format overhead is billable.
  • Re-exports: incremental deliveries that resend prior months, billed again because dedup runs only within each batch. Require dedup against all prior deliveries.
  • Translated or synthetic padding: machine-translated copies or generated paraphrases counted as new tokens.
  • Unlicensed slices: content the seller cannot license for training; dataset licensing metadata is often missing or wrong, with an audit of 1,800+ text datasets finding license omission above 70% and error rates above 50% on popular hosting sites [9]. Confirm the rights grant covers pre-training per pre-training data license rights.

For continued pre-training on a narrow domain, small, clean corpora often matter more than volume; see sourcing domain corpora for continued pre-training. For broader price drivers beyond tokens, see what drives the price of licensed enterprise data.

When per-token pricing fits and when it does not

Per-token pricing fits large, homogeneous text corpora bought for pre-training or continued pre-training, where volume of clean text is the value driver. It fits poorly for small, high-signal datasets such as resolved support threads with outcome labels, engineering incident records or legal workflow documents, where value sits in the structure and labels rather than raw token volume. For those, per-record or flat-fee terms are often easier to verify.

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows, and manages licensing and ongoing purchases. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Buyers can describe the corpus they need to SourceX; requests are not a guarantee of a match, and nothing is contracted until a supplier agrees. More on rights and terms is in the data licensing for AI training hub and the wider AI data guides.

Sourcing a text corpus with a defined token basis

SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company, with pricing and allowed uses agreed in a license per deal. Personal details are removed or replaced before delivery, and delivery runs through private, access-controlled workflows only after an executed agreement. Describe the text corpus you need at sourcex.si/buyers.

Sources

  1. Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini (arXiv; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  2. A. Z. Broder, Proceedings of compression and complexity of sequences, "On the resemblance and containment of documents" (1997). https://ieeexplore.ieee.org/document/666900
  3. arXiv (2311.05741), "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
  4. Together AI documentation, "Fine-tuning pricing". https://docs.together.ai/docs/fine-tuning-pricing
  5. OpenAI Developer Community, "Doesn't understand fine tuned model cost". https://community.openai.com/t/doesnt-understand-fine-tuned-model-cost/80605
  6. arXiv (2506.05209), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
  7. arXiv (2510.13996), "The German Commons: 154 Billion Tokens of Openly Licensed Text for German Language Models" (2025). https://arxiv.org/html/2510.13996v1
  8. arXiv (2401.17377), "Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens" (2024). https://arxiv.org/html/2401.17377v4
  9. Longpre et al. (arXiv 2310.16787), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  10. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data