Data licensing for AI training
Per-token pricing for training data: how to count, audit and compare
Quick answer
A per-token price for licensed training text is only meaningful once the contract fixes four things: the tokenizer and version that does the counting, whether the count is taken before or after deduplication, filtering and PII removal, how the buyer verifies it, and what tolerance triggers an adjustment. Without those, two quotes at the same rate can differ in real cost by a large factor. Pin the counting basis in a schedule, reproduce the count on delivery, and normalize every offer to cost per usable deduplicated token before comparing.
By SourceX Editorial · Updated
Why the same corpus produces different token counts
The token count of a corpus is a property of the corpus and the tokenizer together, so an unqualified number such as "40 billion tokens" is not a unit of sale. A byte-level BPE tokenizer with a 32k vocabulary, a 100k-vocabulary tiktoken encoding and a SentencePiece unigram model will each segment the same text differently. The gap widens for code, tables, non-English text and OCR output, where tokenizers trained mostly on English web text fragment words into more pieces.
Research on adapting models to new languages measures this as fertility, the average number of tokens produced per word, and finds it is much higher where the tokenizer covers a language poorly [3]. For a buyer, high fertility means a seller counting with a poorly matched tokenizer reports more tokens, and a per-token price rewards the seller for that mismatch. Openly licensed corpora publish headline counts, such as 154 billion tokens for the German Commons [7], and those figures also depend on the tokenizer the authors used.
Treat the tokenizer as a contract term. See the glossary entries on tokens and token count for definitions, and the guide to estimating token counts before licensing for pre-offer sizing.
Data price per token is not compute price per token
Most published "price per training token" figures are compute billing for fine-tuning, not the price of licensed data. Fine-tuning providers bill per token processed, typically counting training plus validation tokens and varying by model size and method [4]. That count also multiplies with epochs: three passes over a one-billion-token file bill roughly three billion tokens of compute.
A data license price is a one-time or recurring fee for rights to the content, and it should never scale with how many epochs you run. If a seller's draft says "tokens consumed in training" or "tokens processed", strike it and replace it with "unique tokens delivered, counted as defined in Schedule A". Developer forums show how quickly billed counts diverge from a buyer's own estimate once epochs and tokenizer differences enter [5].
Choosing the counting basis: raw, deduplicated, filtered or delivered
The fairest basis for pre-training buyers is usually tokens remaining after exact and near-duplicate removal, quality filtering and PII handling, because those are the tokens that actually train the model. Raw counts overstate value: Lee et al. found long repeated substrings throughout common language-modeling datasets, including one sentence repeated more than 60,000 times in C4, and showed that deduplication reduces verbatim memorization [1].
Each stage removes tokens, and each needs a definition:
- Raw: everything in the export, including boilerplate, signatures, quoted email chains, templated tickets and log lines.
- Exact-deduplicated: identical documents or lines removed by hash (SHA-256 over normalized text).
- Near-deduplicated: documents above a Jaccard similarity threshold removed, typically via MinHash with locality-sensitive hashing [2]. The threshold (for example 0.8 on 5-gram shingles) and the number of permutations belong in the schedule.
- Filtered: language ID, length floors, quality classifiers, and removal of machine-generated or templated text.
- PII-processed: names, emails, phone numbers and account numbers removed or replaced. Replacement tokens such as
[EMAIL]change the count slightly; decide whether they are billable. See PII redaction for LLM training data.
Operational data from support desks, CRMs and engineering trackers is especially duplicate-heavy because of macros, auto-replies and quoted threads. Pricing it on raw tokens can mean paying for the same canned response thousands of times.
Illustrative token-count schedule for a data license
The counting basis should live in a schedule attached to the license so that both parties can reproduce it. The following is a template you can adapt; it pairs with the data license term sheet and the comparison of pricing structures.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example entry | Why it matters |
|---|---|---|
| Counting tokenizer | cl100k_base via tiktoken, pinned package version and encoding hash | Removes tokenizer ambiguity; version drift changes counts |
| Text normalization | Unicode NFC; strip HTML; collapse whitespace runs; keep code indentation | Normalization alters token counts and dedup matches |
| Unit of record | One support ticket thread = one document | Defines what dedup compares |
| Exact dedup | SHA-256 on normalized document text | Removes copies across exports |
| Near-dedup | MinHash, 128 permutations, 5-gram shingles, Jaccard >= 0.8 | Removes templated and quoted near-copies |
| Filters | Language ID = en; min 50 tokens per document; drop auto-generated notices | Defines "usable" |
| PII handling | Typed placeholders; placeholders count as tokens | Prevents disputes on redacted text |
| Billable count | Tokens after all steps above, in the delivered files only | The number the price multiplies |
| Manifest | Per-file token count, document count, SHA-256 of each file | Lets the buyer reconcile file by file |
| Tolerance | Buyer recount within +/-1.5% of manifest total = accepted | Avoids arguing over rounding |
| Adjustment | Outside tolerance: invoice re-based on buyer count, or seller recounts with shared script | Converts disputes into a procedure |
| Overlap exclusion | Documents with long exact matches in named open corpora excluded from billable count | Avoids paying for freely available text |
Keep the counting script itself, with its dependency lockfile, as an exhibit. A count that cannot be rerun from the same files and the same code is an estimate, not a delivery measure.
Verifying a delivered count within a tolerance band
Verification means rerunning the agreed pipeline on the delivered files and reconciling against the seller's manifest, not trusting a summary total. A workable procedure:
- Check integrity. Confirm every file's SHA-256 against the manifest and that the document count matches.
- Recount with the pinned tokenizer. Run the exact tokenizer version from the schedule over normalized text and compare per-file totals; differences concentrated in a few files usually point to encoding or normalization mismatches.
- Re-run dedup. Apply the same MinHash parameters [2] and measure residual near-duplicate rate. A residual above an agreed ceiling (for example 2% of documents) is a delivery defect, not a pricing nuance.
- Sample for filter quality. Read a stratified sample, such as 400 documents across sources and months, for boilerplate, machine text and missed PII.
- Check overlap with open corpora. Exact n-gram indexes such as infini-gram allow fast lookups against trillion-token public corpora [8], which shows whether a slice of the offer duplicates text you could already use under an open license, such as the roughly 8TB Common Pile [6].
Write the tolerance, the residual-duplicate ceiling and the adjustment mechanism into the license, and pair them with usage-reporting terms from the guide to audit rights in AI data licenses. Record the recount output as part of acceptance; it also helps with disclosure duties such as California's AB 2013, which, as of October 2026, requires developers of generative AI systems offered to Californians to post training-data documentation, including the number of data points, with disclosures due from 1 January 2026 [10].
Converting per-token quotes to comparable units
Normalize every offer to cost per usable deduplicated token under one tokenizer before comparing, then translate to per-document cost to sanity-check against record-priced offers. The method: rebase each quote to your pinned tokenizer using a sample, apply the observed survival rate through your dedup and filter pipeline, then divide total price by surviving tokens.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Offer A | Offer B | Offer C | |
|---|---|---|---|
| Quoted basis | Raw tokens, seller tokenizer | Deduplicated tokens, cl100k_base | Per document |
| Quoted volume | 12.0B tokens | 7.5B tokens | 9.0M documents |
| Rebase factor to your tokenizer (from sample) | 0.88 | 1.00 | 780 tokens per document |
| Survival after your dedup and filters (sample) | 52% | 94% | 81% |
| Usable tokens | 5.49B | 7.05B | 5.69B |
| Total quoted price (index units) | 100 | 100 | 100 |
| Index cost per billion usable tokens | 18.2 | 14.2 | 17.6 |
Offer A looks largest but loses more than half its volume to duplicates and tokenizer inflation. Offer C, priced per record, is competitive once converted. The same normalization logic applies to other modalities; speech offers are rebased per usable hour, as covered in speech data pricing per hour, and multi-vendor comparisons in comparing training data vendor quotes.
Failure modes that inflate a per-token invoice
Most per-token disputes come from a handful of predictable gaps in the counting definition. Watch for these before signing:
- Tokenizer drift: the seller upgrades a tokenizer library between sample and delivery.
- Format tokens: JSON keys, Markdown table pipes, HTML tags or XML wrappers counted as content. Agree whether delivery format overhead is billable.
- Re-exports: incremental deliveries that resend prior months, billed again because dedup runs only within each batch. Require dedup against all prior deliveries.
- Translated or synthetic padding: machine-translated copies or generated paraphrases counted as new tokens.
- Unlicensed slices: content the seller cannot license for training; dataset licensing metadata is often missing or wrong, with an audit of 1,800+ text datasets finding license omission above 70% and error rates above 50% on popular hosting sites [9]. Confirm the rights grant covers pre-training per pre-training data license rights.
For continued pre-training on a narrow domain, small, clean corpora often matter more than volume; see sourcing domain corpora for continued pre-training. For broader price drivers beyond tokens, see what drives the price of licensed enterprise data.
When per-token pricing fits and when it does not
Per-token pricing fits large, homogeneous text corpora bought for pre-training or continued pre-training, where volume of clean text is the value driver. It fits poorly for small, high-signal datasets such as resolved support threads with outcome labels, engineering incident records or legal workflow documents, where value sits in the structure and labels rather than raw token volume. For those, per-record or flat-fee terms are often easier to verify.
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows, and manages licensing and ongoing purchases. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Buyers can describe the corpus they need to SourceX; requests are not a guarantee of a match, and nothing is contracted until a supplier agrees. More on rights and terms is in the data licensing for AI training hub and the wider AI data guides.
Sourcing a text corpus with a defined token basis
SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company, with pricing and allowed uses agreed in a license per deal. Personal details are removed or replaced before delivery, and delivery runs through private, access-controlled workflows only after an executed agreement. Describe the text corpus you need at sourcex.si/buyers.
Sources
- Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini (arXiv; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- A. Z. Broder, Proceedings of compression and complexity of sequences, "On the resemblance and containment of documents" (1997). https://ieeexplore.ieee.org/document/666900
- arXiv (2311.05741), "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
- Together AI documentation, "Fine-tuning pricing". https://docs.together.ai/docs/fine-tuning-pricing
- OpenAI Developer Community, "Doesn't understand fine tuned model cost". https://community.openai.com/t/doesnt-understand-fine-tuned-model-cost/80605
- arXiv (2506.05209), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
- arXiv (2510.13996), "The German Commons: 154 Billion Tokens of Openly Licensed Text for German Language Models" (2025). https://arxiv.org/html/2510.13996v1
- arXiv (2401.17377), "Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens" (2024). https://arxiv.org/html/2401.17377v4
- Longpre et al. (arXiv 2310.16787), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.