Data licensing for AI training
AI data licensing pricing models compared: flat fee, per-record, per-token, subscription and revenue share
Quick answer
AI data licensing pricing models come in six structures: a flat fee for a defined corpus, a per-unit price (per record, document, audio hour or title), a per-token price, a subscription for access and refreshes, a revenue share or per-use fee, and a minimum guarantee recouped against royalties. Choose the structure whose billing unit tracks how you will use the data and how certain that use is, then convert every quote to cost per usable unit at the same rights scope before comparing.
By SourceX Editorial · Updated
Six pricing structures and the unit each one bills
Each structure answers two questions differently: which unit the licensor counts, and who carries the risk that volume or usage differs from the forecast. The AI training data licensing guide maps the rights terms that sit beside price.
| Structure | Billing unit | Fits when | Risk the buyer carries |
|---|---|---|---|
| Flat fee | A defined corpus or snapshot, such as "all closed tickets, 2021 to 2025" | Scope and volume are known after sample review; one training or fine-tuning program | Duplicates and unusable records are paid for; any scope expansion is repriced |
| Per record or unit | Record, thread, document, image, audio hour or title | You can select queues, years or regions, and units are similar in value | A loose unit definition; duplicates and failed units billed |
| Per token | Tokens counted by a named tokenizer | Large text corpora for pre-training or continued pre-training | Tokenizer choice, markup and boilerplate inflate the count |
| Subscription | A term of access, usually with a refresh cadence | Data goes stale (tickets, listings, logs) or you draw from a catalog over time | What you may keep, and keep using, when the term ends |
| Revenue share or per-use fee | A share of defined revenue, or a fee per crawl, retrieval, display or inference | Each use is traceable, as in retrieval and grounding | Attribution, reporting and audit cost; no ceiling unless capped |
| Minimum guarantee plus royalty | An advance recouped against usage royalties | Usage is uncertain and the licensor wants upside | Paying the floor for unused volume; recoupment and true-up rules |
Published examples show suppliers mixing these units. HarperCollins's 2024 program was reported as a per-title fee of $5,000, split evenly between author and publisher, under an opt-in, three-year license for selected nonfiction titles [1][2]. One publisher's licensing page lists a non-exclusive per-book training license and a non-exclusive full-corpus annual license with quarterly refresh side by side [3] (a vendor page, cited as market practice).
The Linguistic Data Consortium's for-profit membership is an annual, non-refundable fee for a January-to-December membership year, with use tied to the sites listed in the agreement [4]. The RSL machine-readable licensing standard, launched in September 2025, lists free, attribution, subscription, pay-per-crawl and pay-per-inference models [5].
Match the structure to the use you are licensing
The intended use, more than the data type, decides which structure is cheapest to administer and fairest to both sides. Volume-based units suit training, where data is consumed once; usage-based units suit retrieval, where each query touches licensed content.
| Intended use | Structures that usually fit | Why | Check before signing |
|---|---|---|---|
| Evaluation only | Flat fee per evaluation set; samples under an evaluation license | Volume is small and no model learns from it | Evaluation grants tend to be narrow: NVIDIA's sample-data license is limited, revocable and solely for evaluating and testing NVIDIA technologies [6]. See evaluation-only license terms |
| Fine-tuning a named model | Flat fee or per record | Volume is known after sample review | Price successor-model rights and model retention separately; see fine-tuning-only licenses |
| Pre-training or continued pre-training | Per token, or flat fee per corpus snapshot | Value scales with text volume | Tokenizer definition and any exclusivity premium; see pre-training rights grants |
| Retrieval and grounding | Subscription plus per-use fees | Content is fetched at query time; a lawyer quoted in industry coverage calls per-use fees the largest part of grounding-deal fees [7] | A grounding license grants no training rights unless it says so; see grounding vs training licenses |
| Refreshed operational data | Subscription with a volume band per delivery | Value decays as records age | Refresh specification and remedies for short or late deliveries; see subscription licenses for refreshed data |
| Product with uncertain revenue | Minimum guarantee plus royalty, or revenue share | Spreads risk across both parties | Reporting duties and caps |
Rights scope often moves price more than the structure does. The same records cost more when the grant adds pre-training, successor models, customer sublicensing or exclusivity, because the licensor gives up more future value; see negotiating exclusivity and per-model vs enterprise-wide licenses. Check what scope a cheap quote grants, and do not rely on a license label alone: an audit of more than 1,800 text datasets found license omission rates above 70% and error rates above 50% on popular hosting sites [8].
SourceX does not publish prices: terms depend on scope, volume, history, rights and exclusivity, and are agreed per deal in writing. Listing your intended uses in a data request to SourceX lets scope shape the terms from the start.
Fixed, recurring or variable: how each structure lands in a budget
Map each structure to a commitment type before comparing totals, because the same spend can be a one-time cost, a recurring line or an open-ended variable.
- Flat fee. One commitment at signature. Split payment across delivery acceptance milestones so cash follows usable data.
- Per record and per token. Set a not-to-exceed amount and a tolerance band for counts, since the invoice depends on what is selected and counted.
- Subscription. A recurring line with a renewal date and escalator. Access can lapse automatically: Delta Sharing recipients authenticate with a bearer token from a profile file that can carry an expiration time [9]. Whether models trained during the term survive is a contract question; see what happens to models when a license ends.
- Revenue share and per-use fees. Forecast from expected usage, then negotiate a cap or tiered rate so the line has a ceiling.
- Minimum guarantee. Budget the floor as fixed and the royalty above it as variable, with a written recoupment schedule.
The AI training data budget guide covers phasing spend across samples, first purchase and refreshes.
Worked example: three quotes for one support-ticket corpus
Converting every quote to a fee per usable unit, at matched scope, is the only fair way to compare a flat fee, a per-record price and a per-token price. The full method is in comparing data vendor quotes.
Illustrative example: invented to show structure; it does not describe an available dataset. Figures are arbitrary amounts in a generic currency, not market prices.
A buyer needs English customer-support threads (customer messages, agent replies, resolution code) to fine-tune a support agent and hold out an evaluation set.
| Offer A: flat fee | Offer B: per thread | Offer C: per token | |
|---|---|---|---|
| Quote | 150,000 for the full export | 0.12 per delivered thread | 0.40 per 1,000 tokens, supplier's count |
| Volume quoted | 1,200,000 threads | 900,000 threads from selected queues | 600 million tokens (about 1,150,000 threads) |
| Duplicates, including auto-replies | 18% | 3% | 15% |
| Non-English or empty after PII removal | 7% | 4% | 6% |
| Usable threads | about 915,000 | about 838,000 | about 919,000 |
| Total fee | 150,000 | 108,000 | 240,000 |
| Fee per usable thread | about 0.164 | about 0.129 | about 0.261 |
| Scope quoted | Fine-tuning and evaluation, one model family, 3 years | Fine-tuning and evaluation, one named model, 2 years | Pre-training, fine-tuning and successor models, no end date |
Under a flat fee the buyer pays for filtering losses, so Offer A's 18% duplicate rate works as a price increase. Offer C's count came from the supplier's tokenizer; the buyer's tokenizer counts 540 million tokens, so billing on the buyer's count would cut the fee to 216,000, or about 0.235 per usable thread. Offer B looks cheapest, but it also grants the least, so ask A and C to quote B's scope, or B to quote C's, before ranking them. Costs outside the license fee, such as legal review and integration, belong in a total cost of ownership model.
Per-record and per-token quotes: define the billable unit first
A unit price is comparable only when both sides define the unit and the point at which it is counted.
- Records. State whether a record is a ticket or each message, a claim or each claim line, a call or each channel-hour. For speech, state whether hours are measured before or after silence trimming; see how speech data is priced.
- Duplicates and overlap. Bill unique units after an agreed deduplication rule, and exclude records that duplicate data you already hold; see overlap checks against your own data.
- Counting point. Bill accepted units rather than delivered units, and tie replacement or credit to the acceptance test; see acceptance sampling for deliveries.
- Tokens. Name the tokenizer and version, and say whether markup, boilerplate and removed PII spans count. Tokens per word ("fertility") vary by tokenizer and language: one study improved fertility by 42% for Hungarian and 73% for Thai by replacing about 10% of a model's vocabulary [10]. See per-token training data pricing and estimating token counts before licensing.
Revenue share and per-use fees: the reporting you sign up for
Usage-based structures move volume risk to the licensor, but the buyer pays with measurement, reporting and audit duties a flat fee never requires. They work when each use can be logged and attributed, and poorly when one dataset is blended into a pre-training mix.
Published examples cluster in content for retrieval. The News/Media Alliance's non-exclusive deal with Bria pays publishers according to how often their content is used in enterprise AI outputs, splitting revenue 50-50 under an attribution model Bria developed [11]. RSL lets publishers declare pay-per-inference terms in machine-readable form [5]. For a dataset absorbed into model weights, there is usually no event log that ties product revenue to one source, so a royalty tends to be a negotiated proxy rather than a measurement.
Before agreeing, pin down:
- The usage event: crawl, retrieval, display, citation or inference, and the log field that records it.
- The revenue base: which product, gross or net, and how bundles and free tiers count (see SourceX's revenue share glossary entry).
- Reporting cadence and format, plus the audit rights the licensor may exercise.
- A cap or tiered rate, so variable cost has a ceiling.
- The share's fate if the product is retired, merged or repriced.
Details are in revenue-share data deals for buyers and usage-based pricing for RAG content.
Headline AI deal values are not a price list
Public AI licensing figures show which structures exist, not what your dataset should cost. They often bundle rights, terms and non-cash value that a quote for one dataset does not include.
Reddit's licensing deal with Google was reported in 2024 at about $60 million a year, with terms undisclosed [12]. Ithaka S+R's tracker records deal type and size only where information is available [13], and a 2023 to 2026 deal map compiles press-reported values, many of them estimates [14]. As of October 2026, the U.S. Copyright Office's Part 3 report on generative AI training remains a May 2025 pre-publication version; it declined to recommend new legislation at that time and said the licensing market for AI training data should be allowed to develop [15].
Use public deals to learn units (per title, per year, per use), then build your estimate from the drivers in SourceX's guide to what drives the price of licensed enterprise data. Why public AI deal values don't price your data explains the gaps in detail.
Price-schedule terms to settle before signature
A pricing structure is only as reliable as the order form that implements it. Check each item against the draft; the AI data license term sheet and the negotiation checklist cover the surrounding terms.
- Billable unit defined in words and by a named counting method (script, query or tokenizer version).
- Counting point stated: delivered, accepted, or after deduplication.
- Credits for duplicates, overlap with data you already hold, and units that fail acceptance.
- Payments tied to acceptance milestones, not only to signature.
- Scope priced explicitly (uses, models, term, territory, affiliates), with options priced now for pre-training, successor models or exclusivity.
- Refresh terms: cadence, volume band per delivery, price per refresh, and the remedy for short or late deliveries.
- Renewal: notice period and an escalator expressed as a formula or cap.
- Variable fees: usage event, revenue base, reporting cadence, cap and audit right.
- Minimum guarantee: recoupment order, carry-forward of unused credit, and the true-up date.
- Add-ons priced separately: annotation, de-identification, format conversion and custom extraction.
- Transfer costs assigned: with an Amazon S3 Requester Pays bucket, the requester pays for requests and downloads and the bucket owner pays for storage [16].
- Termination: refunds or credits for supplier breach, fees that survive, and rights to keep trained models.
Comparing pricing structures for a dataset you need?
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Describe the data and the uses you need licensed: SourceX looks for US businesses that hold it, checks the data and the supplier's licensing permissions, agrees pricing and allowed uses in a license, and coordinates delivery and payment. Talk to SourceX about your requirements.
Sources
- Authors Guild, "HarperCollins AI Licensing Deal" (2024). https://authorsguild.org/news/harpercollins-ai-licensing-deal/
- The Decoder, "Microsoft offers authors $5,000 to train AI on their books" (2024). https://the-decoder.com/microsoft-offers-authors-5000-to-train-ai-on-their-books/
- Source Library, "AI & Data-Mining Licensing" (vendor page, market practice only). https://sourcelibrary.org/licensing
- Linguistic Data Consortium, "LDC For-Profit Membership Agreement". https://Catalog.Ldc.Upenn.Edu/license/ldc-for-profit-membership.pdf
- RSL Collective, "RSL standard" (launch press release, 2025). https://rslstandard.org/press/rsl-standard
- NVIDIA, "NVIDIA Sample Data License for Evaluation" (2026). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
- Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?" (2025). https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- delta-io/delta-sharing documentation, "Delta Sharing protocol: REST APIs". https://www.mintlify.com/delta-io/delta-sharing/protocol/rest-apis
- arXiv, "Efficiently Adapting Pretrained Language Models To New Languages" (2023). https://arxiv.org/pdf/2311.05741
- Digiday, "News/Media Alliance signs AI licensing deal to unlock recurring RAG revenue for small and mid-sized publishers". https://digiday.com/media-buying/news-media-alliance-signs-ai-licensing-deal-to-unlock-recurring-rag-revenue-for-small-and-mid-sized-publishers/
- Engadget, "Reddit is licensing its content to Google to help train its AI models" (2024). https://engadget.com/reddit-is-licensing-its-content-to-google-to-help-train-its-ai-models-200013007.html
- Ithaka S+R, "Generative AI Licensing Agreement Tracker". https://sr.ithaka.org/our-work/generative-ai-licensing-agreement-tracker/
- LLM Pulse, "Every AI Content Licensing Deal, Mapped (2023-2026)" (2026). https://llmpulse.ai/blog/ai-content-licensing-deals/
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- Amazon Web Services, "Using Requester Pays general purpose buckets for storage transfers and usage" (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.