Data licensing for AI training
Revenue-share data deals: attribution, reporting and when buyers should agree
Quick answer
A revenue-share data license pays the supplier a percentage of revenue from the product the data helps build, instead of, or on top of, a fixed fee. Buyers should agree only when four things are pinned down: a narrow revenue base (named product line, net of refunds, taxes and pass-through compute), an attribution rule that does not require proving one dataset's contribution, reporting you can produce from systems you already run, and a cap or buyout. Without those, a flat fee plus minimum guarantee is usually cheaper and safer.
By SourceX Editorial · Updated
Why suppliers ask for revenue share and what it costs the buyer
Suppliers ask for revenue share because they cannot price an operational dataset before they know what it is worth to your model, and a percentage lets them participate in upside. Public AI licensing deal values are mostly undisclosed, and the trackers that follow them show terms drifting from one-off training payments toward usage-linked and attribution-linked payments [1]. Publishers describe this as a move toward usage-based "grounding" or retrieval deals priced on different parameters from a one-time training transfer [3].
For a training-data buyer, the cost is not only the percentage. A revenue share creates a long-tail liability that survives model retraining, an audit relationship with a third party who sees your financials, and a valuation question at your next fundraising or acquisition. Compare that with the reported per-title flat fee model used in book licensing, where the price is fixed per unit and the licensee's revenue is irrelevant [4]. For the wider licensing context, start with the AI training data licensing buyer's guide. See AI data license pricing structures compared for the full set of alternatives and the SourceX insight on how AI data deals are priced.
Define the revenue base before you discuss the percentage
The revenue base matters more than the rate: 2% of gross platform revenue can cost more than 8% of net revenue from one SKU. Negotiate the base first and treat the percentage as the last variable.
A defensible base has five elements. First, a named product or SKU list (for example "the claims-triage model API and the hosted claims assistant"), not "any product that uses or benefits from the data." Second, net revenue: gross invoiced amounts less refunds, credits, chargebacks, sales and VAT taxes, reseller and marketplace fees, and pass-through costs such as third-party inference compute billed at cost. Third, a recognition rule tied to your own revenue policy (for most US startups, ASC 606 as applied in your audited financials) so you are not maintaining a second ledger. Fourth, explicit exclusions: professional services, implementation fees, hardware, and revenue from customers who contractually opt out of models trained on licensed data. Fifth, a bundling rule that allocates bundled contract value to the in-scope product using standalone selling prices, rather than letting the supplier claim the whole bundle.
Watch for three base-expanding phrases: "derived from," "enabled by," and "any successor model." The last ties your share to models you train years later; handle it in the license's derivative and successor model rights clause, not in the payment schedule.
Attribution: how to pay one dataset among many
Attribution is the hardest part, because a production model is trained on many sources and no contractually reliable method isolates one dataset's share of revenue. Pick a rule that is mechanical and auditable, not one that requires model science after every release. If you are still scoping which operational data you need, you can describe the dataset to SourceX before negotiating structure.
Three structures appear in practice:
- Fixed share. The supplier gets a flat percentage of the in-scope base regardless of how much the dataset contributed. Simple to report; overpays if the dataset becomes a small part of later training mixes.
- Pro rata by contribution. The share is scaled by the dataset's proportion of a measurable quantity: training tokens, records, or examples in the fine-tuning mix for the in-scope model. Volume is easy to measure but is a poor proxy for value; a small, high-quality set of resolved support tickets can move evaluation metrics more than millions of boilerplate records.
- Tiered or usage-triggered. The rate steps down as revenue grows, or payment is triggered by retrieval events (each time a RAG system cites a licensed document). Usage-triggered models work for grounding and retrieval; the Benzinga and Dappier arrangement, for example, pairs attribution links in AI answers with an ad-revenue share [2].
Research methods for measuring contribution exist but are not contract-grade. Data Shapley values each source by its marginal effect on model performance across subsets and needs Monte Carlo or gradient approximations at scale [5]; datamodels predict a model's output from which training examples were included [6]. Both depend on the chosen performance metric, the evaluation set and random seeds, so two honest runs can disagree. Use them, if at all, once at signing to set a fixed share, never as a recurring payment calculation. Attribution metadata is also fragile upstream: an audit of more than 1,800 text datasets found widespread license omissions and errors on hosting platforms [8], so confirm in the data rights attestation that the supplier actually controls what it is being paid for.
Worked example: calculating a capped revenue share
The cleanest way to test a proposal is to model it on next year's plan with a worked calculation.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Line | Basis | Amount (USD) |
|---|---|---|
| Gross invoiced revenue, in-scope SKU (FY) | Billing system export | 4,000,000 |
| Less refunds and credits | Credit memos | (120,000) |
| Less sales tax and VAT collected | Tax engine | (180,000) |
| Less marketplace fees (cloud marketplace listing) | Marketplace statements | (200,000) |
| Less pass-through inference compute billed at cost | Usage ledger | (500,000) |
| Net in-scope revenue (base) | 3,000,000 | |
| Pro-rata factor: licensed tokens / fine-tuning mix tokens | 30M / 120M | 25% |
| Contractual rate | Term sheet | 6% |
| Royalty before cap: 3,000,000 x 6% x 25% | 45,000 | |
| Minimum guarantee (credited against royalties) | Term sheet | 30,000 |
| Annual cap | Term sheet | 150,000 |
| Amount payable for the year | max(MG, royalty), capped | 45,000 |
Run the same table at two and five times plan. If the uncapped royalty at five times plan exceeds what a buyout or flat fee would cost, negotiate the cap or a buyout option now, while your revenue is small.
Reporting and audit obligations you will carry
A revenue share turns you into a reporting party, so the reporting clause should match data you already produce for your auditors and board. Ask for quarterly statements within a fixed period after quarter close, showing gross in-scope revenue, each deduction category, the base, the contribution factor and the payment, all reconcilable to your general ledger.
Limit audit rights tightly. Common buyer positions: one audit per year, by an independent CPA firm bound by confidentiality, at the supplier's cost unless underpayment exceeds a threshold (often 5%), limited to the last two or three reported years, and scoped to in-scope revenue records only, not model weights, training logs or customer lists. Your audit and usage-reporting rights clause should also state that the supplier's auditor sees revenue figures, not unit economics by customer.
Mind the supplier's accounting too. Under ASC 606, licensors of intellectual property generally recognize sales- or usage-based royalties only as the underlying sales or usage occur, while a minimum guarantee can be recognized earlier [7]. Suppliers that need predictable revenue therefore push hard for minimum guarantees, which gives you room to trade a guarantee for a lower rate or a cap. Most-favored-nation requests often ride along with revenue shares; see MFN clauses in AI data deals.
Revenue share vs flat fee: when buyers should agree
Agree to a revenue share when the supplier's data is central to one identifiable product and cash is scarce today; prefer a flat fee plus minimum guarantee when the data is one input among many or your revenue base is hard to isolate.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Situation | Better fit | Why |
|---|---|---|
| Pre-revenue vertical AI product, single dataset is the core differentiator | Revenue share with cap and buyout | Preserves cash; payment tracks value |
| Foundation or general-purpose model trained on hundreds of sources | Flat fee or per-record price | Contribution cannot be isolated; base would be the whole company |
| RAG or grounding product that cites documents at answer time | Usage-triggered fee per retrieval | Usage is logged anyway; attribution is native [3] |
| Fine-tuning data for an internal tool with no external revenue | Flat fee | There is no revenue base to share |
| Supplier wants upside but you expect an acquisition within three years | Flat fee plus minimum guarantee, or revenue share with change-of-control buyout | Acquirers discount open-ended royalties |
Whichever you choose, make sure payment terms survive correctly when the license ends: royalties on models already shipped should follow the model retention after license termination clause, not continue indefinitely by default.
Revenue-share term sheet checklist
Put these items in the term sheet before counsel drafts the license; the AI data license term sheet has the surrounding fields.
Illustrative example: invented to show structure; it does not describe an available dataset.
- In-scope products named by SKU; "successor model" coverage stated explicitly
- Net revenue definition with every deduction category listed
- Recognition rule tied to your audited revenue policy
- Attribution method (fixed, pro rata by tokens or records, tiered, or per retrieval) and the data source for the factor
- Rate, tiers, minimum guarantee, crediting of the guarantee, annual cap
- Buyout price or formula, including on change of control
- Reporting cadence, statement fields and payment timing
- Audit frequency, lookback, auditor independence, confidentiality and cost shifting
- Treatment of royalties after termination or deletion
- Interaction with MFN, exclusivity and warranty clauses
Sourcing operational data for a revenue-share deal
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases; pricing and allowed uses are agreed per deal in a license, and nothing is contracted until a supplier agrees. Every dataset is rights-reviewed for ownership and consents before delivery. Describe the data your product needs and the structure you are considering at https://sourcex.si/buyers.
Frequently asked questions
Is a royalty-based data license the same as a revenue share?
They are usually used interchangeably. Some drafts reserve "royalty" for a per-unit payment (per seat, per API call, per retrieval) and "revenue share" for a percentage of revenue; check which base the clause actually uses. The revenue share glossary entry gives the short definition.
Can a supplier demand to see our training mix to verify a pro-rata factor?
Only if you agree to it. A better position is a certified statement of licensed tokens or records and total mix size for the in-scope model, verifiable by the independent auditor, without disclosing other sources' identities.
Should the revenue share apply to model outputs or synthetic data we generate?
Decide explicitly. If you plan to generate synthetic data from licensed records, define whether revenue from models trained on that synthetic data is in the base; see rights to synthetic data derived from licensed data.
Sources
- LLM Pulse, "Every AI Content Licensing Deal, Mapped (2023-2026)" (2025). https://llmpulse.ai/blog/ai-content-licensing-deals/
- AdExchanger, "AI search adoption is boosting Benzinga's data licensing biz". https://www.adexchanger.com/publishers/ai-search-adoption-is-boosting-benzingas-data-licensing-biz
- Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?". https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
- eMarketer, "Microsoft and HarperCollins sign AI licensing deal, but author opt-in still required" (2024). https://www.emarketer.com/content/microsoft-harpercollins-sign-ai-licensing-deal--author-opt-in-still-required
- Ghorbani and Zou, "Data Shapley: Equitable Valuation of Data for Machine Learning" (2019). https://arxiv.org/abs/1904.02868v2
- Ilyas, Park, Engstrom, Leclerc and Madry, "Datamodels: Understanding Predictions with Data and Data with Predictions" (2022). https://proceedings.mlr.press/v162/ilyas22a.html
- Deloitte DART, "12.7 Sales- or Usage-Based Royalties (Roadmap: Revenue Recognition)". https://dart.deloitte.com/USDART/home/codification/revenue/asc606-10/roadmap-revenue-recognition/chapter-12-licensing/12-7-sales-or-usage-based
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://www.nature.com/articles/s42256-024-00878-8
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.