Skip to content

Data licensing for AI training

Revenue-share data deals: attribution, reporting and when buyers should agree

Quick answer

A revenue-share data license pays the supplier a percentage of revenue from the product the data helps build, instead of, or on top of, a fixed fee. Buyers should agree only when four things are pinned down: a narrow revenue base (named product line, net of refunds, taxes and pass-through compute), an attribution rule that does not require proving one dataset's contribution, reporting you can produce from systems you already run, and a cap or buyout. Without those, a flat fee plus minimum guarantee is usually cheaper and safer.

By SourceX Editorial · Updated

Why suppliers ask for revenue share and what it costs the buyer

Suppliers ask for revenue share because they cannot price an operational dataset before they know what it is worth to your model, and a percentage lets them participate in upside. Public AI licensing deal values are mostly undisclosed, and the trackers that follow them show terms drifting from one-off training payments toward usage-linked and attribution-linked payments [1]. Publishers describe this as a move toward usage-based "grounding" or retrieval deals priced on different parameters from a one-time training transfer [3].

For a training-data buyer, the cost is not only the percentage. A revenue share creates a long-tail liability that survives model retraining, an audit relationship with a third party who sees your financials, and a valuation question at your next fundraising or acquisition. Compare that with the reported per-title flat fee model used in book licensing, where the price is fixed per unit and the licensee's revenue is irrelevant [4]. For the wider licensing context, start with the AI training data licensing buyer's guide. See AI data license pricing structures compared for the full set of alternatives and the SourceX insight on how AI data deals are priced.

Define the revenue base before you discuss the percentage

The revenue base matters more than the rate: 2% of gross platform revenue can cost more than 8% of net revenue from one SKU. Negotiate the base first and treat the percentage as the last variable.

A defensible base has five elements. First, a named product or SKU list (for example "the claims-triage model API and the hosted claims assistant"), not "any product that uses or benefits from the data." Second, net revenue: gross invoiced amounts less refunds, credits, chargebacks, sales and VAT taxes, reseller and marketplace fees, and pass-through costs such as third-party inference compute billed at cost. Third, a recognition rule tied to your own revenue policy (for most US startups, ASC 606 as applied in your audited financials) so you are not maintaining a second ledger. Fourth, explicit exclusions: professional services, implementation fees, hardware, and revenue from customers who contractually opt out of models trained on licensed data. Fifth, a bundling rule that allocates bundled contract value to the in-scope product using standalone selling prices, rather than letting the supplier claim the whole bundle.

Watch for three base-expanding phrases: "derived from," "enabled by," and "any successor model." The last ties your share to models you train years later; handle it in the license's derivative and successor model rights clause, not in the payment schedule.

Attribution: how to pay one dataset among many

Attribution is the hardest part, because a production model is trained on many sources and no contractually reliable method isolates one dataset's share of revenue. Pick a rule that is mechanical and auditable, not one that requires model science after every release. If you are still scoping which operational data you need, you can describe the dataset to SourceX before negotiating structure.

Three structures appear in practice:

  • Fixed share. The supplier gets a flat percentage of the in-scope base regardless of how much the dataset contributed. Simple to report; overpays if the dataset becomes a small part of later training mixes.
  • Pro rata by contribution. The share is scaled by the dataset's proportion of a measurable quantity: training tokens, records, or examples in the fine-tuning mix for the in-scope model. Volume is easy to measure but is a poor proxy for value; a small, high-quality set of resolved support tickets can move evaluation metrics more than millions of boilerplate records.
  • Tiered or usage-triggered. The rate steps down as revenue grows, or payment is triggered by retrieval events (each time a RAG system cites a licensed document). Usage-triggered models work for grounding and retrieval; the Benzinga and Dappier arrangement, for example, pairs attribution links in AI answers with an ad-revenue share [2].

Research methods for measuring contribution exist but are not contract-grade. Data Shapley values each source by its marginal effect on model performance across subsets and needs Monte Carlo or gradient approximations at scale [5]; datamodels predict a model's output from which training examples were included [6]. Both depend on the chosen performance metric, the evaluation set and random seeds, so two honest runs can disagree. Use them, if at all, once at signing to set a fixed share, never as a recurring payment calculation. Attribution metadata is also fragile upstream: an audit of more than 1,800 text datasets found widespread license omissions and errors on hosting platforms [8], so confirm in the data rights attestation that the supplier actually controls what it is being paid for.

Worked example: calculating a capped revenue share

The cleanest way to test a proposal is to model it on next year's plan with a worked calculation.

Illustrative example: invented to show structure; it does not describe an available dataset.

LineBasisAmount (USD)
Gross invoiced revenue, in-scope SKU (FY)Billing system export4,000,000
Less refunds and creditsCredit memos(120,000)
Less sales tax and VAT collectedTax engine(180,000)
Less marketplace fees (cloud marketplace listing)Marketplace statements(200,000)
Less pass-through inference compute billed at costUsage ledger(500,000)
Net in-scope revenue (base)3,000,000
Pro-rata factor: licensed tokens / fine-tuning mix tokens30M / 120M25%
Contractual rateTerm sheet6%
Royalty before cap: 3,000,000 x 6% x 25%45,000
Minimum guarantee (credited against royalties)Term sheet30,000
Annual capTerm sheet150,000
Amount payable for the yearmax(MG, royalty), capped45,000

Run the same table at two and five times plan. If the uncapped royalty at five times plan exceeds what a buyout or flat fee would cost, negotiate the cap or a buyout option now, while your revenue is small.

Reporting and audit obligations you will carry

A revenue share turns you into a reporting party, so the reporting clause should match data you already produce for your auditors and board. Ask for quarterly statements within a fixed period after quarter close, showing gross in-scope revenue, each deduction category, the base, the contribution factor and the payment, all reconcilable to your general ledger.

Limit audit rights tightly. Common buyer positions: one audit per year, by an independent CPA firm bound by confidentiality, at the supplier's cost unless underpayment exceeds a threshold (often 5%), limited to the last two or three reported years, and scoped to in-scope revenue records only, not model weights, training logs or customer lists. Your audit and usage-reporting rights clause should also state that the supplier's auditor sees revenue figures, not unit economics by customer.

Mind the supplier's accounting too. Under ASC 606, licensors of intellectual property generally recognize sales- or usage-based royalties only as the underlying sales or usage occur, while a minimum guarantee can be recognized earlier [7]. Suppliers that need predictable revenue therefore push hard for minimum guarantees, which gives you room to trade a guarantee for a lower rate or a cap. Most-favored-nation requests often ride along with revenue shares; see MFN clauses in AI data deals.

Revenue share vs flat fee: when buyers should agree

Agree to a revenue share when the supplier's data is central to one identifiable product and cash is scarce today; prefer a flat fee plus minimum guarantee when the data is one input among many or your revenue base is hard to isolate.

Illustrative example: invented to show structure; it does not describe an available dataset.

SituationBetter fitWhy
Pre-revenue vertical AI product, single dataset is the core differentiatorRevenue share with cap and buyoutPreserves cash; payment tracks value
Foundation or general-purpose model trained on hundreds of sourcesFlat fee or per-record priceContribution cannot be isolated; base would be the whole company
RAG or grounding product that cites documents at answer timeUsage-triggered fee per retrievalUsage is logged anyway; attribution is native [3]
Fine-tuning data for an internal tool with no external revenueFlat feeThere is no revenue base to share
Supplier wants upside but you expect an acquisition within three yearsFlat fee plus minimum guarantee, or revenue share with change-of-control buyoutAcquirers discount open-ended royalties

Whichever you choose, make sure payment terms survive correctly when the license ends: royalties on models already shipped should follow the model retention after license termination clause, not continue indefinitely by default.

Revenue-share term sheet checklist

Put these items in the term sheet before counsel drafts the license; the AI data license term sheet has the surrounding fields.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • In-scope products named by SKU; "successor model" coverage stated explicitly
  • Net revenue definition with every deduction category listed
  • Recognition rule tied to your audited revenue policy
  • Attribution method (fixed, pro rata by tokens or records, tiered, or per retrieval) and the data source for the factor
  • Rate, tiers, minimum guarantee, crediting of the guarantee, annual cap
  • Buyout price or formula, including on change of control
  • Reporting cadence, statement fields and payment timing
  • Audit frequency, lookback, auditor independence, confidentiality and cost shifting
  • Treatment of royalties after termination or deletion
  • Interaction with MFN, exclusivity and warranty clauses

Sourcing operational data for a revenue-share deal

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases; pricing and allowed uses are agreed per deal in a license, and nothing is contracted until a supplier agrees. Every dataset is rights-reviewed for ownership and consents before delivery. Describe the data your product needs and the structure you are considering at https://sourcex.si/buyers.

Frequently asked questions

Is a royalty-based data license the same as a revenue share?

They are usually used interchangeably. Some drafts reserve "royalty" for a per-unit payment (per seat, per API call, per retrieval) and "revenue share" for a percentage of revenue; check which base the clause actually uses. The revenue share glossary entry gives the short definition.

Can a supplier demand to see our training mix to verify a pro-rata factor?

Only if you agree to it. A better position is a certified statement of licensed tokens or records and total mix size for the in-scope model, verifiable by the independent auditor, without disclosing other sources' identities.

Should the revenue share apply to model outputs or synthetic data we generate?

Decide explicitly. If you plan to generate synthetic data from licensed records, define whether revenue from models trained on that synthetic data is in the base; see rights to synthetic data derived from licensed data.

Sources

  1. LLM Pulse, "Every AI Content Licensing Deal, Mapped (2023-2026)" (2025). https://llmpulse.ai/blog/ai-content-licensing-deals/
  2. AdExchanger, "AI search adoption is boosting Benzinga's data licensing biz". https://www.adexchanger.com/publishers/ai-search-adoption-is-boosting-benzingas-data-licensing-biz
  3. Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?". https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
  4. eMarketer, "Microsoft and HarperCollins sign AI licensing deal, but author opt-in still required" (2024). https://www.emarketer.com/content/microsoft-harpercollins-sign-ai-licensing-deal--author-opt-in-still-required
  5. Ghorbani and Zou, "Data Shapley: Equitable Valuation of Data for Machine Learning" (2019). https://arxiv.org/abs/1904.02868v2
  6. Ilyas, Park, Engstrom, Leclerc and Madry, "Datamodels: Understanding Predictions with Data and Data with Predictions" (2022). https://proceedings.mlr.press/v162/ilyas22a.html
  7. Deloitte DART, "12.7 Sales- or Usage-Based Royalties (Roadmap: Revenue Recognition)". https://dart.deloitte.com/USDART/home/codification/revenue/asc606-10/roadmap-revenue-recognition/chapter-12-licensing/12-7-sales-or-usage-based
  8. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://www.nature.com/articles/s42256-024-00878-8

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data