Retrieval, RAG and grounding data
Usage-based pricing for RAG content: per crawl, per retrieval, per display
Quick answer
Usage-based RAG content licenses charge for events rather than for a corpus: each crawl or fetch, each retrieval of a chunk into a prompt, each display or citation shown to a user, or each inference that uses the content. As of October 2026, the usage component often carries most of the value in grounding deals [1]. Budget well by defining the billable event precisely, forecasting each event type separately, and negotiating caps, minimum guarantees and true-ups so a traffic spike does not become an unbudgeted invoice.
By SourceX Editorial · Updated
Why grounding content is priced per use, not per record
Grounding content is priced per use because its value is realized at query time, every time, rather than once at training. A training license typically pays for a fixed copy that is absorbed into weights; a grounding license pays for content that is fetched, ranked and shown repeatedly, and one does not authorize the other [6]. Licensors therefore want fees that scale with exposure, and vendors in the space frame each retrieval as a potential transaction [5].
For a buyer this changes the budgeting problem. Per-record or per-token training pricing, covered in per-token pricing for training data and the pricing benchmarks for licensed enterprise data, is a capital-style purchase. Usage pricing is an operating cost tied to your product's traffic, retrieval depth and answer design. Read the broader rights picture in the RAG content licensing buyer's guide before you model fees.
The five billable units and what each one actually counts
Each pricing unit counts a different event in your pipeline, and they can differ by orders of magnitude for the same user query. Machine-readable licensing has made several of these units explicit: the Really Simple Licensing (RSL) standard, launched in September 2025, lets publishers declare free, attribution, subscription, pay-per-crawl and pay-per-inference terms [2][3].
| Unit | Event that triggers a charge | Where you measure it | Typical failure mode |
|---|---|---|---|
| Per crawl / fetch | Your crawler or connector requests a URL or API object | Crawler logs, HTTP status codes, licensor billing reports | Recrawling unchanged pages to refresh the index |
| Per retrieval | A chunk from the licensed corpus enters a prompt context | Retriever logs (doc_id, chunk_id, rank, query_id) | Top-k of 20 billed when only 3 chunks are used |
| Per display / citation | A snippet, link or attribution is shown to an end user | Front-end render events | Counting hidden or collapsed citations |
| Per inference | A model response is generated using the content | Generation logs joined to retrieval logs | Agent loops calling the model many times per user turn |
| Subscription / flat fee | Access for a period, often with a usage ceiling | Contract and usage reports | Ceiling breached without notice; overage at list rate |
RSL also depends on AI companies choosing to honor declared terms [2], so a negotiated contract still has to define units in its own words. Revenue share is a sixth model that sits beside these: in one reported arrangement, AI answers that use a financial publisher's data link back to the source and share ad revenue from those answers [7]. See revenue share for how the base is defined.
How to forecast usage before you sign
Forecast each billable event separately, because one user query can produce one crawl, twenty retrievals, three displayed citations and five model calls. Start from product telemetry, not from the licensor's assumptions. If the content is not yet in your stack, a pilot for retrieval lift gives you real retrieval and display rates on your own queries.
Four ratios drive the forecast. Retrievals per query depend on top-k and reranker settings; displays per query depend on your citation UI; inferences per query depend on agent design; and fetches per document depend on refresh cadence and caching rights. Model each against low, expected and high traffic, then apply the contract's unit definitions.
Illustrative example: invented to show structure; it does not describe an available dataset.
forecast: "one licensed grounding source, monthly"
traffic:
user_queries: 2_000_000
share_routed_to_source: 0.15 # queries where this source is eligible
ratios_per_eligible_query:
retrievals: 8 # chunks from this source in top-k after rerank
displays: 1.2 # citations rendered to the user
inferences: 2.5 # model calls in the agent loop
crawl:
documents_in_scope: 400_000
refresh_fraction_per_month: 0.10 # only changed docs refetched (ETag / Last-Modified)
billable_events_per_month:
retrievals: 2_400_000 # 2,000,000 x 0.15 x 8
displays: 360_000
inferences: 750_000
fetches: 40_000
scenarios: [low: 0.5x, expected: 1x, high: 3x]
contract_unit: "display" # only one unit billed; others reported
The same traffic yields 2.4 million retrievals but 360,000 displays in this sketch, so the choice of unit can move the bill by a factor of more than six at an identical per-event rate. That is why the unit definition matters more than the headline rate.
Choosing between flat fee, usage-based and hybrid structures
Choose flat fees when usage is high and predictable, usage-based fees when volume is uncertain or small, and hybrids for most production deployments. A flat subscription caps your exposure but makes you pay for headroom you may not use. Pure usage pricing aligns cost with value but transfers traffic risk to your finance team. The general trade-offs across structures are compared in AI data license pricing structures.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Your situation | Structure to propose | Protection to add |
|---|---|---|
| Internal assistant, known headcount | Flat annual fee per seat band | Re-tier at renewal on audited seats |
| New customer-facing AI search | Usage-based on displays | Monthly cap and rate step-downs by volume |
| Agent product with unpredictable loops | Per display or per answer, not per inference | Definition that collapses repeat calls in one session |
| Licensor wants revenue certainty | Minimum guarantee plus usage | Unused minimum credits against next period |
| Bulk index refresh needed | Per crawl with a refresh allowance | Free refetch for unchanged content (304 responses) |
A minimum guarantee is the usual price of predictability for the licensor. Internal and customer-facing deployments carry different risk and often different pricing; see internal vs customer-facing RAG licensing.
Contract terms that keep usage fees predictable
Predictability comes from the unit definition, caps, true-ups and the metering clause, not from the rate. Ask for these terms in writing, alongside the wider clauses in RAG content license terms.
- Billable event definition. Name the event (for example, "a display of a snippet or link from the licensed content to an end user"), and state what does not count: cache hits, retries, evaluation traffic, internal QA, and chunks retrieved but dropped by the reranker.
- Deduplication window. Multiple retrievals of the same document in one session or one answer count once.
- Caps and overage. A monthly or annual cap, with overage at the contract rate rather than list, and notice before the cap is reached.
- Minimum guarantee and true-up. Quarterly reconciliation against metered usage, with any credit carried forward and a defined true-up date.
- Volume tiers. Rate step-downs at stated volumes so growth reduces unit cost.
- Metering source of truth. Whose logs govern, what fields are reported (query_id, doc_id, event type, timestamp), audit rights and dispute windows. Detailed reporting design is in usage reporting and metering for licensed RAG content.
- Caching rights. Fetch fees depend on whether you may cache and for how long; see caching and retention limits.
- Price change and exit. Notice period for rate changes and the ability to terminate if a cap is repeatedly breached.
Reading public deal values without overpaying
Treat publicly reported AI licensing figures as press estimates, not price lists. Most deal values are undisclosed or estimated by reporters, and many bundle training, grounding and display rights into one headline number [4]. A reported annual figure for a large publisher tells you little about the per-display rate for a niche technical corpus.
Better anchors are your own unit economics: the revenue or cost saving per answer that uses the source, measured retrieval lift from a pilot, and the cost of the alternative (a grounding API, direct license or licensed crawling). If a usage rate exceeds the value of the answers it supports at expected traffic, restructure toward a capped hybrid or narrow the licensed scope using a coverage-gap analysis.
Usage pricing for enterprise operational content
Operational content from companies, such as support histories, product manuals or engineering records, is usually licensed as a defined dataset rather than metered per crawl. SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases; each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. SourceX does not publish prices; terms are agreed per deal. You can describe the grounding data you need to SourceX.
Price RAG content against your own usage
SourceX sources datasets on request from US companies for AI teams wherever based, and every release is approved by the supplying company; a request does not guarantee a match. Pricing and allowed uses are agreed in a license, and nothing is contracted until a supplier agrees. Tell SourceX what grounding content you need.
Frequently asked questions
Is pay-per-inference the same as per-retrieval pricing?
No. Per-retrieval counts chunks placed into context; pay-per-inference counts model responses that use the content [3]. An agent that calls the model several times per turn can multiply inference events without retrieving more content, so define which one you pay for.
Does a grounding license let us train on the content?
Usually not. Grounding and training are commonly licensed separately [6]; see grounding license vs training license for what each permits.
Should evaluation and test traffic be billable?
Negotiate to exclude it. Offline evaluation, regression suites and red-teaming can generate retrieval volume comparable to production, so carve out a tagged evaluation environment with its own reporting.
Sources
- Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?". https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
- Wikipedia, "Really Simple Licensing". https://en.wikipedia.org/wiki/Really_Simple_Licensing
- RSL Collective, "RSL Standard press release" (2025). https://rslstandard.org/press/rsl-standard
- LLM Pulse, "Every AI Content Licensing Deal, Mapped (2023-2026)" (2026). https://llmpulse.ai/blog/ai-content-licensing-deals/
- Supertab, "What AI monetization means for retrieval-augmented systems". https://supertab.co/blog/what-ai-monetization-means-for-retrieval-augmented-systems
- Newstex, "Editorial content licensing for AI training and grounding". https://www.newstex.com/blog/editorial-content-licensing-for-ai-training-and-grounding
- AdExchanger, "AI search adoption is boosting Benzinga's data licensing biz". https://www.adexchanger.com/publishers/ai-search-adoption-is-boosting-benzingas-data-licensing-biz
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.