Retrieval, RAG and grounding data
RAG content licensing and retrieval data: a buyer's guide
Quick answer
RAG content licensing is permission to copy third-party content into a retrieval index, turn it into chunks and embeddings, and show or summarize it in AI answers at query time. It is a separate grant from a training license, and neither one implies the other [1]. A RAG product usually needs three inputs with different rights and sources: grounding content, retrieval training data (queries, relevance labels, hard negatives) for embedders and rerankers, and evaluation data tied to a pinned corpus snapshot.
By SourceX Editorial · Updated
Three inputs a RAG system licenses, mapped
Grounding content, retrieval training data and evaluation data come from different holders, raise different rights questions and fail differently, so scope and license each separately. (For the basic mechanism, see the glossary entry on retrieval-augmented generation.)
| Input | What it is | Typical origin | Deciding rights question | Common failure | Go deeper |
|---|---|---|---|---|---|
| Published grounding content | News, reference and trade content | Publishers, collecting societies, content APIs | May you index, cache, embed, display and send it to a model vendor? | A training-only license used for live answers | Grounding license vs training license |
| Enterprise grounding content | Wikis, manuals, tickets, drive files | Operational systems inside companies | May the holder share the records, with personal data handled? | Toy corpora that hide permission and version problems | Multi-source enterprise search corpora |
| Retrieval training data | Query-document pairs, graded labels, hard negatives | Search logs, ticket-to-article links, annotation | Is commercial model training permitted? | Building on a non-commercial benchmark | Training data for domain embedding models |
| Retrieval evaluation data | Topics, qrels, answer nuggets, frozen corpus | Assessors, pooled runs, versioned snapshots | May you keep the snapshot after the term? | Judgments that no longer match the indexed corpus | Pinned corpus snapshots |
Grounding rights: name every operation your pipeline performs
A grounding license has to authorize each thing your retrieval pipeline does to the content, not just "use in AI": storing copies, transforming them, displaying excerpts and passing text to the model that writes the answer. A content licensing vendor notes that licenses can set rules for how long content may be retained and when it must be deleted when rights expire [1]. One commentator argues RAG suits licensing because each answer traces to the content behind it, and licensed RAG content feeds prompts rather than model weights [2].
| Pipeline step | Right involved | What to get in writing |
|---|---|---|
| Crawl or bulk ingest | Reproduction | Which content sets, refresh frequency, full text or metadata only |
| Chunking, embeddings, summaries | Adaptation | Vectors, chunk text and summaries as permitted derived artifacts (embedding and vector index rights) |
| Cache and index retention | Reproduction over time | Maximum cache age, re-fetch rules (caching and retention limits) |
| Answer display | Display | Verbatim limits, summary-only modes, attribution, link-back (excerpt display rights) |
| Prompting a hosted model | Sublicensing | Whether content may reach a third-party model API, on what retention terms (model API flow-down terms) |
| Term end | Deletion | Delete-by-document-ID across source store, vector index, payloads and caches (removing content from vector indexes) |
Training rights carry obligations that grounding-only use may not. In the EU, providers of general-purpose AI models must put in place a policy to comply with Union copyright law, including honoring rights reservations under Article 4(3) of Directive (EU) 2019/790, and publish a summary of training content [3]. Keep grounding-only content out of training pipelines unless the license grants training, so your training-content records stay accurate.
In the US, building a search tool on editorial content is not automatically fair use. On 29 September 2026 the US Court of Appeals for the Third Circuit affirmed that Westlaw headnotes were copyrightable and that ROSS Intelligence's use of Westlaw material to train a non-generative legal-research tool was not fair use [4]. The ruling concerns a non-generative tool and does not decide how courts will treat retrieval at answer time. The RAG content license terms guide covers the clauses to negotiate.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How grounding deals are priced and metered
Reported publisher grounding deals increasingly pay per use (per crawl, retrieval or display) rather than a one-time fee, so the license and your logs must agree on what counts as a use. In Digiday's August 2025 explainer, a lawyer said the per-usage element had become the largest part of deal fees; the piece notes "grounding", "RAG" and "content inference compute" are used almost interchangeably [5]. A News/Media Alliance deal with Bria was reported as non-exclusive, paying publishers by how often their content appears in enterprise AI outputs, with revenue split 50-50 under an attribution model [6].
Machine-readable terms are emerging next to negotiated contracts. The Really Simple Licensing (RSL) standard, launched in September 2025, lists free, attribution, subscription, pay-per-crawl and pay-per-inference license models [7]. IAB Tech Lab released its CoMP (Content Monetization Protocol) specification v1.0 for public comment on 10 March 2026, with comments due by 9 April 2026; as of October 2026 our sources do not confirm a final version [8].
Australia's Copyright Agency began consulting members in 2026 on licensing grounding, fine-tuning and training as separate activities [9]. TollBit describes its own RAG licenses as ranging from summarization use cases to full-article display (vendor self-description), so state the display scope you need when comparing offers [10].
Three buyer decisions drive cost: the metered unit (retrieved chunk, cited document or displayed excerpt), whether internal copilots and customer-facing answers are priced differently, and the usage report you owe. See usage-based pricing for RAG content, usage reporting and metering and internal vs customer-facing RAG licensing. SourceX does not publish prices for the operational records it sources; terms depend on scope, volume, history, rights and exclusivity and are agreed per deal in writing.
Where grounding corpora come from
Grounding content comes from four kinds of holders, each with its own rights chain, freshness profile and failure mode. Publisher and feed content is licensed through publisher deals, collecting societies or grounding APIs and licensed crawling; enterprise and customer content needs record-level rights review.
- Publisher and reference content. News archives, reference works and trade titles. Ask whether the licensor holds AI rights for every contributor and syndicated item (third-party content inside licensed corpora).
- Real-time feeds. One vendor guide notes that for news search, freshness is controlled via query parameters and native ranking [11]. Measure both against your own queries; see API and streaming feeds vs batch delivery.
- Enterprise operational knowledge. Confluence or SharePoint wikis, Google Drive and Box folders, Zendesk or Salesforce ticket histories, and product manuals and service documentation. These hold what makes enterprise RAG hard: access controls, superseded versions, near-duplicates and personal data. Public benchmarks rarely capture it: EnterpriseRAG-Bench is generated for a fictional company with deliberately injected drafts and near-duplicates, per its vendor-hosted page [12], and the W-RAC preprint's RAG-Multi-Corpus uses 236 documents from five fictional organizations [13].
- Customer content. Files your customers upload are governed by your agreements with them; no third-party content license widens those.
SourceX works on the enterprise row. It sources operational datasets from US companies on request, including support and sales histories, engineering records and documents; these are kinds of data it sources, not inventory under contract. Every dataset goes through rights review. Personal details such as names, emails and account numbers are removed or replaced before delivery, with the method recorded per dataset, though no de-identification method is perfect.
SourceX does not source scraped public web content, and published news or reference content is not among the data kinds it describes; use the publisher routes above for that row. If real company knowledge is your gap, describe the corpus you need to SourceX; its page on RAG evaluation datasets from real company documents shows the document types involved.
Retrieval training data: queries, labels and hard negatives
Training or fine-tuning an embedding model or reranker needs real queries paired with passages that answer them, hard negatives that look relevant but do not answer, and ideally graded labels; several widely used public sets do not allow commercial use. Microsoft states that MS MARCO datasets are for non-commercial research only and extend no license or IP rights [14]. The authors of ORCAS, its click-log companion, note that click logs are usually withheld because they can reveal personally or commercially sensitive information, so ORCAS was aggregated and filtered with a k-anonymity requirement [15].
Check license labels at the source. BEIR's component licenses range from CC BY-SA and GPL to CC BY-NC, copyright-reserved and custom agreements, and 4 of its 19 datasets report no license [16]. A Hugging Face mirror of MS MARCO lists it as available for research and commercial use, contradicting Microsoft's terms [17]. This is consistent with the Data Provenance Initiative's finding of license omission above 70% and error rates above 50% on popular hosting sites [18]. See public retrieval datasets that allow commercial use.
Operational systems produce relevance signals as a by-product: an agent linking a ticket to the article that resolved it, a site-search click, a merged duplicate question. These make strong positives (support tickets linked to knowledge articles, search query and click logs) once query text is de-identified. Hard negatives need screening because they sit close to the query, which makes them more likely to be unlabeled positives, or false negatives [19]; see hard negatives for retriever training and reranker training data.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"query_id": "q-00481",
"query_text": "invoices stopped syncing after we changed our fiscal year start",
"query_origin": "support_ticket_subject",
"query_deidentified": true,
"positives": [
{"doc_id": "kb-2291", "anchor": "sec-3", "doc_version": "v14", "grade": 2, "label_source": "agent_linked_article"}
],
"hard_negatives": [
{"doc_id": "kb-2290", "anchor": "sec-1", "doc_version": "v9", "screening": "cross_encoder_below_threshold"},
{"doc_id": "kb-1874", "anchor": "sec-2", "doc_version": "v3", "screening": "assessor_confirmed_not_relevant"}
],
"negative_mining": "bm25_top50_minus_positives",
"corpus_snapshot_id": "helpcenter-2026-09-30",
"acl_scope": "public_help_center",
"license_use": ["retriever_training", "reranker_training", "internal_evaluation"],
"split": "train",
"split_key": "query_cluster_id"
}
The snapshot ID and document version keep labels valid when articles change; license_use states what the record may be used for. Splitting by query cluster keeps paraphrases out of both train and test.
Evaluation data: pinned snapshots, qrels and answer nuggets
A retrieval or RAG test set is reproducible only when the questions, the relevance judgments and the exact corpus version ship together, and when the license lets you keep that snapshot for repeat testing. The WixQA authors state that end-to-end RAG benchmarks need not only question-answer pairs but the knowledge-base snapshot the answers came from [20]. NIST's TREC data pages say the document collection and the qrels (query relevance judgments) must match, and describe judgments as complete only in a limited sense: enough results were judged to assume most relevant documents were found [21].
Specify the judging design instead of accepting "labeled" at face value. The TREC 2022 NeuCLIR track pooled the top 25 documents from baseline runs and the top 50 from other runs, and mapped four assessor categories to grades of 3, 1 and 0 [22]. In a pilot that served as development data for the TREC 2025 RAGTIME track, some question-answer pairs had answers linked to documents [23].
Ask for nugget-to-document links, near-duplicate distractors and unanswerable questions; see buying relevance judgments, a private domain test collection, distractor documents and question-answer-citation triples. SourceX's knowledge retrieval evaluation page lists business records suited to such test sets.
Metadata to require before a corpus is indexed
Deletion, metering, attribution and access control all key off per-document metadata, so require it at delivery rather than reconstructing it after indexing. Use this list in the request and at acceptance:
- Stable document ID that survives re-delivery, plus version and effective or superseded dates
- Source system and original path or URL, for attribution and link-back
- Per-document rights fields: permitted uses, display limits, licensor, expiry date
- Access-control scope (public, team, role) carried from the source system
- Content hash for deduplication across refreshes and syndicated copies
- Structure preserved (HTML, Markdown or DOCX next to any PDF) so your chunker can split on headings and tables
- A record of how personal data was removed or replaced, and in which fields
- Machine-readable dataset description; Croissant, for example, is a schema.org-based JSON-LD vocabulary for files and record structure [24]
Metadata a licensed RAG corpus should ship with and delivery formats that survive chunking expand each item; the delivery hub covers transfer.
Mistakes that leave a RAG product unlicensed or unmeasurable
RAG data mistakes usually surface after launch, when a usage report, deletion request or score regression exposes a gap in the contract or dataset.
- Metering a unit your logs cannot count. A per-retrieval price cannot be reconciled if your logs record only final citations. Log retrieval events with document IDs before you sign.
- Keying labels to chunk IDs. Changing chunk size creates new IDs and silently invalidates positives and qrels. Key labels to document ID plus a stable section anchor.
- Tuning on the test snapshot. Adjusting chunking, prompts or rerankers against the evaluation corpus inflates scores. Freeze a held-out split by query.
- Testing without permissions. A corpus stripped of access controls hides cross-user leakage; see permission-aware RAG evaluation.
- De-identifying away the search terms. Replacing product names, error codes or account-type labels with generic placeholders can remove exactly what queries match on; see de-identifying a RAG corpus without breaking retrieval.
Reading next: the licensing hub, collective licenses from collecting societies, synthetic vs real queries, pilot-testing a source for retrieval lift and SourceX's enterprise data licensing explainer.
Need operational records for retrieval or grounding?
Describe the corpus or retrieval data you need on the SourceX buyers page: source systems (help center, ticketing, wiki, shared drives), history, volume, metadata fields and the uses to license, such as indexing, answer display or retriever training. SourceX looks for US companies that hold that data, checks the data and each supplier's licensing permissions, manages the license and coordinates delivery; nothing is contracted until a supplier agrees. Specify your retrieval dataset with SourceX.
Guides in this section
- Buying Relevance Judgments (Qrels) for Retrieval EvaluationHow to specify, commission or license graded relevance judgments (qrels): pooling depth, TREC format, corpus versioning, assessor QA and judgment budgets.
- Deleting Licensed Content From Vector Indexes at Term EndHow to purge licensed content from RAG systems at license end: source text, vectors, payloads, caches, compaction, verification and deletion records.
- Domain Embedding Model Training Data: Pairs and NegativesHow to source query-passage positive pairs and hard negatives for domain embedding fine-tuning: pair sources, positive rules, splits, licenses.
- Embedding and Vector Index Rights for Licensed ContentDo embeddings of licensed content need a license? The rights to embed, re-embed, index, retain and share vectors, plus when to buy pre-computed embeddings.
- Enterprise Search Benchmark Data: Wikis, Chat, TicketsHow to source a licensed, cross-application enterprise search corpus from a real company: systems, linkage, ACLs, qrels, rights review and license scope.
- Grounding License vs Training License: What Each PermitsGrounding licenses cover query-time retrieval and display; training licenses cover model weights. Compare scope, pricing, overlap and the clauses to add.
- Metadata for RAG Documents: Fields to Require From SuppliersPer-document and per-chunk metadata a licensed RAG corpus should ship with: stable IDs, versions, timestamps, rights, ACLs and citation anchors.
- Product Manuals and Service Docs as RAG CorporaHow to source product manuals and service documentation for RAG: versions, applicability, question pairing, formats and rights checks for support copilots.
- RAG Content License Terms: Clauses to NegotiateClause-by-clause checklist for RAG and grounding content licenses: index, cache, embed, display, attribution, takedown, reporting and exit rights.
- Reranker Training Data: Formats, Negatives and VolumesHow to build reranker training data: pointwise, pairwise and listwise formats, hard negatives, distillation labels, query-level splits and volume checks.
- Retrieval Datasets for Commercial Use: A License AuditMS MARCO, ORCAS and BEIR licenses checked for commercial retriever training: what the official terms say, where mirrors mislead, and when to license data.
- Search Query and Click Logs for Retriever TrainingHow to license real search query and click logs for commercial retriever and reranker training: fields to request, position bias, privacy and license terms
- Support Tickets Linked to KB Articles as Relevance LabelsHow to source support tickets with linked knowledge-base article IDs and a pinned KB snapshot to train and evaluate support retrieval and RAG assistants.
- Usage-Based RAG Content Pricing: Crawl, Retrieval, DisplayHow pay-per-use RAG content licenses are priced: per crawl, per retrieval, per display and per inference, with a fee model, caps and minimum guarantees.
- Attribution and Link-Back Terms in AI Content LicensesHow citation, attribution and link-back clauses work in RAG and grounding licenses, what to negotiate, and how to build source IDs that keep you compliant.
- Best Document Formats for RAG Chunking and CitationsWhich delivery formats survive RAG chunking: originals plus normalized Markdown or JSONL with headings, tables, page numbers and section paths kept intact.
- Caching and Retention Limits for Licensed RAG ContentHow long a RAG system may cache licensed content: storage rights, TTLs, pointer vs full-text indexes, re-fetch duties and audit logs to negotiate.
- Cross-Language Retrieval Datasets: Sourcing and JudgmentsHow to source cross-language retrieval data: query-document pairs, bilingual relevance judgments, multilingual hard negatives and RAG evaluation sets.
- Duplicate-Question and Answer-Reuse Data for RetrievalHow to source duplicate-question pairs and approved-answer reuse logs for RFP, security-questionnaire and internal Q&A retrieval, with labels and rights.
- Grounding API vs Content License vs Licensed CrawlingCompare grounding APIs, direct content licenses and licensed crawling for RAG and agents by rights, freshness, cost and control, with a decision table.
- Hard Negatives for Retriever Training: Mining and ControlHow to mine, verify and buy hard negatives for dense retriever training, control false negatives, and specify negative data you request from a supplier.
- How to Anonymize Search Query Logs Before LicensingWhat a search query log supplier should apply before licensing: PII scans, k-anonymity, frequency thresholds, session cuts, and the retrieval cost of each.
- Internal RAG vs Customer-Facing AI: Content License ScopeWhich content license covers an internal RAG copilot and which covers customer-facing AI answers: users, entities, outputs, display, metering and clauses.
- Licensed Content in Third-Party LLM APIs: RAG TermsCan licensed content go to hosted LLM, embedding and vector-database APIs in RAG? The sublicense, processor and retention terms buyers should negotiate.
- LLM vs Human Relevance Labels: Where Each BelongsWhen LLM relevance judgments can replace human assessors for retrieval data, where they fail, and how to calibrate automatic qrels before you trust them.
- Pilot-Testing a RAG Content Source for Retrieval LiftDesign a sample-based pilot that measures whether a candidate corpus improves retrieval and grounded answers on your own queries before you license it.
- Private Domain Retrieval Test Collections: How to Build OneBuild a domain-specific retrieval benchmark: freeze a corpus snapshot, source real-user topics, grade qrels and link RAG nuggets to supporting documents.
- Product Search Relevance Data: Queries, Catalogs, LabelsHow to license product search relevance data: real queries, click logs, query-time catalog snapshots and graded labels for ranking and evaluation.
- RAG Content Usage Reporting: Metering Licensed SourcesHow to meter licensed RAG content: an event taxonomy, a per-answer metering record, attribution rules, licensor usage reports, reconciliation and audits.
- Real Distractor Documents for RAG Retrieval TestingHow to source corpora with real version chains, drafts and near-duplicates so RAG tests measure stale and superseded retrieval, not synthetic noise.
- Snippet and Excerpt Display Rights for AI AnswersHow much of a licensed source an AI answer may show: display tiers, verbatim limits in measurable units, and the output controls that enforce them.
- Synthetic vs Real Queries for Retriever TrainingWhen LLM-generated queries are enough to train retrievers and embedding models, and when licensed real user queries are worth the cost and review effort.
Sources
- Newstex (content licensing vendor blog), "Editorial content licensing for AI training and grounding". https://www.newstex.com/blog/editorial-content-licensing-for-ai-training-and-grounding
- AI Copyright newsletter (Substack), "Has Axel Springer really \"set the template\" for licensing deals with AI companies?". https://aicopyright.substack.com/p/has-axel-springer-really-set-the
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?" (2025). https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
- Digiday, "News/Media Alliance signs AI licensing deal to unlock recurring RAG revenue for small and mid-sized publishers". https://digiday.com/media-buying/news-media-alliance-signs-ai-licensing-deal-to-unlock-recurring-rag-revenue-for-small-and-mid-sized-publishers/
- RSL Collective, "RSL Standard press release" (2025). https://rslstandard.org/press/rsl-standard
- IAB Tech Lab (via PR Newswire), "IAB Tech Lab Announces CoMP Framework to Ensure LLMs Have Commercial Agreements with Publishers Before Content Crawling" (2026). https://tools.prnewswire.com/en-us/live/20823/release/20260310EN06615
- Copyright Agency (Australia), "Copyright Agency member consultations on AI licensing" (2026). https://www.copyright.com.au/2026/03/copyright-agency-member-consultations-on-ai-licensing/
- TollBit (vendor page), "Bot & agent paywall: licensed RAG". https://tollbit.com/licensed-rag/
- Firecrawl (vendor blog), "Best news API (buyer's guide)". https://firecrawl.dev/blog/best-news-api
- Onyx (vendor-hosted project page), "EnterpriseRAG-Bench". https://onyx.app/enterpriserag-bench
- arXiv:2604.04936 (preprint), "Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems" (2026). https://arxiv.org/pdf/2604.04936
- Microsoft (MS MARCO), "MS MARCO datasets". https://microsoft.github.io/msmarco/Datasets.html
- Craswell et al. (arXiv:2006.05324), "ORCAS: 18 Million Clicked Query-Document Pairs for Analyzing Search" (2020). https://arxiv.org/pdf/2006.05324
- Thakur et al. (arXiv:2104.08663), "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
- Hugging Face (ContextualAI/msmarco mirror), "msmarco / README.md". https://huggingface.co/datasets/ContextualAI/msmarco/blob/main/README.md
- Longpre et al. (arXiv:2310.16787), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Wang et al., AAAI 2024 (ML Anthology), "Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence Regularization" (2024). https://mlanthology.org/aaai/2024/wang2024aaai-mitigating
- arXiv:2505.08643, "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/html/2505.08643v1
- NIST, Text REtrieval Conference (TREC), "Data - English Relevance Judgements". https://trec.nist.gov/data/reljudge_eng.html
- Lawrie et al. (arXiv:2304.12367), "Overview of the TREC 2022 NeuCLIR Track" (2023). https://arxiv.org/pdf/2304.12367
- arXiv:2602.10024, "Overview of the TREC 2025 RAGTIME Track" (2026). https://arxiv.org/pdf/2602.10024
- Akhtar et al., MLCommons Croissant working group (arXiv:2403.19546), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.