Skip to content

Retrieval, RAG and grounding data

RAG content licensing and retrieval data: a buyer's guide

Quick answer

RAG content licensing is permission to copy third-party content into a retrieval index, turn it into chunks and embeddings, and show or summarize it in AI answers at query time. It is a separate grant from a training license, and neither one implies the other [1]. A RAG product usually needs three inputs with different rights and sources: grounding content, retrieval training data (queries, relevance labels, hard negatives) for embedders and rerankers, and evaluation data tied to a pinned corpus snapshot.

By SourceX Editorial · Updated

Three inputs a RAG system licenses, mapped

Grounding content, retrieval training data and evaluation data come from different holders, raise different rights questions and fail differently, so scope and license each separately. (For the basic mechanism, see the glossary entry on retrieval-augmented generation.)

InputWhat it isTypical originDeciding rights questionCommon failureGo deeper
Published grounding contentNews, reference and trade contentPublishers, collecting societies, content APIsMay you index, cache, embed, display and send it to a model vendor?A training-only license used for live answersGrounding license vs training license
Enterprise grounding contentWikis, manuals, tickets, drive filesOperational systems inside companiesMay the holder share the records, with personal data handled?Toy corpora that hide permission and version problemsMulti-source enterprise search corpora
Retrieval training dataQuery-document pairs, graded labels, hard negativesSearch logs, ticket-to-article links, annotationIs commercial model training permitted?Building on a non-commercial benchmarkTraining data for domain embedding models
Retrieval evaluation dataTopics, qrels, answer nuggets, frozen corpusAssessors, pooled runs, versioned snapshotsMay you keep the snapshot after the term?Judgments that no longer match the indexed corpusPinned corpus snapshots

Grounding rights: name every operation your pipeline performs

A grounding license has to authorize each thing your retrieval pipeline does to the content, not just "use in AI": storing copies, transforming them, displaying excerpts and passing text to the model that writes the answer. A content licensing vendor notes that licenses can set rules for how long content may be retained and when it must be deleted when rights expire [1]. One commentator argues RAG suits licensing because each answer traces to the content behind it, and licensed RAG content feeds prompts rather than model weights [2].

Pipeline stepRight involvedWhat to get in writing
Crawl or bulk ingestReproductionWhich content sets, refresh frequency, full text or metadata only
Chunking, embeddings, summariesAdaptationVectors, chunk text and summaries as permitted derived artifacts (embedding and vector index rights)
Cache and index retentionReproduction over timeMaximum cache age, re-fetch rules (caching and retention limits)
Answer displayDisplayVerbatim limits, summary-only modes, attribution, link-back (excerpt display rights)
Prompting a hosted modelSublicensingWhether content may reach a third-party model API, on what retention terms (model API flow-down terms)
Term endDeletionDelete-by-document-ID across source store, vector index, payloads and caches (removing content from vector indexes)

Training rights carry obligations that grounding-only use may not. In the EU, providers of general-purpose AI models must put in place a policy to comply with Union copyright law, including honoring rights reservations under Article 4(3) of Directive (EU) 2019/790, and publish a summary of training content [3]. Keep grounding-only content out of training pipelines unless the license grants training, so your training-content records stay accurate.

In the US, building a search tool on editorial content is not automatically fair use. On 29 September 2026 the US Court of Appeals for the Third Circuit affirmed that Westlaw headnotes were copyrightable and that ROSS Intelligence's use of Westlaw material to train a non-generative legal-research tool was not fair use [4]. The ruling concerns a non-generative tool and does not decide how courts will treat retrieval at answer time. The RAG content license terms guide covers the clauses to negotiate.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

How grounding deals are priced and metered

Reported publisher grounding deals increasingly pay per use (per crawl, retrieval or display) rather than a one-time fee, so the license and your logs must agree on what counts as a use. In Digiday's August 2025 explainer, a lawyer said the per-usage element had become the largest part of deal fees; the piece notes "grounding", "RAG" and "content inference compute" are used almost interchangeably [5]. A News/Media Alliance deal with Bria was reported as non-exclusive, paying publishers by how often their content appears in enterprise AI outputs, with revenue split 50-50 under an attribution model [6].

Machine-readable terms are emerging next to negotiated contracts. The Really Simple Licensing (RSL) standard, launched in September 2025, lists free, attribution, subscription, pay-per-crawl and pay-per-inference license models [7]. IAB Tech Lab released its CoMP (Content Monetization Protocol) specification v1.0 for public comment on 10 March 2026, with comments due by 9 April 2026; as of October 2026 our sources do not confirm a final version [8].

Australia's Copyright Agency began consulting members in 2026 on licensing grounding, fine-tuning and training as separate activities [9]. TollBit describes its own RAG licenses as ranging from summarization use cases to full-article display (vendor self-description), so state the display scope you need when comparing offers [10].

Three buyer decisions drive cost: the metered unit (retrieved chunk, cited document or displayed excerpt), whether internal copilots and customer-facing answers are priced differently, and the usage report you owe. See usage-based pricing for RAG content, usage reporting and metering and internal vs customer-facing RAG licensing. SourceX does not publish prices for the operational records it sources; terms depend on scope, volume, history, rights and exclusivity and are agreed per deal in writing.

Where grounding corpora come from

Grounding content comes from four kinds of holders, each with its own rights chain, freshness profile and failure mode. Publisher and feed content is licensed through publisher deals, collecting societies or grounding APIs and licensed crawling; enterprise and customer content needs record-level rights review.

  • Publisher and reference content. News archives, reference works and trade titles. Ask whether the licensor holds AI rights for every contributor and syndicated item (third-party content inside licensed corpora).
  • Real-time feeds. One vendor guide notes that for news search, freshness is controlled via query parameters and native ranking [11]. Measure both against your own queries; see API and streaming feeds vs batch delivery.
  • Enterprise operational knowledge. Confluence or SharePoint wikis, Google Drive and Box folders, Zendesk or Salesforce ticket histories, and product manuals and service documentation. These hold what makes enterprise RAG hard: access controls, superseded versions, near-duplicates and personal data. Public benchmarks rarely capture it: EnterpriseRAG-Bench is generated for a fictional company with deliberately injected drafts and near-duplicates, per its vendor-hosted page [12], and the W-RAC preprint's RAG-Multi-Corpus uses 236 documents from five fictional organizations [13].
  • Customer content. Files your customers upload are governed by your agreements with them; no third-party content license widens those.

SourceX works on the enterprise row. It sources operational datasets from US companies on request, including support and sales histories, engineering records and documents; these are kinds of data it sources, not inventory under contract. Every dataset goes through rights review. Personal details such as names, emails and account numbers are removed or replaced before delivery, with the method recorded per dataset, though no de-identification method is perfect.

SourceX does not source scraped public web content, and published news or reference content is not among the data kinds it describes; use the publisher routes above for that row. If real company knowledge is your gap, describe the corpus you need to SourceX; its page on RAG evaluation datasets from real company documents shows the document types involved.

Retrieval training data: queries, labels and hard negatives

Training or fine-tuning an embedding model or reranker needs real queries paired with passages that answer them, hard negatives that look relevant but do not answer, and ideally graded labels; several widely used public sets do not allow commercial use. Microsoft states that MS MARCO datasets are for non-commercial research only and extend no license or IP rights [14]. The authors of ORCAS, its click-log companion, note that click logs are usually withheld because they can reveal personally or commercially sensitive information, so ORCAS was aggregated and filtered with a k-anonymity requirement [15].

Check license labels at the source. BEIR's component licenses range from CC BY-SA and GPL to CC BY-NC, copyright-reserved and custom agreements, and 4 of its 19 datasets report no license [16]. A Hugging Face mirror of MS MARCO lists it as available for research and commercial use, contradicting Microsoft's terms [17]. This is consistent with the Data Provenance Initiative's finding of license omission above 70% and error rates above 50% on popular hosting sites [18]. See public retrieval datasets that allow commercial use.

Operational systems produce relevance signals as a by-product: an agent linking a ticket to the article that resolved it, a site-search click, a merged duplicate question. These make strong positives (support tickets linked to knowledge articles, search query and click logs) once query text is de-identified. Hard negatives need screening because they sit close to the query, which makes them more likely to be unlabeled positives, or false negatives [19]; see hard negatives for retriever training and reranker training data.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "query_id": "q-00481",
  "query_text": "invoices stopped syncing after we changed our fiscal year start",
  "query_origin": "support_ticket_subject",
  "query_deidentified": true,
  "positives": [
    {"doc_id": "kb-2291", "anchor": "sec-3", "doc_version": "v14", "grade": 2, "label_source": "agent_linked_article"}
  ],
  "hard_negatives": [
    {"doc_id": "kb-2290", "anchor": "sec-1", "doc_version": "v9", "screening": "cross_encoder_below_threshold"},
    {"doc_id": "kb-1874", "anchor": "sec-2", "doc_version": "v3", "screening": "assessor_confirmed_not_relevant"}
  ],
  "negative_mining": "bm25_top50_minus_positives",
  "corpus_snapshot_id": "helpcenter-2026-09-30",
  "acl_scope": "public_help_center",
  "license_use": ["retriever_training", "reranker_training", "internal_evaluation"],
  "split": "train",
  "split_key": "query_cluster_id"
}

The snapshot ID and document version keep labels valid when articles change; license_use states what the record may be used for. Splitting by query cluster keeps paraphrases out of both train and test.

Evaluation data: pinned snapshots, qrels and answer nuggets

A retrieval or RAG test set is reproducible only when the questions, the relevance judgments and the exact corpus version ship together, and when the license lets you keep that snapshot for repeat testing. The WixQA authors state that end-to-end RAG benchmarks need not only question-answer pairs but the knowledge-base snapshot the answers came from [20]. NIST's TREC data pages say the document collection and the qrels (query relevance judgments) must match, and describe judgments as complete only in a limited sense: enough results were judged to assume most relevant documents were found [21].

Specify the judging design instead of accepting "labeled" at face value. The TREC 2022 NeuCLIR track pooled the top 25 documents from baseline runs and the top 50 from other runs, and mapped four assessor categories to grades of 3, 1 and 0 [22]. In a pilot that served as development data for the TREC 2025 RAGTIME track, some question-answer pairs had answers linked to documents [23].

Ask for nugget-to-document links, near-duplicate distractors and unanswerable questions; see buying relevance judgments, a private domain test collection, distractor documents and question-answer-citation triples. SourceX's knowledge retrieval evaluation page lists business records suited to such test sets.

Metadata to require before a corpus is indexed

Deletion, metering, attribution and access control all key off per-document metadata, so require it at delivery rather than reconstructing it after indexing. Use this list in the request and at acceptance:

  • Stable document ID that survives re-delivery, plus version and effective or superseded dates
  • Source system and original path or URL, for attribution and link-back
  • Per-document rights fields: permitted uses, display limits, licensor, expiry date
  • Access-control scope (public, team, role) carried from the source system
  • Content hash for deduplication across refreshes and syndicated copies
  • Structure preserved (HTML, Markdown or DOCX next to any PDF) so your chunker can split on headings and tables
  • A record of how personal data was removed or replaced, and in which fields
  • Machine-readable dataset description; Croissant, for example, is a schema.org-based JSON-LD vocabulary for files and record structure [24]

Metadata a licensed RAG corpus should ship with and delivery formats that survive chunking expand each item; the delivery hub covers transfer.

Mistakes that leave a RAG product unlicensed or unmeasurable

RAG data mistakes usually surface after launch, when a usage report, deletion request or score regression exposes a gap in the contract or dataset.

  • Metering a unit your logs cannot count. A per-retrieval price cannot be reconciled if your logs record only final citations. Log retrieval events with document IDs before you sign.
  • Keying labels to chunk IDs. Changing chunk size creates new IDs and silently invalidates positives and qrels. Key labels to document ID plus a stable section anchor.
  • Tuning on the test snapshot. Adjusting chunking, prompts or rerankers against the evaluation corpus inflates scores. Freeze a held-out split by query.
  • Testing without permissions. A corpus stripped of access controls hides cross-user leakage; see permission-aware RAG evaluation.
  • De-identifying away the search terms. Replacing product names, error codes or account-type labels with generic placeholders can remove exactly what queries match on; see de-identifying a RAG corpus without breaking retrieval.

Reading next: the licensing hub, collective licenses from collecting societies, synthetic vs real queries, pilot-testing a source for retrieval lift and SourceX's enterprise data licensing explainer.

Need operational records for retrieval or grounding?

Describe the corpus or retrieval data you need on the SourceX buyers page: source systems (help center, ticketing, wiki, shared drives), history, volume, metadata fields and the uses to license, such as indexing, answer display or retriever training. SourceX looks for US companies that hold that data, checks the data and each supplier's licensing permissions, manages the license and coordinates delivery; nothing is contracted until a supplier agrees. Specify your retrieval dataset with SourceX.

Guides in this section

Sources

  1. Newstex (content licensing vendor blog), "Editorial content licensing for AI training and grounding". https://www.newstex.com/blog/editorial-content-licensing-for-ai-training-and-grounding
  2. AI Copyright newsletter (Substack), "Has Axel Springer really \"set the template\" for licensing deals with AI companies?". https://aicopyright.substack.com/p/has-axel-springer-really-set-the
  3. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  4. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  5. Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?" (2025). https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/
  6. Digiday, "News/Media Alliance signs AI licensing deal to unlock recurring RAG revenue for small and mid-sized publishers". https://digiday.com/media-buying/news-media-alliance-signs-ai-licensing-deal-to-unlock-recurring-rag-revenue-for-small-and-mid-sized-publishers/
  7. RSL Collective, "RSL Standard press release" (2025). https://rslstandard.org/press/rsl-standard
  8. IAB Tech Lab (via PR Newswire), "IAB Tech Lab Announces CoMP Framework to Ensure LLMs Have Commercial Agreements with Publishers Before Content Crawling" (2026). https://tools.prnewswire.com/en-us/live/20823/release/20260310EN06615
  9. Copyright Agency (Australia), "Copyright Agency member consultations on AI licensing" (2026). https://www.copyright.com.au/2026/03/copyright-agency-member-consultations-on-ai-licensing/
  10. TollBit (vendor page), "Bot & agent paywall: licensed RAG". https://tollbit.com/licensed-rag/
  11. Firecrawl (vendor blog), "Best news API (buyer's guide)". https://firecrawl.dev/blog/best-news-api
  12. Onyx (vendor-hosted project page), "EnterpriseRAG-Bench". https://onyx.app/enterpriserag-bench
  13. arXiv:2604.04936 (preprint), "Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems" (2026). https://arxiv.org/pdf/2604.04936
  14. Microsoft (MS MARCO), "MS MARCO datasets". https://microsoft.github.io/msmarco/Datasets.html
  15. Craswell et al. (arXiv:2006.05324), "ORCAS: 18 Million Clicked Query-Document Pairs for Analyzing Search" (2020). https://arxiv.org/pdf/2006.05324
  16. Thakur et al. (arXiv:2104.08663), "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
  17. Hugging Face (ContextualAI/msmarco mirror), "msmarco / README.md". https://huggingface.co/datasets/ContextualAI/msmarco/blob/main/README.md
  18. Longpre et al. (arXiv:2310.16787), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  19. Wang et al., AAAI 2024 (ML Anthology), "Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence Regularization" (2024). https://mlanthology.org/aaai/2024/wang2024aaai-mitigating
  20. arXiv:2505.08643, "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/html/2505.08643v1
  21. NIST, Text REtrieval Conference (TREC), "Data - English Relevance Judgements". https://trec.nist.gov/data/reljudge_eng.html
  22. Lawrie et al. (arXiv:2304.12367), "Overview of the TREC 2022 NeuCLIR Track" (2023). https://arxiv.org/pdf/2304.12367
  23. arXiv:2602.10024, "Overview of the TREC 2025 RAGTIME Track" (2026). https://arxiv.org/pdf/2602.10024
  24. Akhtar et al., MLCommons Croissant working group (arXiv:2403.19546), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data