Skip to content

Retrieval, RAG and grounding data

Which public retrieval datasets allow commercial use?

Quick answer

Fewer than most teams assume. MS MARCO and ORCAS, the default training sources for dense retrievers and rerankers, are offered for non-commercial research only, with no license granted [1][2]. BEIR bundles 19 datasets (18 zero-shot evaluation sets plus MS MARCO) with mixed terms, including non-commercial, share-alike, GPL and unreported licenses [4]. Some components, such as CC BY 4.0 sets, permit commercial reuse with attribution. Audit each component at its official source, separate training from evaluation use, and license proprietary data where public terms fall short.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

MS MARCO and ORCAS: research only, with no license granted

MS MARCO is not licensed for commercial use. Microsoft's dataset page says the MS MARCO datasets are intended for non-commercial research purposes only and are provided free of charge without extending any license or other intellectual property rights [1]. That language covers the passage and document corpora, the training queries, the qrels.train.tsv files and the triples commonly used for contrastive training. The same page notes the original passage ranking task has been retired, which does not change the terms on existing copies [1].

ORCAS, the query-to-document click log released alongside the MS MARCO document corpus, sits under the same program and terms [2]. Teams use it for query augmentation and weak supervision, so it often enters a training mix quietly. If your retriever, embedding model or cross-encoder reranker will ship in a commercial product, treat both as research-only inputs unless you obtain separate written rights.

The practical answer to "can I train a commercial embedding model on MS MARCO?" is that the official terms do not grant that right. Many published checkpoints were fine-tuned on MS MARCO, so the question also applies to any base model you inherit; check the model card's training-data section, not just the weights license.

Mirror metadata does not override official terms

The official publisher's terms control, not a mirror's license tag. A widely used Hugging Face mirror of MS MARCO lists research and commercial use in its card metadata, which contradicts Microsoft's own statement [3][1]. On the Hub, the dataset card is the repository's README.md, and its YAML license: field drives search filters [9], so a wrong tag propagates into every filtered search.

This is a known, systemic problem. The Data Provenance Initiative found that licenses on aggregator platforms such as GitHub and Hugging Face are frequently unspecified or miscategorized relative to the original source [8]. Benjamin et al. reached a similar conclusion: publicly available datasets often carry ad hoc terms, and commercial use requires tracing the license chain back through each upstream source [7].

BEIR is a bundle, so audit it component by component

BEIR has no single license; each of its datasets keeps its own terms. The BEIR paper's appendix lists the licenses it could identify: ArguAna and Touché-2020 under CC BY 4.0, some sets under CC BY-SA, one under GPL, SciFact under CC BY-NC, several copyrighted or custom, and 4 of 19 with no reported license [4]. The project wiki points to each dataset's own source for downloads [5]. Those license notes date from 2021, so recheck each component at its origin before relying on them.

Three distinct risks hide in that list. Non-commercial (NC) terms bar commercial use outright. Share-alike (SA) and GPL terms raise questions about what derived artifacts must carry the same license, and counsel should decide whether a trained model or an index is a derivative. "Unknown" is not "permissive"; it means you have no grant at all.

BEIR's MS MARCO component inherits Microsoft's research-only terms [1][4]. A team that "only uses BEIR" for zero-shot evaluation can still end up training on it if hard-negative mining or distillation scripts pull in-domain splits.

Qrels, topics and corpus can carry different terms

Within one benchmark, the relevance judgments and the documents they point to can sit under different terms. TREC Deep Learning topics and qrels are described as public domain US Government work, while the MS MARCO corpus they judge remains non-commercial [6]. Public-domain qrels are of limited use without lawful access to the passages they reference.

Dataset publishers also may not own the underlying documents. Web-derived passages, scientific abstracts, forum posts and news articles keep their original copyright regardless of how the benchmark is labeled [7]. Record three license layers for every dataset: the compilation or release terms, the annotation layer (queries, qrels, answers), and the underlying content.

Training use and evaluation use carry different risk profiles

Training and evaluation are different uses, and the license may treat them differently. Training uses the content to update model weights and, for retrievers, often persists a vector index of the training corpus; evaluation typically runs a model over a test set and reports metrics. A non-commercial term can still bar internal evaluation that supports a commercial product, so do not assume evaluation is automatically exempt.

Statutory exceptions rarely close the gap. The UK text and data analysis exception in CDPA section 29A covers computational analysis for non-commercial research only [10]. In the EU, providers of general-purpose AI models must maintain a policy to comply with Union copyright law under AI Act Article 53(1)(c), applicable [11]. A dataset register like the one below is the evidence that policy relies on. For a broader treatment of evaluation terms, see checking eval dataset licenses for commercial use.

A component-level license register for retriever data

Keep one row per dataset component per layer, with the official source URL and the use you intend. The register below shows the structure; statuses for real datasets must come from the official pages cited above, rechecked at the time of use.

Illustrative example: invented to show structure; it does not describe an available dataset.

componentlayerofficial terms sourcelicense as statedmirror tagintended usedecision
example-passage-corpus-v2corpus textpublisher site, terms pagenon-commercial researchlicense: mit (conflicts)hard-negative miningexclude from training
example-passage-corpus-v2train qrelspublisher sitesame as corpusnonecontrastive pairsexclude from training
example-trec-style-topicstopics + qrelstrack overview paperpublic domain (US Gov)noneoffline evalallowed; corpus access still needed
example-argument-setfull releasedataset paper appendixCC BY 4.0cc-by-4.0eval and trainingallowed with attribution notice
example-sci-claimsfull releasedataset paper appendixCC BY-NCcc-by-nc-4.0eval onlycounsel review
example-qa-forum-dumpfull releasenot reportedunknownothernoneexclude; no grant

Useful fields to add: retrieval date, file hash of the copy you hold, the script or pipeline that consumes it, and the checkpoints trained from it. That last column is what lets you answer a removal or audit request later. Pair the register with the broader AI training data due diligence checklist.

Common audit failures in retriever pipelines

Most license problems in retrieval stacks come from indirect ingestion, not from a deliberate choice. Typical failure modes:

  • Inherited checkpoints. A base embedding model fine-tuned on MS MARCO triples carries that history into your fine-tune.
  • Synthetic query generation. Generating queries for a research-only corpus does not change the corpus terms.
  • Hard-negative caches. Mined negatives from BM25 runs over a restricted corpus usually store passages copied from that corpus.
  • Distillation targets. Cross-encoder scores computed on restricted passages still depend on those passages.
  • Mirror substitution. A data loader pointed at a mirror ID instead of the official release inherits the mirror's wrong tag [3][8].

When licensed alternatives make sense

License proprietary data when the public options are research-only, unknown, or do not match your domain. Enterprise retrievers rarely benefit from web-search click logs alone; support tickets with resolved answers, internal search logs, product documentation and engineering records are closer to production queries. Related guides cover search query and click logs for retriever training, buying relevance judgments and getting commercial rights for research-only datasets.

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and documents, and manages the licensing process. Nothing is held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. You can describe the retrieval data you need without naming specific companies. For the wider trade-offs, see licensed vs synthetic vs scraped training data and the retrieval and RAG data hub.

Get licensed retrieval training data for commercial use

SourceX looks for US businesses that hold the data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Personal details such as names, emails and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked. Start a buyer request at sourcex.si/buyers.

Frequently asked questions

Can I use MS MARCO only to evaluate a commercial retriever?

The official terms describe the datasets as intended for non-commercial research and grant no license [1]. Whether internal evaluation for a commercial product fits that description is a question for counsel. Many teams build a private domain retrieval test collection to remove the dependency.

Does a CC BY 4.0 component of BEIR need anything beyond attribution?

CC BY 4.0 permits commercial reuse with attribution, but the license covers only what the releaser had rights to grant [4][7]. Confirm the underlying documents were licensable, keep the attribution notice with the data, and record it in your register.

If a Hugging Face card says "commercial use," can I rely on it?

No. The card's license field is metadata entered by the uploader [9], and audits show it is often missing or wrong [8]. The MS MARCO mirror case shows a direct conflict with the publisher's terms [3][1].

Sources

  1. Microsoft, "MS MARCO Datasets". https://microsoft.github.io/msmarco/Datasets.html
  2. Microsoft, "ORCAS: Open Resource for Click Analysis in Search". https://microsoft.github.io/msmarco/ORCAS.html
  3. ContextualAI on Hugging Face, "msmarco / README.md (Hugging Face mirror)". https://huggingface.co/datasets/ContextualAI/msmarco/blob/main/README.md
  4. Thakur et al., arXiv, "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
  5. beir-cellar on GitHub, "Datasets available (BEIR wiki)". https://github.com/beir-cellar/beir/wiki/Datasets-available
  6. arXiv, "AcuRank: Uncertainty-Aware Adaptive Computation for Listwise Reranking" (2025). https://arxiv.org/pdf/2505.18512
  7. Benjamin et al., arXiv, "Can I use this publicly available dataset to build commercial AI software? Most likely not" (2021). https://arxiv.org/abs/2111.02374v4
  8. Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/pdf/2310.16787.pdf
  9. Hugging Face, "Dataset Cards". https://huggingface.co/docs/hub/en/datasets-cards
  10. UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 (consolidated), section 29A". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  11. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data