Skip to content

Retrieval, RAG and grounding data

Pilot-testing a content source for retrieval lift before you license

Quick answer

To evaluate a content source for RAG before licensing, index a representative sample alongside your current corpus and run a fixed set of your own production queries against two arms: baseline index, and baseline plus sample. Compare recall@k, nDCG@10 and grounded-answer correctness per query, split by queries the sample should cover and queries it should not. License only if the lift on covered queries is real, survives extrapolation to the full corpus, and does not degrade everything else.

By SourceX Editorial · Updated

Why a vendor demo cannot tell you the lift

A demo shows the source performing on queries someone else chose, so it measures the content's best case rather than its contribution to your system. Retrieval quality is strongly domain-dependent: the BEIR benchmark found that models tuned on one collection often lose ground on others, and that plain BM25 remained a hard baseline across heterogeneous domains [1]. A reranker fine-tuned on a single collection, such as one trained on Natural Questions, documents its gains on that collection, which is not evidence of gains on your queries [2].

Buyer comparisons of news APIs for LLM grounding make a similar point in market terms: freshness, coverage and enrichment differ by provider and are worth checking on your own query patterns before you commit [3]. The question a pilot answers is narrow and commercial: does adding this corpus to your existing index, with your retriever, chunker and generator, change outcomes on the questions your users actually ask? For the general mechanics of running a supplier pilot, see how to run a data pilot with a supplier; this page covers the retrieval-specific design.

Build the query set before you see the sample

The query set is the instrument, so freeze it before any candidate content arrives. Pull 300 to 1,000 real queries from search logs, chat transcripts or support tickets, deduplicate near-identical phrasings, and strip any personal data the logs contain. Then label each query into one of three strata.

  • Target queries: questions the candidate source claims to answer, ideally drawn from your coverage-gap analysis of unanswered queries.
  • Control queries: questions your current corpus already answers well. These detect displacement, where new documents push correct existing passages out of the top k.
  • Unanswerable queries: questions neither corpus should answer, to check that the new content does not induce confident wrong answers. See testing abstention with unanswerable questions.

For each target and control query, record a reference answer and the gold passages that support it. Ragas-style test records use exactly this shape: question, retrieved contexts, generated answer and ground truth [4]. If you lack in-house judges, buying relevance judgments (qrels) is a separate procurement decision worth making first.

Design the arms so the only variable is the content

A clean pilot holds the retriever, embedding model, chunking rules, reranker, prompt and generator constant across arms. Change one thing, the corpus, and the difference in metrics is attributable to the content. Run at least two arms, and add a third when you are comparing vendors in a bake-off.

Illustrative example: invented to show structure; it does not describe an available dataset.

ArmIndex contentsWhat it isolates
A: baselineCurrent production corpus, pinned snapshotYour starting point
B: baseline + sampleCurrent corpus plus candidate sample, same chunkerNet lift, including displacement
C: sample onlyCandidate sample aloneRaw coverage of target queries, without competition
B2: baseline + vendor 2 sampleCurrent corpus plus second candidateHead-to-head bake-off on identical queries

Pin the baseline with a dated snapshot so reruns are reproducible; pinned corpus snapshots for RAG evaluation explains why drifting indexes invalidate comparisons. Process the sample with exactly the same parser as production, because a vendor's pre-chunked Markdown will retrieve differently from the PDFs and HTML you would actually ingest. If the production corpus contains superseded drafts and near-duplicates, the pilot index should too; realistic distractor documents keep the test honest.

The retrieval lift metrics that matter

Retrieval lift is the per-query change in a ranking metric between arm B and arm A, averaged within each stratum. Report it per stratum, never as a single blended number, because a large gain on target queries can hide a regression on controls.

  • Recall@k at the k your generator actually consumes (often 5 to 20): did a gold passage reach the context window?
  • nDCG@10: does the gold passage rank near the top? This is the standard cross-domain comparison metric in BEIR [1].
  • MRR: the rank of the first relevant hit, useful when the prompt favors the top passage.
  • Source share: the fraction of top-k slots filled by the new corpus, per stratum. High share on control queries with flat or falling recall signals displacement.
  • Grounded-answer correctness and citation precision: whether the final answer matches the reference and cites a passage that supports it.

Measure answers as well as rankings, because better retrieval does not translate one-to-one into better answers. Models use evidence less reliably when it lands in the middle of a long context [5], so a source that adds many marginally relevant passages can raise recall while lowering answer quality. Nugget-based scoring, used in the TREC 2025 RAGTIME track to judge whether generated output covers key facts [6], is a practical way to grade answers on target queries.

Treat lift as a paired comparison: compute the difference for each query, then use a paired bootstrap or permutation test to put an interval around the mean. With a few hundred queries per stratum, a two-point nDCG gain can easily sit inside the noise.

A worked readout from a two-arm pilot

A readout should let a procurement committee make a license decision without rerunning anything. The example below shows the shape: per-stratum metrics, a decision rule written before the pilot, and a verdict.

Illustrative example: invented to show structure; it does not describe an available dataset.

pilot: field-service-manuals-sample
frozen_query_set: q-2026-09-v3   # 420 queries
sample: 4,800 documents, random draw across product lines (supplier-attested method)
arms: [A_baseline, B_baseline_plus_sample]
results:
  target (n=180):   recall@10 0.41 -> 0.63 | nDCG@10 0.33 -> 0.52 | answer_correct 38% -> 57%
  control (n=190):  recall@10 0.78 -> 0.76 | source_share 0% -> 22%
  unanswerable (n=50): false_answer_rate 12% -> 14%
paired_bootstrap_95ci:
  target nDCG@10 delta: [+0.14, +0.24]
  control recall@10 delta: [-0.05, +0.01]
pre_registered_rule: "license if target nDCG delta CI excludes 0 AND control recall loss < 0.03"
verdict: proceed to terms; add product-line filter to cut control displacement

Note what the readout does not claim: that the full corpus will perform identically. That requires the extrapolation step below.

Make sure the sample predicts the full corpus

A sample only predicts the licensed corpus if it was drawn the way the corpus will be delivered. Ask the supplier how the sample was selected: a random or stratified draw across the dimensions that matter (product line, document type, date range) is informative; a hand-picked "best of" set is not. Request the full corpus's distribution of those same dimensions so you can reweight sample results.

Check four failure modes before extrapolating. Overlap: deduplicate the sample against your existing index, since documents you already hold produce false lift. Freshness: confirm the full corpus's update cadence matches the sample's dates. Metadata: verify the fields you filter on (version, effective date, product, language) exist at full scale, using a RAG corpus metadata checklist. Quality measures from ISO/IEC 5259-2, such as completeness and consistency, give a shared vocabulary for writing these checks into acceptance criteria [7].

Coverage lift scales roughly with the share of target queries the full corpus answers, not with document count. If the sample covers 4 of 12 product lines, expect target lift on the other 8 only after you test a slice of them.

Paper the pilot before content touches your index

Pilot content gets embedded, cached and sometimes sent to third-party model APIs, so the evaluation terms need to say what is allowed and how it ends. Ask the supplier for written evaluation terms that cover the following.

  • Permitted use limited to internal evaluation, with no production serving or end-user display.
  • Whether sending sample text to an external embedding or generation API is allowed; see sublicense and processor terms for third-party model APIs.
  • Deletion at pilot end covering raw files, chunks, embeddings, caches and logs, with a vector index deletion procedure you can certify.
  • Whether pilot results may be shared internally and whether the sample is representative of the deliverable, stated in writing.
  • How personal data in the sample was removed and how that was checked.

If the pilot succeeds, the same metrics become the acceptance test in the full license, and the clause list in RAG content license terms picks up from there.

Where SourceX fits in a retrieval pilot

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and documents, and manages the commercial process through licensing and ongoing purchases. Nothing is held in stock and a request does not guarantee a match; buyers describe the content they need, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows after an executed agreement. You can describe the corpus you want to test to SourceX; nothing is contracted until a supplier agrees, so plan the pilot design above around what the supplying company approves. For the wider buying picture, start at the RAG content licensing guide or the AI data hub.

Request content to pilot for retrieval lift

If your coverage-gap analysis points to operational content held by US businesses, describe the documents, fields and query types you need. SourceX looks for companies that hold that data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Start a buyer request.

Sources

  1. Thakur et al., arXiv, "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
  2. Hugging Face, "bge-reranker-base-nq-ft model card". https://huggingface.co/sfczaa/bge-reranker-base-nq-ft
  3. Firecrawl, "Best news API". https://firecrawl.dev/blog/best-news-api
  4. Ragas documentation, "Prepare your test dataset". https://docs.ragas.io/en/v0.1.21/getstarted/prepare_data.html
  5. Liu et al., arXiv, "Lost in the Middle: How Language Models Use Long Contexts" (2023). https://arxiv.org/pdf/2307.03172
  6. NIST TREC, arXiv, "Overview of the TREC 2025 RAGTIME Track" (2026). https://arxiv.org/pdf/2602.10024
  7. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and machine learning, Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data