Retrieval, RAG and grounding data
Synthetic queries vs real user queries for retriever training
Quick answer
Synthetic queries are usually enough to bootstrap a retriever, embedding model or reranker when your corpus is private and labeled pairs are scarce: an LLM writes queries for each passage, a filter discards weak ones, and you train on the survivors [1]. Real user queries earn their licensing cost when production phrasing, intent mix and failure cases differ from what a generator writes, and above all as the held-out evaluation set that tells you whether synthetic training transferred at all [3][5].
By SourceX Editorial · Updated
How synthetic query generation for retrievers actually works
Synthetic query generation inverts the retrieval problem: start from a passage, ask an LLM for a query that the passage answers, and treat the pair as a positive. NVIDIA's Nemotron reranker fine-tuning recipe, for example, generates synthetic training data for domains where labeled query-document pairs are scarce [1]. Published academic recipes such as InPars and Promptagator follow the same shape and differ mainly in prompting and filtering.
Prompting is usually few-shot: a handful of example query-passage pairs, then the target passage. The examples set the style of every generated query, so exemplars drawn from MS MARCO produce MS MARCO-like questions, while exemplars drawn from your own users produce something closer to your traffic. This is the cheapest place to spend a small real-query budget.
Filtering decides which pairs survive. Common filters keep queries with high generation likelihood, keep queries for which a retriever trained on the synthetic data ranks the source passage in its top results (round-trip consistency), or keep pairs that an existing cross-encoder reranker scores highly. Each filter has a bias: likelihood favors generic queries, and round-trip and reranker filters favor queries that are easy for models similar to the one you are training.
The practical consequence is that the generator model is rarely the bottleneck. Prompt design, filtering, hard-negative mining and, most of all, how closely the generated queries match the queries your system will actually receive decide the outcome. For negatives, see hard negatives for retriever training.
Where generated queries diverge from real user queries
Generated queries diverge from real ones in predictable ways, because the generator sees the answer before it writes the question. The result is query-document lexical overlap that real users rarely produce, which flatters BM25-style matching and teaches dense models a shortcut. Real users type fragments, product codes, misspellings, internal jargon and multi-intent requests, and they often ask questions the corpus cannot answer.
Typical gaps a buyer should test for:
- Lexical leakage. Generated queries reuse the passage's rare terms, so recall looks high in training and drops on paraphrased production queries.
- Answerability bias. Every synthetic query has a positive by construction. Real logs contain unanswerable and out-of-scope queries, which matter for abstention and for coverage-gap analysis.
- Intent mix. Generators over-produce factoid "what is" questions. Production traffic in support, legal or engineering search skews toward navigational lookups, troubleshooting and comparisons.
- Length and form. Real queries include two-word keyword strings and pasted error messages; generated ones tend toward well-formed sentences.
- Temporal drift. Real queries reference new products, releases and policy versions that never appear in a static corpus snapshot, which pairs with superseded and near-duplicate document testing.
Benchmark designers already treat this as a measurable difference. WixQA, an enterprise RAG benchmark, ships a synthetic QA set derived from each knowledge-base article alongside sets built from real user questions, which lets you compare a system on both [3]. If your retriever scores well on the synthetic split and poorly on the real one, the generator, not the model, is the likely weak point.
Evidence on domain transfer and query distribution shift
Retrieval gains are domain-specific, so the useful question is not "synthetic or real" in the abstract but "does this training set match the distribution I will serve". BEIR was built to measure exactly this zero-shot transfer across heterogeneous domains, and synthetic query generation papers commonly report results on it for that reason [2]. A model that leads on one BEIR dataset can trail on another, because argument retrieval, scientific claim checking and biomedical search reward different query styles.
Fine-tuning results on public model hubs show the same pattern at small scale. Primary reranker research, such as the BGE C-Pack paper, describes models fine-tuned on diverse datasets like Natural Questions [5]; these results show in-domain tuning, not transfer, so measure out-of-domain performance on your own queries before relying on it. Production systems that can access real traffic often train on it directly; one domain-specific QA system described in the literature fine-tuned its retriever using real user search and click data [4]. For how click data is structured and licensed, see search query and click logs for retriever training.
The working hypothesis most teams adopt is simple: synthetic queries cover the corpus, real queries cover the users. You need corpus coverage to train and user coverage to know whether training worked.
Decision table: when synthetic queries are enough
Synthetic queries are enough when the corpus defines the task and user phrasing is close to document language; real queries justify the cost when user behavior is the hard part.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Situation | Synthetic queries alone | Add licensed real queries | Why |
|---|---|---|---|
| New private corpus, no traffic yet | Yes, for training | For evaluation once available | Nothing real exists to license internally; bootstrap, then measure |
| Technical docs where users paste error strings or part numbers | Weak | Yes | Generators rarely reproduce raw logs, stack traces or SKU formats |
| Support or helpdesk search with jargon and abbreviations | Partial | Yes | Customer phrasing differs from article phrasing |
| Embedding model for a general-purpose domain | Often sufficient | Small eval set | Public benchmarks plus synthetic pairs cover much of the space |
| Reranker trained by distillation from a larger model | Often sufficient | Eval set and hard cases | Labels come from the teacher; the query set still sets the distribution |
| Abstention or "no answer" behavior | No | Yes | Synthetic queries are answerable by construction |
| Regulated domain (health, finance) | Depends on corpus rights | Only if de-identified and licensed | Real queries carry personal data and need review before use |
For reranker-specific label formats, see reranker training data. For whether public sets can carry the evaluation load, check which public retrieval datasets allow commercial use first; many popular ones restrict commercial use.
A hybrid recipe: synthetic training, licensed real-query evaluation
The most cost-effective pattern is to generate training queries at scale and spend licensing budget on a smaller, carefully documented set of real queries used mainly for evaluation. That set also seeds the few-shot examples that Promptagator-style generation relies on, so a handful of real queries improves the synthetic set too.
A workable sequence:
- Snapshot the corpus and record its version, so query-document pairs stay valid.
- Generate synthetic queries with an open mid-sized model, using real queries as few-shot exemplars where you have them.
- Filter with round-trip consistency or a reranker score, and drop pairs with high n-gram overlap with the source passage.
- Mine hard negatives and check them for false negatives before training.
- Hold out real queries with relevance judgments (qrels) and never use them as exemplars or training data.
- Report both splits: recall@k and nDCG@10 on synthetic queries and on real queries, the way WixQA separates its sets [3]. A large gap is a signal to license more real queries, not to train longer.
Relevance labels for the real holdout can come from human assessors or LLM judges; the trade-offs are covered in LLM vs human relevance labels and buying relevance judgments. The broader principle of validating generated data against a licensed real sample is covered in using a real-data holdout to validate synthetic data and the guide to combining licensed and synthetic data.
What to specify when you license real queries
A real-query set is only useful if its provenance, distribution and privacy treatment are documented well enough that you can trust the evaluation numbers. Ask for a datasheet covering source system, time window, sampling method and preparation, in the spirit of Data Cards, which document upstream sources, collection methods, intended use and decisions that affect model performance [6]. Real queries often contain names, account numbers and ticket IDs, so require a recorded de-identification method; see de-identifying search query logs.
Illustrative example: invented to show structure; it does not describe an available dataset.
Request template: real-query set for retriever evaluation
- Use: evaluation holdout and few-shot exemplars for retriever, embedding and reranker training.
- Source system: internal search box, helpdesk ticket subject lines, or chat-assistant first turns (state which).
- Fields per record:
query_id,query_text(de-identified),timestamp_bucket(month),channel,locale,session_position,clicked_doc_idsorresolved_doc_idif available,result_count_zeroflag. - Corpus link: document IDs that match a dated corpus snapshot, or the snapshot itself under the same license.
- Sampling: stratified by channel and month; include zero-result and abandoned queries; state the deduplication rule.
- Size target: enough queries per intent stratum to report per-stratum recall with useful confidence intervals.
- Privacy: method used to remove or replace personal details, plus a reviewed sample.
- Rights: confirmation that the holder may license the queries and the corpus for model training and evaluation, and the permitted uses and term.
How SourceX fits a real-query request
SourceX sources operational datasets from US companies on request and manages the licensing process; it does not hold query logs in stock, and a request does not guarantee a match. Relevant material can include support and sales histories, engineering records and documents, which are common sources of real search and question traffic. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. Buyers describe the data they need on the SourceX buyers page, not the businesses that might hold it.
More retrieval guidance sits in the retrieval and RAG data hub and the wider AI data guides. For terminology, see the synthetic data glossary entry.
Request real queries for retriever evaluation
If synthetic training needs a real-query holdout, describe the queries, corpus and allowed uses you need. SourceX looks for US businesses that hold matching operational data and handles rights review and licensing; nothing is contracted until a supplier agrees. Start at sourcex.si/buyers.
Sources
- NVIDIA, "Nemotron reranker fine-tuning recipe (README)". https://docs.nvidia.com/nemotron/nightly/nemotron/rerank/README.html
- arXiv (2104.08663), "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
- arXiv (2505.08643), "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/html/2505.08643v1
- arXiv (2404.14760), "Retrieval Augmented Generation for Domain-specific Question Answering" (2024). https://arxiv.org/pdf/2404.14760
- arXiv (2309.07597), "C-Pack: Packaged Resources To Advance General Chinese Embedding" (2023). https://arxiv.org/pdf/2309.07597
- Google Research (FAccT 2022; arXiv 2204.01075), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.