Retrieval, RAG and grounding data
Training data for domain-specific embedding models
Quick answer
Training data for a domain-specific embedding model is a set of query-passage positive pairs drawn from real interactions in the target domain, plus hard negatives that are close to the query but do not answer it. The strongest positives come from natural signals such as search clicks, support tickets resolved by a linked article, or analyst questions answered by a cited document. Buyers should specify the pair source, the positive rule, deduplication, query-level splits and a license that permits commercial model training.
By SourceX Editorial · Updated
What an embedding training record contains
An embedding training record is a query, at least one passage judged to answer it, and optionally one or more passages that look relevant but are not. Contrastive objectives such as in-batch softmax (often called multiple-negatives ranking loss) pull the query toward its positive and away from every other passage in the batch, so a wrong positive is actively reinforced rather than merely ignored. If you are new to the vector side, the embeddings glossary entry covers the basics; this page is about the data that shapes them.
Three record shapes cover most purchases:
- Pairs (query, positive): the minimum for in-batch negative training.
- Triplets (query, positive, hard negative): adds a mined or labeled near-miss per query.
- Groups (query, positives[], negatives[], graded labels): needed if the same data will also train a cross-encoder. Group formats overlap with reranker training data, so negotiate one license that covers both models.
Where natural positives come from in a domain
Natural query-document interactions are the most reliable positives because a real user, not a generator, decided the passage answered the question. One domain question-answering study trained its retriever on query-document click pairs from product help content [1], and earlier work argued that domain click logs can supply the richness and scale neural retrievers need [2]. The practical question is which operational systems in your target domain already record that a query was satisfied by a document.
| Pair source | Positive signal | Typical failure mode |
|---|---|---|
| Site or help-center search logs | Click with dwell, no reformulation | Position bias: top results get clicks regardless of quality |
| Support tickets with linked KB article | Agent attached article and ticket closed | Agents link a "closest" article to meet a macro, not the true answer |
| Internal wiki or Confluence search | Click plus copy or share event | Few queries per page; long tail is thin |
| Engineering issues linked to runbooks or commits | Issue resolved referencing the doc | Link points to a whole repo, not a passage |
| Contract or policy Q&A in legal ops | Answer cites clause ID | Citation granularity at document, not clause, level |
For tickets in particular, see support tickets linked to knowledge articles, and for click-based pairs see search query and click logs for retriever training.
Why public QA pairs are a baseline, not a substitute
Public QA pair datasets are useful for warm-starting and ablations, but they rarely match the vocabulary, document structure or licensing a commercial domain model needs. WebFAQ 2.0 is a large multilingual FAQ-style collection built specifically for dense retrieval training, with mined hard negatives included [3]. ORCAS offers 18 million clicked query-document connections across 10 million distinct queries and 1.4 million TREC Deep Learning documents, but it is released under non-commercial research terms [6].
Two checks decide whether a public set belongs in your mix. First, the license: an audit of more than 1,800 text datasets found license information omitted in over 70% of cases on popular hosting sites and error rates above 50% [7], so confirm terms at the original publisher, not the mirror. Second, domain match: a reranker fine-tuned on Natural Questions reports gains that do not carry out of domain [5], and the same logic applies to bi-encoders. The page on which public retrieval datasets allow commercial use tracks the licensing side.
Defining a positive before you buy
A positive definition is a written rule that says exactly which interaction counts as "this passage answers this query", and it should be agreed before any data moves. Without it, two deliveries from the same supplier can use different thresholds and silently shift your training distribution. Write the rule against the supplier's actual fields, for example click_dwell_seconds >= 30 AND next_query_within_60s = false, or ticket.status = 'solved' AND kb_link.added_by_role = 'agent'.
Specify passage granularity at the same time. Embedding models train on chunks, so the supplier should deliver either pre-chunked passages with stable passage_id values or full documents with character offsets that point at the answering span. Ask how chunk boundaries were chosen (heading-aware, fixed token windows, overlap) and require the same chunker for the corpus you will index at inference.
Hard negatives and the false-negative problem
Hard negatives sharpen an embedding model, but mined negatives frequently contain passages that do answer the query, and training on those teaches the model to push correct answers away. Recent work describes this false-negative problem and relabels mined negatives with open-source LLMs that check whether the candidate actually contains the answer [4]. In domain corpora the risk is higher, because policies, manuals and KB articles exist in near-duplicate versions.
Practical controls to require or apply yourself:
- Exclude negatives that share a canonical document ID, version family or near-duplicate cluster (MinHash or SimHash) with the positive.
- Mine negatives with a different model from the one you are training, and skip the top few ranks where false negatives concentrate.
- Have a domain expert or LLM judge audit a stratified sample, and record the false-negative rate per source.
- Keep superseded versions and drafts as a separate, labeled pool; distractor and near-duplicate documents are valuable for testing but dangerous as unchecked negatives.
Real queries, synthetic queries and how to mix them
Real user queries carry the misspellings, internal codes and terse phrasing your production traffic will contain, while synthetic queries generated from passages give coverage of documents nobody has searched for yet. Most domain programs combine both, using real pairs as the anchor and for evaluation, and synthetic pairs to fill long-tail coverage. Tag every record with query_origin so you can ablate the mix; the trade-offs are covered in synthetic vs real queries for retriever training.
Never evaluate on synthetic queries produced by the same generator used for training. Hold out a set of real queries with expert judgments, ideally structured as a private domain retrieval test collection.
Deduplication, splits and leakage
Splits for embedding data must be made at the query level and, where possible, at the document-family level, or the evaluation numbers will overstate generalization. Click and ticket logs repeat the same intent with small wording changes ("reset 2fa", "reset two factor"), so normalize and cluster queries before splitting. Assign each cluster wholly to train, validation or test.
Apply the same discipline to passages. If a KB article appears in train as a positive and its next revision appears in test, the model is being tested on memorized text. Ask the supplier for document_family_id and version fields so you can enforce this.
Spec sheet for an embedding pair request
The request below is a checklist you can send to any supplier of domain pairs; it fits on one page and makes deliveries comparable.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | What to specify |
|---|---|
| Domain and corpus | e.g., B2B SaaS support KB plus resolved tickets, English, 2022 to 2026 |
| Pair source | Search clicks, ticket-to-article links, or both, with system names |
| Positive rule | Exact field-level rule and passage granularity |
| Negatives | None, mined (method and model), or labeled; false-negative audit rate |
| Volume | Target unique queries and unique passages, not just pair count |
| Dedup | Query normalization, near-duplicate passage clustering method |
| Splits | Query-cluster and document-family level, with held-out real-query test set |
| Personal data | Names, emails, phones and account numbers removed or replaced; method recorded |
| Metadata | Croissant JSON-LD or equivalent describing files, fields and record counts [8] |
| License | Commercial model training, embedding storage, term, deletion on termination |
Illustrative example: invented to show structure; it does not describe an available dataset.
{"query_id": "q_018442", "query": "sso login loops back to sign-in page idp", "query_origin": "search_log",
"positive": {"passage_id": "kb_2231#p4", "document_family_id": "kb_2231", "version": "2025-03", "text": "If users are redirected to sign-in after SAML assertion..."},
"positive_signal": "click_dwell_42s_no_reformulation",
"hard_negatives": [{"passage_id": "kb_1907#p2", "mined_by": "bm25", "fn_audit": "not_relevant"}],
"split": "train", "pii_method": "ner_replace_v3"}
Licensing terms specific to embedding models
An embedding license has to cover two things that a generic fine-tuning license may not: training the encoder and storing vectors derived from the licensed passages. If the passages will also be indexed for retrieval, the vector index itself may be a derivative the licensor cares about; see embedding and vector index rights for licensed content. If you only need weights adapted to a domain, a narrower fine-tuning-only data license may be enough.
Query logs are the highest-risk input because users type personal data into search boxes. Before licensing any log, read de-identifying search query logs and confirm how names, emails, account numbers and free-text identifiers were handled. For LLM rather than encoder adaptation, the owner page on domain-specific fine-tuning data is the better starting point.
How SourceX approaches embedding pair requests
SourceX sources operational datasets, including support histories, engineering records and documents, from US companies and manages the licensing process; it does not train models and does not source scraped web content. Buyers describe the pairs they need, and SourceX looks for US businesses that hold that data. Every dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs only after an executed agreement and supplier approval. Data is sourced on request rather than held in stock, so a request does not guarantee a match; you can describe your embedding data need to SourceX to start.
Source licensed query-passage pairs for your embedding model
If your domain encoder needs real query-passage positives from support, engineering or document workflows, describe the domain, pair source and positive rule you need. SourceX looks for US companies that hold matching data, reviews rights, and handles the license and ongoing purchases, with every release approved by the supplying company. Start at sourcex.si/buyers.
For the wider context, return to the retrieval and RAG data buyer's guide or the AI data hub.
Sources
- arXiv, "Retrieval Augmented Generation for Domain-specific Question Answering" (2024). https://arxiv.org/pdf/2404.14760
- arXiv (Rekabsaz et al.), "TripClick: The Log Files of a Large Health Web Search Engine" (2021). https://arxiv.org/abs/2103.07901v2
- arXiv, "WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval" (2026). https://arxiv.org/pdf/2602.17327
- arXiv, "ARHN: Answer-Centric Relabeling of Hard Negatives with Open-Source LLMs for Dense Retrieval" (2026). https://arxiv.org/pdf/2604.11092
- Hugging Face, "bge-reranker-base fine-tuned on Natural Questions (train split)". https://huggingface.co/sfczaa/bge-reranker-base-nq-ft
- Microsoft (MS MARCO), "ORCAS: Open Resource for Click Analysis in Search". https://microsoft.github.io/msmarco/ORCAS.html
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.