Retrieval, RAG and grounding data
Duplicate-question and answer-reuse data for retrieval
Quick answer
Duplicate-question retrieval data pairs a new question with earlier questions that mean the same thing, and the strongest commercial version also records which approved answer a person reused. Public sets such as Quora and CQADupStack are useful baselines but carry mixed or unreported licenses [1]. For RFP, security-questionnaire and internal Q&A assistants, the better source is an answer library's own reuse log: question variants, the canonical answer chosen, edits made, and who approved it.
By SourceX Editorial · Updated
What duplicate-question data has to contain for answer reuse
A usable dataset links each incoming question to one canonical question and one approved answer version, not just to a "similar" label. Classic duplicate-question benchmarks give binary pairs (is A a duplicate of B), which train a similarity model but say nothing about which answer was safe to reuse. Answer-reuse assistants need three hops: incoming question, matched library entry, and the answer text as it stood on the day it was reused.
That distinction matters because answer libraries change. A SOC 2 control answer approved in 2024 may be superseded after a new audit period, and a model trained on the old pairing will recommend stale text with confidence. Ask for answer version IDs and effective dates, so your retriever can learn "same question, current answer" rather than "same question, any answer."
Duplicate-question data is also different from ticket-to-article links. A support ticket mapped to a knowledge base article is a query-to-document label; owners of that content type are covered on the knowledge base articles page. Question-to-question labels sit closer to the paraphrase problem and need their own evaluation set.
Public duplicate-question and FAQ baselines and their license limits
Public sets are good for warm-starting a bi-encoder, but most were not built for commercial redistribution and none reflect enterprise questionnaire language. BEIR packages two duplicate-question tasks, CQADupStack (StackExchange subforums) and Quora question pairs, and its license appendix lists CQADupStack under Apache 2.0 while Quora's license is not reported [1]. Several other BEIR datasets restrict commercial use, so "it is in BEIR" is not a license [1].
FAQ-mined collections add scale. WebFAQ 2.0 builds multilingual question-answer pairs from FAQ markup on websites and mines hard negatives for dense retrieval training [2]. That makes it a reasonable baseline for FAQ matching, but web-mined content inherits the terms of the sites it came from, and its questions are written for website visitors rather than procurement analysts.
The broader record is not reassuring: the Data Provenance Initiative audited more than 1,800 text datasets and found license omissions above 70% and license errors above 50% on popular hosting sites [3]. Before a public set enters a commercial training run, trace it to its original source and terms. Our guide to public retrieval datasets that allow commercial use walks through that check.
Answer-library reuse logs as a natural positive label
The most valuable signal for answer recommendation is a logged reuse event, because a reviewer deliberately chose an existing answer for a new question. RFP and questionnaire tools typically store a content library of question-answer entries, alternate phrasings, tags such as product or framework, owners, and review dates. When a writer accepts a suggestion or searches and inserts an entry, many tools log that action. Whether a given system exports those logs, and with which fields, varies by product and must be confirmed with the holder.
Treat reuse events as strong but noisy positives. Writers sometimes insert the nearest available answer and then rewrite most of it, so capture the edit distance between the inserted and submitted text. A reuse with a 5% edit is a near-certain match; a reuse followed by a full rewrite is closer to a hard negative than a positive.
Rejected suggestions are equally useful. If the assistant showed three candidates and the writer picked the second, the first and third are in-domain hard negatives, which is the same shape that powers reranker training data. The logic mirrors search query and click logs for retriever training: implicit feedback, with position bias to correct.
Labeling, deduplication and evaluation splits
Duplicate-question data fails most often through leakage, because paraphrases of the same question land in both training and test splits. Questionnaires reuse standard frameworks (SIG, CAIQ and custom vendor-risk forms), so the same question can recur across many documents with trivial wording changes. Near-duplicate detection such as MinHash reduces this overlap in training corpora [4]; for question data, split by canonical question cluster, not by row.
Label quality needs its own budget. Even widely used benchmarks carry an estimated average label error rate of at least 3.3%, enough to reorder model rankings [5]. Have two reviewers adjudicate a sample of reuse-derived pairs, especially "near miss" pairs such as "Do you encrypt data at rest?" versus "Do you encrypt backups at rest?", which share most tokens but need different answers.
For evaluation, hold out a set of incoming questions with human-verified canonical matches and graded relevance, following the approach in buying relevance judgments (qrels). Report recall at k for the canonical question and, separately, whether the top answer version was current on the question's date.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
event_id | rv_000184 | Stable key for joins and deletion requests |
incoming_question | "Is customer data encrypted at rest, and with what key management?" | Query text as received |
source_doc_type | security_questionnaire | Separates RFP, questionnaire and internal Q&A |
matched_entry_id / entry_version | kb_0412 / v7 | Ties the match to the answer as it stood |
alt_phrasings_count | 14 | Signals how well-covered the cluster is |
action | inserted_suggestion_rank_2 | Implicit positive plus rank for bias correction |
shown_not_chosen | [kb_0398, kb_0577] | In-domain hard negatives |
edit_ratio | 0.06 | Low edit means a strong positive |
approver_role | security_lead | Shows the answer was reviewed |
event_date | 2025-03-11 | Needed for version-aware evaluation |
redaction_method | client names replaced with [CLIENT_n] | Documents de-identification |
Confidentiality and redaction in RFP and questionnaire content
Answer libraries hold confidential material, so redaction is a licensing condition rather than a cleanup step. Incoming questions often name the issuing client, the deal, and internal system names, while answers describe security architecture, subprocessors and pricing positions. Ask the holder to replace client names, project codenames and personal contact details with consistent placeholders, and to drop or generalize answers that disclose specific infrastructure.
Consistent placeholders matter for retrieval. If "Acme Bank" becomes [CLIENT_1] everywhere, cluster structure survives; if it is deleted, duplicate pairs can collapse into false matches. The trade-offs are covered in de-identifying a RAG corpus without breaking retrieval and de-identifying search query logs.
The submitted proposal text itself is a different asset; if you want full RFP responses for generation training, the owner page is RFP responses and proposals. Duplicate-question data needs only the question side, the match, and the approved answer, which usually reduces what the holder must clear.
Buyer checklist for a duplicate-question data request
A clear request names the label you need, the systems it should come from, and the rights you need for training versus evaluation. Use the checklist below when you describe the data to a supplier or broker.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Task: question-to-question retrieval, answer recommendation, or both; state whether you need graded or binary labels.
- Source systems: answer library or RFP tool, security-questionnaire portal, internal Q&A or helpdesk macros.
- Required fields: incoming question, canonical entry ID, answer version, reuse action, shown-but-not-chosen candidates, edit ratio, event date.
- Coverage: frameworks and domains (security, privacy, product, legal), time span, languages.
- Splits: cluster-level train and test split, held-out human-verified evaluation set.
- Redaction: client and person names replaced with consistent placeholders; method documented.
- Documentation: a dataset card describing source, collection, known biases and intended uses [6].
- Rights: confirmation that the holder owns the library content and may license it; allowed uses spelled out for training, evaluation and retrieval indexing, as discussed in grounding license vs training license.
Where SourceX fits for answer-reuse data
SourceX sources operational datasets from US companies on request and manages the commercial process, including the licensing agreement. Relevant kinds of data include support and sales histories, documents, and finance and legal workflows. Nothing is held in stock and a request does not guarantee a match; you describe the data you need, SourceX looks for US businesses that hold it, and every release is approved by the supplying company. You can describe your question-matching data needs to SourceX.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. For the wider retrieval picture, start at the retrieval and RAG data hub or the AI data hub.
Request duplicate-question and answer-reuse data
If public duplicate-question sets do not match your questionnaire domain or license needs, describe the reuse logs, fields and allowed uses you need. SourceX looks for US companies that hold that data, reviews ownership and consents, and agrees pricing and allowed uses in a license before anything is delivered. Start a buyer request.
Sources
- arXiv (Thakur et al.), "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
- arXiv, "WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval" (2026). https://arxiv.org/pdf/2602.17327
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- arXiv (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.