Skip to content

Fine-tuning and post-training data

Retrieval-augmented fine-tuning data: questions, documents, distractors and cited answers

Quick answer

A retrieval-augmented fine-tuning dataset pairs each question with the documents a retriever would return: one or more oracle documents that contain the answer, plus realistic distractors that do not. The target answer reasons over that context and quotes the supporting span. Some records deliberately omit the oracle, so the model learns to abstain instead of guessing. This is the pattern behind RAFT-style training [9], and its quality depends on distractor realism, citation accuracy and training rights to the underlying corpus.

By SourceX Editorial · Updated

What a RAFT-style training record contains

A usable record has five parts: a question, a context bundle, oracle labels, a cited target answer, and metadata that lets you audit and split the set. Think of it as an open-book exam: the model must learn to read the right pages and ignore the rest. The oracle is the document, or set of documents for multi-hop questions, from which the answer can be derived; distractors are retrieved alongside it but do not contain the full answer. A common target style reasons step by step and quotes the supporting oracle span verbatim, which makes the citation checkable.

Most teams store these as JSONL in the chat format their trainer expects, with the context bundle rendered into the user turn and the cited answer as the assistant turn. Keep the structured fields alongside the rendered prompt, because you will need them to re-render with a different number of distractors or a new citation style without re-annotating. For prompt and loss-masking conventions, see chat fine-tuning data format.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "raft-000412",
  "question": "What is the grace period before a late fee applies on net-30 invoices?",
  "context": [
    {"doc_id": "policy-ar-2024-07#s3", "role": "oracle", "text": "...Invoices on net-30 terms incur a late fee after a 10-day grace period..."},
    {"doc_id": "policy-ar-2022-01#s3", "role": "distractor", "distractor_type": "superseded_version"},
    {"doc_id": "policy-ap-2024-02#s5", "role": "distractor", "distractor_type": "same_topic_wrong_entity"},
    {"doc_id": "faq-billing-2023#q14", "role": "distractor", "distractor_type": "lexical_overlap"}
  ],
  "oracle_present": true,
  "context_order": "shuffled",
  "answer": "The current AR policy sets a 10-day grace period. ##begin_quote## Invoices on net-30 terms incur a late fee after a 10-day grace period ##end_quote## [policy-ar-2024-07#s3]. The 2022 policy is superseded.",
  "answer_type": "extractive_with_reasoning",
  "citations": [{"doc_id": "policy-ar-2024-07#s3", "char_start": 3, "char_end": 72}],
  "split_group": "policy-ar",
  "source_system": "document management export",
  "rights_ref": "license-schedule-A-row-7"
}

Why distractors decide whether the fine-tune helps

Distractors are the part of the dataset that teaches discrimination, so they must look like what your production retriever actually returns. Random passages from unrelated documents are easy to ignore and teach almost nothing. Realistic near-misses come from the same corpus and the same index you deploy, ideally mined by running your real retriever (BM25, a dense embedding model, or a hybrid with a reranker) and taking high-ranked chunks that annotators confirm do not answer the question.

Useful distractor types for enterprise corpora include:

  • Superseded versions: last year's policy, an older contract amendment, a deprecated runbook.
  • Same topic, wrong entity: the accounts-payable policy when the question is about accounts receivable; another customer's ticket with the same error code.
  • Lexical overlap: chunks that share rare terms with the question but answer a different question.
  • Partial evidence: a chunk that contains half of a multi-hop answer, which tests whether the model admits the gap.

Two failure modes recur. Unverified "hard negatives" are often actually relevant, which trains the model to ignore correct evidence; sample and adjudicate them. And if the oracle always appears in the same position, the model learns position instead of content, so shuffle the bundle and log the order.

Mixing records with and without the oracle

RAFT-style sets hold out the oracle from a share of training examples, so the model sees some bundles made only of distractors. Treat the oracle-present fraction as a hyperparameter to sweep against your own evaluation set rather than a fixed constant, and record it in the dataset card. Too few oracle-absent records and the model trusts whatever it is given; too many and it learns to refuse answerable questions.

What the target says when the oracle is missing is a design choice you must make explicitly. The OpenAI Cookbook RAG fine-tuning example trains on SQuAD samples where the answer is not in the context, with a target that says the model does not know [1]. The alternative keeps the reference answer on oracle-absent records, which pushes some domain knowledge into the weights but rewards answering without evidence. For enterprise assistants where an unsupported answer is worse than none, an abstention target is usually safer; the trade-offs are covered in fine-tuning data that reduces hallucinations.

Writing cited target answers that a grader can check

A cited answer is only useful as training signal if every citation resolves to a real span in the bundle and that span supports the claim. Store citations as document ID plus character offsets, not just a bracketed label, so an automated check can confirm the quoted text matches the source byte for byte. Explicit begin and end quote markers inside the reasoning make parsing reliable; whatever convention you choose, use the same one in training targets, the system prompt and your output parser.

Run these checks before training:

  • Quote fidelity: each quoted span is an exact substring of the cited chunk after whitespace normalization.
  • Citation coverage: every factual sentence in the answer carries at least one citation.
  • No distractor citations: answers never cite a chunk labeled distractor, unless the answer explains why it is superseded.
  • Answer-type balance: extractive, multi-hop, numeric and no-answer records are each represented.

Model-generated targets are common and fast, but they inherit the generator's citation habits. If a stronger model drafted the answers, have domain reviewers verify a stratified sample and record the acceptance rate; see due diligence for synthetic fine-tuning data.

Building questions from real enterprise documents

Questions should mirror what your users ask, which is why logged queries, support tickets and internal search histories are the best seeds. The Lenovo Press guide describes building a structured fine-tuning dataset from a company's own documents and notes that base-model responses can be verbose, hallucinate, or miss the desired tone, which is exactly what grounded targets correct [2]. Synthetic generation is a reasonable fill; CRAFT, for example, retrieves human-written documents and augments them into task samples [4]. Still, generated questions over-represent questions whose wording echoes the passage, which makes retrieval look easier than production.

Chunking must match deployment. If production splits on headings with a fixed token window and overlap, the training bundles should use the same chunker, embedding model and top-k, or the model learns to read context shapes it will never see. If you also intend to update the retriever, joint approaches such as RA-DIT fine-tune both the language model and the retriever [3], which means you also need query-to-relevant-chunk labels, not only answer targets.

Splits, leakage and evaluation

Split by document group, not by record, or the same oracle passage will appear in both training and test and inflate scores. Near-duplicate text is common in real corpora, and models trained on duplicated data reproduce training text verbatim [6], so run MinHash or embedding-based near-duplicate detection across policy versions, templated emails and boilerplate before assigning splits. Hold out whole document families (for example, every version of one policy) to measure generalization to unseen documents.

Keep the fine-tuning set and your RAG evaluation set separate in both content and authorship. Evaluation triples are a distinct asset with their own guide: question-answer-citation triples for RAG evaluation. Public enterprise benchmarks such as WixQA, built from a company knowledge base [5], are useful as a second, out-of-distribution check. Track at least answer accuracy, citation precision, abstention rate on oracle-absent items, and accuracy when the oracle is ranked low in the bundle.

Rights: training on a corpus is not the same as retrieving from it

Permission to index documents for retrieval does not automatically include permission to train on them. Content licensing practice treats training and grounding as separate categories, where a grounding agreement does not grant the right to train on a corpus [7]. RAFT records embed full document chunks in the training prompt and verbatim quotes in the target, so the corpus itself becomes training data, not just the questions.

Open datasets do not remove this question: an audit of more than 1,800 text datasets found license omission above 70% and license error rates above 50% on popular hosting sites [8]. Before purchase, confirm in writing that the license covers fine-tuning on document text, derivative question-answer generation, quoting in targets, and the model's eventual deployment scope. More on recording these terms is in the AI training data provenance guide.

Illustrative example: invented to show structure; it does not describe an available dataset.

Buyer checklist itemWhat to ask forWhy it matters for RAFT data
Corpus scopeDocument types, systems of origin, date range, version historySuperseded versions are your best distractors
Chunk-level IDsStable document and section IDsCitations must resolve after re-chunking
Training rightsLicense language covering fine-tuning on document textRetrieval-only rights do not cover chunks in training prompts
PII handlingMethod recorded, sample checkedQuotes copy text into targets verbatim
Query seedsReal user questions or tickets, with consent basisSynthetic questions skew toward easy retrieval
Version metadataEffective dates and supersession linksLets you label stale documents as distractors

How SourceX helps source document corpora for RAG-aware fine-tuning

SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and it manages the licensing agreement and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, which is where corpus training rights should be stated. Datasets are sourced on request, not held in stock, so describe the corpus you need on the SourceX buyer page.

Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked; no method is perfect, so run your own scan as described in scanning a training corpus for PII before fine-tuning. Related context: RAG evaluation datasets from real company documents, knowledge retrieval evaluation, the retrieval-augmented generation glossary entry, the supervised fine-tuning glossary entry and the fine-tuning and post-training datasets buyer's guide.

Request document corpora for retrieval-augmented fine-tuning

Describe the documents, systems and question types you need, and SourceX looks for US businesses that hold that data; every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Start a request on the SourceX buyer page.

Sources

  1. OpenAI Cookbook, "Fine-tuning for retrieval augmented generation (RAG) with Qdrant". https://cookbook.openai.com/examples/fine-tuned_qa/ft_retrieval_augmented_generation_qdrant
  2. Lenovo Press, "Making LLMs Work for Enterprise Part 2: RAG Fine-Tuning Dataset Creation". https://lenovopress.lenovo.com/lp1954-making-llms-work-for-enterprise-part-2-rag-fine-tuning-dataset-creation
  3. LlamaIndex, "Improving RAG effectiveness with Retrieval-Augmented Dual Instruction Tuning (RA-DIT)". https://llamaindex.ai/blog/improving-rag-effectiveness-with-retrieval-augmented-dual-instruction-tuning-ra-dit-01e73116655d
  4. arXiv, "CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation" (2024). https://arxiv.org/html/2409.02098v2
  5. arXiv, "WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation" (2025). https://arxiv.org/html/2505.08643v1
  6. arXiv (Lee et al.; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2022). https://arxiv.org/abs/2107.06499v1
  7. Newstex, "Editorial content licensing for AI training and grounding" (2024). https://www.newstex.com/blog/editorial-content-licensing-for-ai-training-and-grounding
  8. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  9. arXiv (Zhang et al.), "RAFT: Adapting Language Model to Domain Specific RAG" (2024). https://arxiv.org/abs/2403.10131

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data