Skip to content

Fine-tuning and post-training data

Fine-tuning data that reduces hallucinations: abstention and grounded answers

Quick answer

Fine-tuning reduces hallucinations when the training data teaches a model when not to answer, not just what to answer. Filter supervised targets to facts the base model already knows or that a supplied context states, add unanswerable, out-of-scope and insufficient-context prompts paired with correct abstentions, and build preference pairs that rank a grounded refusal above a fluent guess. Then measure over-refusal on answerable items so the model does not learn to decline everything.

By SourceX Editorial · Updated

Why ordinary SFT data can increase hallucination

Supervised fine-tuning on facts the base model does not already know can increase hallucinations, as such examples are learned slower and linearly increase the tendency to invent answers [9]. Filtering SFT data to facts the base model already knows, by rewriting targets toward supplied evidence or toward what the model can recall, reduced factual errors in adapted models [1]. The working explanation in post-training research is that models acquire most factual knowledge in pre-training, while fine-tuning mainly teaches them how to use it; treat that as a hypothesis to confirm on your own checkpoint.

The practical consequence for a data buyer is uncomfortable. A large, accurate, expert-written SFT set can still be a hallucination source if many targets assert facts outside the base model's knowledge and outside any context in the prompt. The standard pipeline of SFT on demonstrations followed by reinforcement learning from human rankings of outputs [6] does not correct this by itself; preference tuning tends to build on whatever answering policy the demonstrations encode.

Quality also beats volume here. LIMA fine-tuned a 65B model on 1,000 curated prompt-response pairs [7], which supports spending budget on a smaller set of carefully labeled answerable and unanswerable items rather than a larger set with unchecked targets.

Four example types every factuality mixture needs

A hallucination-reducing dataset combines grounded answers with three distinct kinds of abstention, because each failure mode needs its own training signal. Mixing them under a single "I don't know" label hides which behavior the model actually learned.

  • Grounded answerable. The answer is either stated in the provided context or is a fact the base model reliably answers closed-book. Targets cite or quote the supporting span where a context exists.
  • Unanswerable from context. The context is on-topic but does not contain the answer. SQuAD 2.0 showed that extractive readers guess in exactly this case and built adversarial unanswerable questions that look answerable [3]; your items should be equally plausible, not trivially off-topic.
  • Out of scope. The question falls outside the assistant's declared domain or tool access. The target names the boundary and, where appropriate, routes to a human or another system.
  • Insufficient information. The question is answerable in principle but needs a missing parameter (account, date range, jurisdiction, version). The target asks a specific clarifying question instead of assuming a default.

Research on training objectives points the same way: a CCN 2024 poster reportedly proposes an objective that rewards accurate answers and acknowledging uncertainty, which only works if the data contains cases where uncertainty is the right answer [2]. The core recipe is that the label depends on the model, not only on the question.

Probing the base model before you label

Answerability labels for closed-book items must be computed against the specific base checkpoint you will fine-tune, because "known" is a property of the model. A target that is correct but unknown to the model is a likely hallucination-training example under the reasoning in [1].

A workable probe runs each candidate question through the base model several times with sampling, scores answers against the gold answer with exact match or a normalized-alias match, and buckets items by how often the model is correct. Items the model never answers correctly are candidates for conversion to abstention targets or for moving into a retrieval-grounded format where the fact appears in the context. Re-run the probe whenever the base checkpoint changes; labels do not transfer across model families.

For context-grounded items, the probe is a different check: confirm the supporting span exists in the supplied documents and that a second annotator can locate it. Our page on retrieval-augmented fine-tuning data with distractors and cited answers covers document packing and distractor selection in more depth.

Where natural abstention examples come from

Operational business records contain abstention behavior written by people under real constraints, which makes it more varied than templated refusals. Support tickets escalated to tier 2, "need more info" replies that request an order number or log file, sales emails that defer a pricing question to a specialist, and engineering issues closed as "cannot reproduce, please attach version" are all natural insufficient-information and out-of-scope examples (a working hypothesis to test on your own evals, not a published result).

These records need conversion before use. Strip agent signatures and internal ticket macros, rewrite the agent's next action into a model-appropriate target, and keep the original resolution so you can verify that the clarifying question was actually the right one. Personal details in the thread must be removed or replaced before training; for health-related threads, HIPAA de-identification under Safe Harbor or Expert Determination applies [8].

Synthetic refusals are faster to make but tend to collapse into a few phrasings, which teaches style rather than judgment. If you buy synthetic abstention data, apply the checks in our guide to due diligence for purchased synthetic fine-tuning data.

Illustrative record schema for abstention SFT and preference data

Each record should carry its answerability label, the evidence for that label, and the probe result, so you can rebalance the mixture without relabeling.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "id": "abst-000417",
  "split": "train",
  "answerability": "insufficient_information",
  "domain": "b2b_saas_support",
  "messages": [
    {"role": "system", "content": "You answer billing questions using the provided account context only."},
    {"role": "user", "content": "Why was I charged twice this month?"}
  ],
  "context_docs": [{"doc_id": "kb-billing-12", "text": "Duplicate charges can occur when a plan change and a renewal fall on the same day..."}],
  "target": "I can check that, but I need the account email or the last four digits of the invoice number. Duplicate charges often come from a plan change on a renewal date, so knowing whether you changed plans this month will also help.",
  "evidence_span": null,
  "probe": {"base_model": "ckpt-2026-08", "samples": 8, "correct": 0},
  "rejected": "You were charged twice because your card was processed by two payment gateways. A refund of the second charge has been issued.",
  "rejected_failure": "fabricated_action_and_cause",
  "provenance": {"source_type": "support_ticket", "pii_method": "replace_with_placeholders", "reviewer": "annotator_b"}
}

The rejected field turns the same record into a preference pair. DPO fits the policy directly to such chosen-rejected pairs without a separate reward model [5], so the quality of the rejected answer matters: it should be the kind of fluent fabrication your current model actually produces, ideally sampled from it, not a strawman.

Balancing abstention against over-refusal

Abstention data reduces hallucination only up to the point where the model starts declining questions it could answer, so the mixture needs an explicit over-refusal budget. Too many refusal targets teach a cheap policy: decline whenever uncertain, which scores well on hallucination metrics and fails users.

Illustrative example: invented to show structure; it does not describe an available dataset.

Mixture leverStarting point to testSignal that it is too highSignal that it is too low
Share of abstention targets15 to 30 percent of factual itemsRefusal rate on probed-known items risesUnsupported claims on unanswerable eval stay flat
Split across unanswerable / out-of-scope / insufficient-infoRoughly equal, then weight toward your top production failureOne phrasing dominates outputsModel guesses defaults instead of asking
Preference pairs with refusal as chosenPaired with an equal number where a grounded answer beats a refusalModel refuses answerable context questionsModel still prefers fluent fabrications
Closed-book items with probe correct = 0Converted to abstention or moved to grounded formatNot applicableHallucination rate grows with training steps

The ratios are starting hypotheses for ablation, not published constants. Track three numbers per checkpoint: accuracy on answerable items, abstention rate on unanswerable items, and refusal rate on answerable items. A calibration curve that plots confidence against correctness on the same eval tells you whether refusals concentrate where the model is actually wrong.

Training objectives that pair with abstention data

Data composition does most of the work, but some objectives make grounded behavior easier to learn. Faithful Finetuning (F2) adds explicit faithfulness losses for question answering during fine-tuning and reports gains over vanilla fine-tuned models [4]. Preference optimization over chosen grounded answers versus rejected fabrications [5] is a lighter-weight option that uses the same records.

Whatever the objective, keep the evaluation set separate from training and include held-out unanswerable items written by different annotators, so the model cannot pass by memorizing refusal templates. Our page on faithfulness evaluation sets with answers labeled grounded or unsupported covers measurement; this page covers the training side.

Buyer checklist for hallucination-reducing fine-tuning data

Before licensing an abstention or grounded-QA dataset, confirm the supplier can answer these questions in writing.

  • Is every record labeled with one of the answerability classes, with the evidence span or reason recorded?
  • Were closed-book items probed against a named base model, and can you re-run the probe on yours?
  • Are unanswerable items adversarial (on-topic, plausible) rather than random off-topic questions [3]?
  • Do refusal targets vary in phrasing and include next steps, or are they one template?
  • Is there a matched set of answerable items to measure over-refusal?
  • Are rejected answers in preference pairs realistic model failures?
  • What is the source of each record, what consents and ownership back it, and what uses does the license allow?
  • How were personal details removed, and was a sample checked after removal?

For general acceptance testing, see how to evaluate a fine-tuning dataset before you buy it, and for the wider cluster, the fine-tuning and post-training datasets buyer's guide. Teams combining abstention with domain adaptation can also review training data for domain-specific fine-tuning and the RAG evaluation datasets from real company documents use case.

Where SourceX fits for abstention and grounded-answer data

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records and documents, which carry natural escalation and clarification behavior. Nothing is held in stock, and a request does not guarantee a match; buyers describe the data they need, and every release is approved by the supplying company. You can describe the records you need for factuality training to start the Find and Assess steps.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. SourceX does not train models and does not source scraped web content.

Sourcing data to teach a model to say "I don't know"

If your factuality work needs real escalations, clarification requests and grounded answers from business records, SourceX can look for US companies that hold them and manage the licensing through to delivery. Diligence materials on source, rights, preparation and allowed use are prepared per dataset, and terms are agreed per deal. Tell SourceX what grounded and abstention data you need.

Sources

  1. The Model Wire, "Knowledge-aligned fine-tuning reduces hallucinations in adapted models". https://themodelwire.com/article/knowledge-aligned-fine-tuning-reduces-hallucinations-in-adapted-models-01M1DGGK1HT9QK4HMVYPHGZCT9
  2. Cognitive Computational Neuroscience (CCN) 2024, "CCN 2024 poster on training objectives that acknowledge uncertainty" (2024). https://2024.ccneuro.org/poster/?id=473
  3. Rajpurkar, Jia and Liang (arXiv), "Know What You Don't Know: Unanswerable Questions for SQuAD" (2018). https://arxiv.org/abs/1806.03822
  4. arXiv, "Mitigating Large Language Model Hallucination with Faithful Finetuning" (2024). https://arxiv.org/html/2406.11267v1
  5. Rafailov et al., NeurIPS 2023 (arXiv), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  6. Ouyang et al., OpenAI (arXiv), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  7. Zhou et al., Meta AI, NeurIPS 2023 (arXiv), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  8. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  9. [Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?](https://arxiv.org/abs/2405.05904), arXiv (Gekhman et al.), 2024.

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data