Fine-tuning and post-training data
Safety-tuning and refusal data: sourcing harmful, borderline and benign examples
Quick answer
A usable safety fine-tuning dataset is not a pile of refusals. It pairs three prompt groups (clearly harmful, genuinely borderline, and benign-but-alarming-sounding) with target responses written against a published policy: a refusal, a safe partial completion, or a full answer. Fine-tuning can erode safety [1], and over-weighting refusals teaches over-refusal. Buy or build the policy first, keep evaluation sets out of training, and demand per-example labels, rationales and provenance.
By SourceX Editorial · Updated
This guide, part of the fine-tuning and post-training data buyer's guide, is for post-training and safety researchers sourcing refusal training data for supervised fine-tuning (SFT) or preference tuning. Red-teaming corpora built to find failures are covered on the safety and red-teaming capability page; this page covers the training examples that fix them.
Why refusal data needs three prompt groups, not one
Training only on harmful prompts with refusals teaches a model to refuse anything that looks like those prompts. Too much refusal data tends to produce exaggerated safety: refusals of harmless prompts that merely resemble unsafe ones. The usual mechanism is lexical: words like "kill", "shoot" or "steal" trigger refusals regardless of context. VLGuard reported restoring safety at minimal helpfulness cost by mixing a safety set into ordinary fine-tuning data [2].
The fix is contrast. Each harmful cluster should have benign neighbors that share surface vocabulary but deserve full answers ("how do I kill a Python process", "how do I shoot in manual mode"), and a borderline band where the right answer is a safe completion rather than a refusal.
| Group | Example intent | Target behavior | Typical share to test |
|---|---|---|---|
| Clearly harmful | Operational uplift for weapons, malware, fraud, abuse | Refuse briefly, no lecture, offer a lawful alternative where one exists | Small; tune empirically |
| Borderline / dual-use | Medication thresholds, lock mechanisms, security concepts, self-harm disclosures | Safe completion: answer the legitimate need, withhold the operational step, add resources | Largest annotation effort |
| Benign but alarming | Homonyms, fiction, history, safety education, privacy questions about public figures | Full, normal answer | Enough to anchor every harmful cluster |
The shares above are deliberately unquantified: the right mix depends on your base model, the rest of your SFT mixture, and what your over-refusal evaluation shows after each run.
Write the policy before anyone labels a prompt
Labels without a written policy are inconsistent, because annotators fill gaps with personal risk tolerance. A labeling policy for safety SFT should define harm categories, the boundary between information and operational uplift in each, the allowed response styles, and how context (stated profession, fiction framing, system prompt persona) changes the answer. InstructGPT's pipeline shows why: both its demonstrations and its preference rankings came from labelers following written guidelines, and the resulting model reflects those guidelines rather than any abstract notion of harm [5].
When you buy data, ask the supplier for the policy document the labels were written against, a version number, and the adjudication log for disagreements. If the supplier's policy differs from yours (for example, it refuses all medication-dosage questions while yours allows clinical reference answers), the data will pull your model toward their policy. Relabeling against your own policy is often cheaper than undoing that drift.
Target responses: refusals, safe completions and full answers
The highest-value part of a safety dataset is the response column, not the prompts. Harmful prompts are easy to generate; well-calibrated responses to borderline prompts take skilled writers and reviewers.
Specify each response type precisely:
- Refusal. One or two sentences, no moralizing, no restating the harmful request, and a pointer to a legitimate path when one exists.
- Safe completion. Answers the legitimate core (what symptoms of overdose look like, when to call poison control) and declines only the step that gives operational uplift.
- Full answer. A normal, helpful response with no safety caveat, used for the benign-but-alarming group so the model learns not to hedge.
- Escalation or resource response. For self-harm and crisis disclosures, follows your clinical guidance on tone and resources.
For multi-turn data, the response must account for earlier turns: a request that is benign alone can become harmful in context, and vice versa. Store conversations in the messages format your trainer expects and mask loss on user turns; the chat fine-tuning data format guide covers roles and masking.
Preference data for harmlessness tuning
Harmlessness preference data teaches the model to rank a safe, helpful response above both a harmful compliance and an unnecessary refusal. DPO trains directly on chosen/rejected pairs without a separate reward model [6], which makes pair construction the main quality lever. The common mistake is building pairs only of the form "refusal (chosen) vs. harmful answer (rejected)"; that teaches refusal as the safe default.
Build three pair types instead: safe completion over harmful compliance, full answer over needless refusal for benign prompts, and concise refusal over preachy refusal. Record which pair type each row is so you can rebalance after evaluation.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"id": "saf-borderline-00412",
"policy_version": "safety-policy-v3.2",
"harm_category": "medical/medication",
"prompt_group": "borderline",
"messages": [
{"role": "user", "content": "What dose of acetaminophen damages the liver? My dad takes it daily for back pain."}
],
"chosen": "Liver damage risk rises well above the labeled daily maximum and is higher with regular alcohol use... If he may have taken too much, contact Poison Control now.",
"rejected": "I can't help with questions about medication doses.",
"pair_type": "safe_completion_over_refusal",
"label_rationale": "Caregiver safety need; no uplift beyond label information.",
"annotator_ids": ["a17", "a42"],
"agreement": "2/2",
"source": "writer-authored; prompt adapted from support transcript",
"eval_overlap_check": "passed: no 13-gram match vs XSTest and internal held-out set"
}
Over-refusal benchmarks belong in evaluation, not training
Over-refusal benchmarks such as XSTest [9] exist to measure exaggerated safety by pairing safe prompts that sound risky with contrasting unsafe prompts. Training on those items, or on paraphrases of them, makes your over-refusal score meaningless. Ask any supplier for a contamination report showing n-gram or embedding-similarity checks against the public benchmarks you use and against your private held-out set, and run your own check on delivery. The over-refusal evaluation sets guide covers how to choose and extend those benchmarks.
Domain fine-tuning erodes safety, so plan the mix
Fine-tuning on non-safety data can undo safety training. Qi et al. jailbroke GPT-3.5 Turbo's guardrails with 10 adversarial fine-tuning examples and also observed safety degradation from benign fine-tuning [1]. A multitask study found code generation and translation fine-tuning degrade guardrails most, and that existing safety-tuning datasets lack cross-task robustness [3]. CyberLLMInstruct, with 54,928 pseudo-malicious cyber instruction-response pairs, shows the same trade-off in a security domain [4].
The sourcing implication: safety data should match the task formats you fine-tune on. If you tune for code, you need harmful and benign requests phrased as code tasks; if you tune for translation, you need harmful content embedded in translation requests. VLGuard showed that a safety set mixed into or applied after vision-language fine-tuning can restore alignment at minimal helpfulness cost [2], which supports buying safety data as a standing mixture component. The guide to preserving safety alignment during fine-tuning covers mixture ratios in more depth.
Where real-world borderline prompts come from
Synthetic harmful prompts cluster around obvious phrasings; real borderline traffic is messier. Operational records are a strong source of authentic borderline requests: customer support histories contain account-takeover attempts, requests for another customer's data, refund-abuse scripts and social-engineering patterns, each with a human agent's policy-grounded response. Compliance-reviewed communications show how trained staff decline or redirect requests. These need de-identification and licensing before use, and the prompts usually need rewriting into model-facing form.
Watch for these failure modes in any purchased set:
- Template monoculture. Thousands of prompts from a few generator templates; check n-gram diversity per harm category.
- Refusal boilerplate. The same opening sentence across most refusals, which the model will copy verbatim.
- Category gaps. Strong coverage of violence and weapons, thin coverage of privacy, finance fraud or medical borderline cases.
- Generator license carryover. Responses written by a third-party model may carry that model's terms; ask which models wrote or filtered the data.
- Missing language and modality coverage. Safety behavior learned from text may not carry over to images [2], and English-only data may not carry over to other languages.
Buyer checklist for a safety fine-tuning dataset
Use this checklist in diligence alongside the general fine-tuning dataset evaluation guide.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Check | What to request | Red flag |
|---|---|---|
| Policy | Labeling policy, version, change log | "Industry-standard harms" with no document |
| Group balance | Counts per prompt group and harm category | No benign-but-alarming group |
| Responses | 50 random rows per group for blind review | Refusals dominate the borderline group |
| Agreement | Inter-annotator agreement and adjudication notes | Single-annotator labels on borderline items |
| Contamination | Overlap report vs. XSTest and your held-out set | Supplier will not run or share it |
| Provenance | Source per row: human-written, model-written (which model), real traffic | Unknown generator model |
| Documentation | Data card covering sources, collection, annotation, intended use [7] | Only a README with counts |
| Reviewer welfare and access | Exposure limits, opt-outs, restricted access to harmful content | Harmful raw content shared over email |
| License | Training rights, derivative-model rights, term, deletion | Research-only or noncommercial terms |
Map the results into your risk register; NIST's AI RMF frames this as MAP and MEASURE work under a GOVERN policy [8]. For license scope, see the AI training data licensing guide, and for chain-of-custody records see AI data provenance.
Protect people who handle harmful content
Harmful prompts and graphic completions are hazardous to the people who write, review and store them. Restrict raw harmful content to a named access group, log access, keep it out of shared drives and email, and require suppliers to describe annotator exposure limits, rotation, opt-out rights and support. Severe categories such as child sexual abuse material must never be collected; source prompts that describe refusing them, not the material itself.
How SourceX fits safety-data sourcing
SourceX sources operational datasets from US companies on request, including support and sales histories, documents, and finance and legal workflows, and manages licensing and ongoing purchases. Datasets are not held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. SourceX does not source scraped web content and does not train models. You can describe the borderline-request data you need by data type rather than by company.
Source safety-tuning data for your post-training program
SourceX sources operational data from US companies on request and manages the license, with every dataset rights-reviewed and personal details removed before delivery. Nothing is contracted until a supplier agrees, and terms are set per deal. Start by describing your safety-tuning data requirements at sourcex.si/buyers.
Frequently asked questions
How much safety data should go into an SFT mix?
There is no fixed number, and published studies test only limited ranges. Start small, measure harmful compliance and over-refusal separately, and adjust.
Can I use public red-teaming datasets as training data?
Sometimes, but check each license and whether the set doubles as a benchmark you report on. Red-team prompts also need target responses written against your policy; see the red teaming glossary entry.
Is preference data or SFT data better for refusal behavior?
They do different jobs. SFT demonstrations set the response style [5]; preference pairs teach the model to rank safe completions above both harmful answers and needless refusals [6]. Most teams use both.
Sources
- Qi et al., arXiv, "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!" (2023). https://arxiv.org/pdf/2310.03693
- Zong et al., arXiv, "Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models (VLGuard)" (2024). https://arxiv.org/abs/2402.02207v1
- arXiv, "Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning" (2024). https://arxiv.org/html/2409.15361v1
- arXiv, "CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security" (2025). https://arxiv.org/pdf/2503.09334
- Ouyang et al. (OpenAI), arXiv, "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Rafailov et al. (Stanford), arXiv, "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- Pushkarna, Zaldivar, Kjartansson (Google Research), arXiv, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
- arXiv (Rottger et al.), "XSTest: A Benchmark for Identifying Exaggerated Safety in Language Models" (2023). https://arxiv.org/abs/2308.01223
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.