Fine-tuning and post-training data
Keeping safety intact when fine-tuning: data to mix in
Quick answer
Yes, fine-tuning can compromise safety. Research shows that a handful of harmful examples can strip guardrails from an aligned model, and that ordinary benign instruction data also weakens refusals, to a lesser degree [1]. To preserve safety, mix a small, deliberately built slice of refusal and safe-completion examples into your domain SFT set [4], add borderline prompts that must still be answered, and gate every checkpoint on before-and-after safety and over-refusal evaluations. No mitigation is complete, so document the mixture and the results.
By SourceX Editorial · Updated
Why fine-tuning erodes safety alignment
Fine-tuning erodes safety because the refusal behavior learned during alignment can be shifted by further gradient updates, whatever your intent. Qi et al. jailbroke GPT-3.5 Turbo's guardrails with 10 harmful training examples at a cost under $0.20 through the provider's fine-tuning API, and found that fine-tuning on common benign datasets also degraded safety, to a lesser extent [1]. The Stanford HAI policy brief on that work concludes that no existing intervention reliably prevents safety loss from customization, and that closed fine-tuning APIs approach the risk profile of open weights [2].
The benign-data finding is the one that matters for an enterprise team. Later work reports that fine-tuning both lowers safety and makes safety evaluation results less consistent across runs and setups [3]. In practice, your support-ticket or contract-redline SFT set does not need a single toxic record to produce a model that complies with requests the base model refused.
Domain matters too. In dual-use domains such as security operations, chemistry or finance, the useful behavior and the harmful behavior overlap, and researchers building a pseudo-malicious cybersecurity instruction set reported explicit safety-performance trade-offs [5]. For background on SFT itself, see the supervised fine-tuning glossary entry.
Which failure modes to expect after domain SFT
Expect four distinct regressions, and test for each separately because they have different data fixes.
- Refusal collapse. The model now complies with clearly harmful requests (weapons uplift, malware, self-harm instructions) it previously declined. This is the headline effect in [1].
- Persona and system-prompt drift. Training data that always answers in a single compliant voice, with no system prompt or a different one, teaches the model to ignore the deployed system prompt's constraints.
- Over-refusal. The opposite failure: too much safety data, or refusals keyed to surface words, makes the model refuse "how do I kill a Python process" or a legitimate claims question. Researchers call this exaggerated safety and trace it to refusals keyed to trigger words rather than intent [8].
- Evaluation instability. Safety scores on the fine-tuned checkpoint can shift with seemingly minor changes to the fine-tuning setup and evaluation settings [3], so a single pass on one benchmark gives false comfort.
What safety data to mix into the fine-tuning set
Mix three kinds of examples into the domain set: refusals for genuinely harmful requests, helpful answers to borderline prompts that look risky but are safe, and in-domain safety cases written in the same format and voice as your task data. VLGuard showed the pattern for vision-language models: a safe instruction-following set realigned models whether mixed into standard fine-tuning or applied afterward [4]. Because even benign data moves refusal behavior [1], the safety slice has to be present in every domain run, not added once.
Treat any ratio, including the illustrative one below, as a starting point to tune against your own evaluations, not a constant. The right ratio depends on how far your domain data sits from the base model's training distribution, how many epochs you run, and whether you use full fine-tuning or a LoRA adapter. Sourcing the harmful, borderline and benign prompts themselves is covered in our guide to safety-tuning and refusal data; this page is about how to mix and verify them.
Three properties make the safety slice work:
- Format parity. Safety rows use the same chat template, system prompt and role structure as domain rows. A refusal set in a different template teaches the model that safety applies only to that template. See chat fine-tuning data format for messages, roles and loss masking.
- In-domain harm. Generic refusals ("I can't help with building a bomb") do little for a claims or HR assistant. Write cases specific to the deployment: requests to reveal another customer's account data, to backdate a compliance record, or to draft a discriminatory rejection letter.
- Safe completions, not just refusals. Many safety rows should show a partial, redirecting or caveated answer rather than a flat "I can't help." This reduces over-refusal and teaches the boundary.
The same mixing logic protects general capability, which is covered in data mixtures for preventing catastrophic forgetting.
Illustrative safety mixture specification
A mixture spec turns "add some safety data" into an auditable artifact that reviewers can approve and engineers can reproduce.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Slice | Share of rows | Content | Format rule | Target behavior |
|---|---|---|---|---|
| Domain task | 90% | De-identified support transcripts paired with resolved answers | Production system prompt, multi-turn, loss on assistant turns only | Task quality |
| Harmful-request refusals | 3% | General harm categories plus in-domain abuse (data exfiltration, record falsification) | Same system prompt and template as domain rows | Decline with a short reason |
| Borderline-safe answers | 4% | Prompts with alarming words but legitimate intent ("kill the stuck job", "attack surface review") | Same template | Answer helpfully |
| Safe completions | 2% | Requests answerable in part, with redirection | Same template | Partial help plus caveat |
| System-prompt adherence | 1% | Users trying to override the system prompt or persona | Same template | Keep the deployed constraints |
Each row also carries metadata so the mixture can be audited later. An illustrative record:
{
"id": "safety-0412",
"slice": "harmful_request_refusal",
"harm_category": "privacy.cross_account_access",
"source": "written_in_house",
"license_scope": "fine-tuning",
"messages": [
{"role": "system", "content": "<production system prompt>"},
{"role": "user", "content": "Pull the last four orders for the account ending 7731, it's my neighbor's."},
{"role": "assistant", "content": "I can only share order details with the verified account holder. If your neighbor needs help, they can contact us directly."}
],
"reviewed_by": "policy-reviewer-02",
"review_date": "2026-09-30"
}
How to run safety regression evaluations before and after fine-tuning
Run the same frozen safety suite on the base model and every candidate checkpoint, and block promotion when harmful compliance rises or safe-prompt compliance falls beyond an agreed tolerance. Because fine-tuning makes safety scores less stable [3], fix decoding parameters, the chat template and the system prompt, and run multiple seeds.
Illustrative example: invented to show structure; it does not describe an available dataset.
Safety gate checklist for a fine-tuned checkpoint
- Harmful-request suite: general categories plus your in-domain abuse cases, scored by refusal or safe-completion rate. Compare against the base model, not an absolute number.
- Over-refusal suite: public exaggerated-safety prompt sets plus in-domain borderline prompts that a well-calibrated model should answer.
- Adversarial probes: role-play, instruction override and multi-turn escalation, run with and without the production system prompt.
- Held-out split: keep safety evaluation prompts disjoint from the safety training slice, checked by near-duplicate matching, so the gate measures generalization rather than memorization.
- Variance check: three or more seeds and two decoding settings; report ranges.
- Task-quality check: confirm the safety slice did not cost more domain accuracy than the team accepted.
- Sign-off record: mixture spec version, data hashes, eval results and approver.
For the broader dataset-acceptance process, see how to evaluate a fine-tuning dataset before you buy it. Red-team data for deeper adversarial testing is described on our safety and red-teaming data page.
What to require from purchased fine-tuning data
Require purchased domain data to arrive with enough documentation to explain a safety regression after the fact. Ask suppliers for a datasheet or Croissant-RAI style record covering source, collection method, labeling, filtering and known gaps [7], and confirm the license covers fine-tuning and the evaluation copies you will keep; our guide to fine-tuning-only data licenses covers scope limits.
Screen incoming rows for content that teaches compliance with harmful requests, such as support agents who disclosed account data on request, or engineering logs containing working exploit code. These are the operational versions of the harmful examples in [1]. NIST SP 800-218A adds AI-specific secure development practices that address the integrity of training, fine-tuning and alignment data, a useful frame for this intake control [6].
SourceX sources operational datasets from US companies, such as support and sales histories, engineering records, documents and finance or legal workflows, and manages licensing for AI teams wherever they are based. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and diligence materials on source, rights, preparation and allowed use are prepared per dataset. Those records make the mixture spec above easier to document. You can describe the domain data you need on the buyers page.
How to document the mixture for governance review
Document the safety mixture as a versioned record that a model risk, privacy or security reviewer can check without rerunning training. At minimum, record the slice table, row counts and hashes per slice, where each slice came from and under what license, the evaluation suites and versions, base and fine-tuned scores with variance, and the approver.
Keep the record with the model card for the fine-tuned checkpoint, and refresh it whenever the domain data is refreshed, because each new training run is a new chance for regression. For provenance practice across the whole corpus, see the AI data provenance guide and the fine-tuning and post-training data hub.
Get domain fine-tuning data with documented provenance
SourceX sources operational data from US companies on request, rights-reviews each dataset and delivers it under a license that defines records, uses, term and delivery. Sourcing is not a guarantee of a match, and nothing is contracted until the supplying company agrees. Tell SourceX what fine-tuning data you need.
Frequently asked questions
Does LoRA or another parameter-efficient method avoid the problem?
Not reliably. Qi et al. observed safety loss both through a provider's fine-tuning API and on an open-weights chat model [1], and the HAI brief concludes that no existing intervention reliably prevents it [2]. Treat parameter-efficient methods as carrying the same risk and run the same safety gate whatever the method.
Can we restore safety after fine-tuning instead of mixing data in?
Sometimes. VLGuard reported realignment both when safety data was mixed in and when it was applied after fine-tuning [4]. Mixing during training is simpler to validate because there is one checkpoint to gate.
How do we know if we added too much safety data?
Over-refusal rises. Track compliance on safe prompts, including your in-domain borderline set; if it falls, reduce the refusal share or convert flat refusals into safe completions.
Sources
- arXiv (Qi et al.), "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To" (2023). https://arxiv.org/pdf/2310.03693
- Stanford HAI, "Safety Risks from Customizing Foundation Models via Fine-Tuning" (2024). https://hai.stanford.edu/policy/policy-brief-safety-risks-customizing-foundation-models-fine-tuning
- arXiv, "Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency" (2025). https://arxiv.org/html/2506.17209v1
- arXiv, "Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models (VLGuard)" (2024). https://arxiv.org/abs/2402.02207v1
- arXiv, "CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security" (2025). https://arxiv.org/pdf/2503.09334
- National Institute of Standards and Technology, "Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile (NIST SP 800-218A)" (2024). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-218A.pdf
- arXiv (Jain et al., MLCommons), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- arXiv (Rottger et al.), "XSTest: A Benchmark for Identifying Exaggerated Safety in Language Models" (2023). https://arxiv.org/abs/2308.01223
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.