Fine-tuning and post-training data
Data mixtures for fine-tuning: preventing catastrophic forgetting
Quick answer
Catastrophic forgetting in fine-tuning is the loss of general skills (instruction following, safety refusals, reasoning, output formats) when a model is trained on a narrow domain set. The most dependable data-side fix is a deliberate mixture: domain examples blended with replayed general instruction data and safety data, dominant task types capped, and the ratio chosen by a sweep against regression evaluations run before and after training. Parameter-efficient methods such as LoRA can reduce forgetting but do not remove it.
By SourceX Editorial · Updated
How catastrophic forgetting shows up after domain SFT
Forgetting after domain SFT usually appears first as regressions in behaviors the domain set never exercises, not as a collapse in the target task. In an empirical study of continual instruction tuning, Luo et al. observed forgetting across models from 1B to 7B parameters, and severity grew with scale in that range [1]. A 2026 study reports that forgetting severity correlates strongly with how similar the new task is to prior tasks (r = 0.87) and with gradient alignment, which is a direct argument for designing the mixture rather than hoping the base model holds [2].
The failure modes a post-training team should expect are specific:
- Instruction-following drift: the model ignores length, tone or format constraints because every domain target used one template.
- Format collapse: every answer becomes the domain output shape, such as a claim note or a JSON object, even for unrelated prompts.
- Safety regression: refusal behavior weakens. Qi et al. showed fine-tuning aligned models can compromise safety even on benign data [3], and later work found fine-tuning also makes safety evaluation results less consistent [4].
- Reasoning and knowledge loss: multi-step math or code ability degrades when the domain set contains only short extractive answers.
- Multilingual erosion: non-English quality drops when the domain set is English-only.
Practitioner guidance converges on the same point: narrow datasets forget more than diverse ones, and forgetting must be measured explicitly rather than inferred from target-task gains [7].
What belongs in a forgetting-resistant SFT mixture
A forgetting-resistant mixture has four components: domain data, general instruction replay, safety data and format-preserving examples. Each one protects a different capability, so dropping one tends to open a matching regression.
Domain data carries the new skill. Its quality matters more than its size; LIMA fine-tuned a 65B model on 1,000 curated prompt-response pairs and argued that most capability comes from pre-training, with SFT mainly teaching style and format [6]. For sourcing criteria see how to source supervised fine-tuning data.
General instruction replay is a sample of broad, diverse instruction data resembling what the base or instruct model was originally tuned on. If you do not have the original post-training data, an openly documented mixture such as the Tulu 3 SFT set is a common reference point because its composition and evaluation suite are published [5]. Check the license of every component first; open instruction datasets differ on commercial use.
Safety data pairs harmful or borderline prompts with appropriate refusals, plus benign look-alike prompts with helpful answers so the model does not over-refuse. Given the findings in [3] and [4], treat this as mandatory for any instruct model you will deploy.
Format-preserving examples cover the chat template, system prompts, multi-turn context and tool calls the production model must still handle. Mismatched role tags or loss masking cause their own regressions; see chat fine-tuning data format.
Choosing the domain-to-general ratio
There is no universal ratio; the defensible approach is a small sweep of mixture weights scored against both target metrics and a regression suite. The right share of general replay depends on base model size, how far the domain sits from the base distribution, learning rate, number of epochs and whether you use full fine-tuning or adapters.
A practical sweep runs three to five mixtures that vary only the general share, with identical hyperparameters and seeds. Plot target-task score against the worst regression on the general suite and pick the point where target gains flatten before regressions climb. Re-run the sweep when you change the base model or the learning rate, because the trade-off moves with both.
Count the ratio in the unit your trainer consumes. For SFT that is usually examples per task family, weighted by loss-bearing tokens, because a few long-document examples can dominate gradient updates even when they are a small share of rows. Packing and truncation settings change effective token weights, so record them with the mixture.
Task balancing inside the domain set
Capping dominant task types inside the domain set is as important as the domain-to-general split. Operational data is heavily skewed: a support archive may be mostly password resets, and an engineering record set mostly one ticket type. Without caps, the model overfits the majority pattern and rare but valuable tasks are barely learned.
Use a per-task cap (for example, a maximum number of examples per intent label or document type), then upsample rare task families with a bounded multiplier so you do not memorize a handful of rare examples. Deduplicate near-identical prompts before capping; templated records inflate counts without adding signal. The filtering methods in instruction-tuning data quality filtering apply directly here.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Mixture component | Source type | Share of examples (run B) | Per-task cap | Protects |
|---|---|---|---|---|
| Domain: claim-note drafting | Licensed operational records | 30% | 4,000 per claim type | Target skill |
| Domain: policy Q&A with citations | Licensed documents plus written answers | 15% | 2,500 per product line | Target skill, grounding |
| Domain: rare escalations | Licensed records, upsampled at most 3x | 5% | none (floor of 300) | Long-tail tasks |
| General instruction replay | Open, commercially licensed mixture | 35% | 1,500 per category | Instruction following, reasoning |
| Safety and over-refusal pairs | Curated refusal and benign look-alike sets | 10% | n/a | Safety alignment |
| Format and tool-call examples | Internal templates | 5% | n/a | Chat template, tool use |
Run A and run C in the sweep would shift the general replay share down and up while holding the other proportions in the same relative balance.
Replay in continued pre-training versus SFT mixtures
Replay in continued pre-training mixes raw token streams, while an SFT mixture is balanced by task; the principles transfer but the knobs differ. In continued pre-training, replay means mixing a fraction of tokens from the original pre-training distribution into the new-domain stream, usually alongside learning-rate re-warming and re-decay. In that setting, replay is measured as a fraction of tokens from the original corpus [10].
For SFT you rarely replay raw pre-training text. You replay instruction-shaped examples, balanced by task category, and you measure success with behavioral evaluations rather than held-out perplexity. If your plan includes both stages, design two separate mixtures. The trade-offs between stages are covered in fine-tuning vs RAG vs continued pre-training data.
Training choices that interact with the mixture
Parameter-efficient fine-tuning can reduce forgetting but tends to trade away some target-domain accuracy. Published comparisons generally report that LoRA-style adapters preserve more out-of-domain ability than full fine-tuning while learning less of the target domain; verify this on your own model rather than assuming it [9]. A LoRA run may therefore tolerate a smaller general share, but it still needs regression testing.
Other levers move the same trade-off: lower peak learning rate, fewer epochs over the domain set, and early stopping on the regression suite rather than on domain loss alone. Change one lever at a time so you can attribute a regression to the mixture or to the optimizer.
Regression evaluations before and after fine-tuning
Forgetting is only visible if you run the same regression suite on the base checkpoint and every candidate checkpoint. The suite should include:
- Your own held-out domain tasks, split by task family so rare tasks are scored separately.
- General instruction-following checks with explicit constraints on length, format and language.
- Safety and over-refusal probes, run several times because results vary across runs [4].
- Reasoning and code checks matched to what the base model could already do.
- Format and tool-call conformance on the production chat template.
Decontaminate before you trust any score. Public benchmarks can leak into training mixtures; as of October 2026, OpenAI no longer reports SWE-bench Verified scores because it judged gains increasingly reflected training exposure [8]. Hash or n-gram match every training component, including purchased domain data, against your evaluation prompts. For acceptance criteria on bought data, see how to evaluate a fine-tuning dataset before you buy it.
What to specify when buying domain data for a mixture
When you buy domain data for a mixture, specify the task-level metadata you need to balance it, not only the volume. Ask suppliers for a task or intent label per record, document or ticket type, language, date range, record length distribution, a deduplication method, and a held-out split that has never been shared elsewhere. Without those fields you cannot cap dominant tasks or build a clean regression set.
Confirm that the license covers fine-tuning and evaluation use, including mixing with third-party data; the fine-tuning-only data license guide covers scope limits. The broader landscape is in the fine-tuning and post-training datasets buyer's guide, and the SourceX overview of training data for domain-specific fine-tuning describes the domain side. Definitions for supervised fine-tuning and fine-tuning are in the glossary.
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows, which are the raw material for the domain share of a mixture. If you can describe the records and task labels you need, start a request on the SourceX buyers page.
Source domain data for your SFT mixture
SourceX looks for US businesses that hold the data you describe, rights-reviews each dataset for ownership and consents, and delivers it under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Data is sourced on request, a request does not guarantee a match, and nothing is contracted until a supplier agrees; describe your domain data needs to SourceX.
Sources
- arXiv (Luo et al.), "An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning" (2023). https://arxiv.org/abs/2308.08747v1
- Olaf Yunus Laitinen Imanov (arXiv:2601.18699), "Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning" (2026). https://arxiv.org/abs/2601.18699v1
- arXiv (Qi et al.), "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!" (2023). https://arxiv.org/pdf/2310.03693
- arXiv (Fraser et al.), "Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency" (2025). https://arxiv.org/abs/2506.17209v1
- arXiv (Allen Institute for AI), "Tulu 3: Pushing Frontiers in Open Language Model Post-Training" (2024). https://arxiv.org/pdf/2411.15124
- arXiv (Zhou et al.), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Sabr Research, "Catastrophic forgetting in fine-tuning". https://www.sabrresearch.com/cookbooks/catastrophic-forgetting-in-fine-tuning
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- arXiv (Biderman et al.), "LoRA Learns Less and Forgets Less" (2024). https://arxiv.org/abs/2405.09673v2
- arXiv (Ibrahim et al.), "Simple and Effective Continual Learning for Large Language Models" (2024). https://arxiv.org/abs/2402.16402
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.