skip to content

After SFT on one narrow task, general answers degrade — what happened and what do you do?

level: seniorimportance: must knowfreq 58%

answer

  1. one weight set holds everything
  2. narrow data pulls the whole model
  3. alignment is learned behaviour too
  4. rehearsal: mix the old back in
  5. compare against the base, not the task

basics

~20 s

That is catastrophic forgetting: gradients from a narrow dataset overwrite general capability, and often loosen refusal behaviour too. The standard mitigation is mixing general instruction and safety examples back into the training set, plus tracking a capability and safety baseline against the base model.

solid answer

~50 s

Fine-tuning moves weights that encode everything the model already knew, so training exclusively on one narrow distribution — say structured disposition summaries from call transcripts — drags the model toward that distribution and away from everything else. General question answering, formatting variety and instruction following all degrade, and the effect is not limited to capability: refusal and safety behaviour learned during alignment erodes too, sometimes on topics unrelated to your data. You detect it by scoring the tuned model against the base model on a general-capability suite and a refusal set, not only on your task metric; a several-point drop is the signal. The mitigations are to mix a slice of general instruction data and explicit safety examples back into the SFT set, to keep the update small and short rather than training the narrow data to convergence, and to prefer an adapter over a full-weight update so the base behaviour remains recoverable.

go deeper

for a junior

Know the term catastrophic forgetting and the basic cause: fine-tuning changes the same weights that hold everything the model already knew, so training on one narrow task can make it worse at others.

for a middle

Explain why narrow data causes drift, and give the standard mitigation — mixing general instruction and safety examples into the training set, and keeping the update short rather than training to convergence.

for a senior

Show the production discipline: a fixed pre/post comparison against the base model on general-capability and refusal sets, sampling outside the task, and treating a compliance where the base refused as a release blocker rather than a curiosity.

for a principal

Own the policy. Define a regression budget agreed before training, mandate safety evaluation on every fine-tune regardless of how innocuous the data looks, and decide the organizational default — reversible adapters over full-weight updates — so a regression can be rolled back in minutes.

## Why forgetting is the default, not a bug A pretrained, instruction-tuned model is a single set of weights carrying everything it can do: language, world knowledge, formatting, instruction following, and the refusal behaviour installed by alignment training. Supervised fine-tuning does gradient descent on those same weights using a new objective. Nothing in that objective preserves prior behaviour — it only rewards fitting your examples. So the model drifts toward your data distribution, and abilities that your data never exercises drift away. This is **catastrophic forgetting**, and it is the expected outcome of narrow training, not an anomaly. The more homogeneous your dataset, the stronger the effect. Twelve thousand contact-centre transcripts that all produce the same structured summary is about as narrow as it gets: one domain, one input shape, one output shape, one register. The model learns that mapping very well and gets worse at almost everything else. ## The two distinct regressions **Capability regression** is the familiar half. The model answers general questions less well, follows unusual formatting instructions less reliably, and loses some of its range. It typically shows up as a several-point drop on broad knowledge-and-reasoning evaluations relative to the base model. **Safety regression** is the half people forget, and it is the one that will end up in an incident review. Alignment behaviour — refusing harmful requests, hedging on medical or legal advice, declining to impersonate — is itself learned behaviour stored in the same weights, and narrow fine-tuning erodes it. As of mid-2026 there is a well-documented and unsettling stronger version: fine-tuning on data that looks entirely innocuous, in a narrow domain, can induce broad misalignment far outside that domain. This has been reproduced across several open-weight model families, which makes safety regression testing a non-optional part of a fine-tuning pipeline rather than a nicety. ## Detecting it The trap is that your task metric goes *up* while everything else goes down. If the only number you watch is task quality, forgetting is invisible. So you measure against the base model on axes your training data does not cover: - **A general-capability suite** run identically on base and tuned checkpoints. Absolute scores matter less than the delta; a drop of a couple of points on a broad benchmark reads as real forgetting. - **A refusal/safety set** of prompts the base model correctly declines, plus borderline ones. Any request the base refused and the tuned model complies with is a regression — score it as such. - **A general instruction-following set** covering formats and tasks your fine-tune never touches. - **Qualitative sampling.** Ask the tuned model something completely outside its new task and read the answer. Models that have over-fit a narrow output format often try to force unrelated questions into that format, which is instantly visible and never appears in an aggregate score. Run all of these as a fixed pre/post pair for every fine-tune, so the comparison is automatic rather than remembered. ## Mitigating it **Mix general and safety data back in.** This is the standard, well-established move: blend a slice of broad instruction-following data and explicit safety/refusal examples into the SFT mixture alongside your task data. It is a form of rehearsal — the objective now includes keeping the old behaviour, so gradients stop being free to discard it. The mixing ratio is empirical; the point is that some non-trivial fraction of general and safety data is present rather than none. **Keep the update small.** Forgetting scales with how far you move the weights. Training for fewer passes, and stopping when the task metric plateaus rather than pushing it to convergence, preserves more of the base behaviour. There is a real tension here: the run that maximises task score is usually not the run that best preserves everything else, and picking the checkpoint is a judgement call informed by both numbers. **Prefer an adapter to a full-weight update.** Constraining the update to a small set of added parameters both limits drift and — crucially — keeps the original weights intact, so the base behaviour is one toggle away. Reversibility is an operational property, not just a training one: it means a safety regression discovered in production can be undone immediately. **Retain a representative slice of the original behaviour in evaluation, permanently.** Forgetting compounds across successive fine-tunes of the same lineage; each round on narrow data pushes further. A model on its third task-specific tune, each without rehearsal data, can be markedly narrower than anyone intended. ## What a strong answer sounds like Name the phenomenon, distinguish capability from safety regression, insist on base-versus-tuned comparison on axes the training data does not cover, and give the mitigation stack in order: mix rehearsal and safety data, keep the update small, prefer a reversible adapter. Saying explicitly that a narrow, innocuous-looking dataset can still loosen alignment — and that you therefore run a refusal set on every fine-tune — is the detail that marks someone who has shipped this rather than read about it.

  • Why is safety regression treated separately from capability regression?
    Because refusal behaviour is learned during alignment and stored in the same weights, narrow fine-tuning can erode it even when the training data is entirely benign — an effect reproduced across several open-weight families as of mid-2026. A capability benchmark will not surface it, so you need a dedicated refusal set scored as base-refused-versus-tuned-complied, run on every fine-tune.
  • How does mixing general data back in actually prevent forgetting?
    It changes the objective. With only narrow data, nothing penalises losing general behaviour, so gradients discard it freely. Adding general instruction and safety examples means the loss now includes those behaviours, so the optimizer must keep them. It is rehearsal: the model is reminded of the old distribution while learning the new one, at the cost of a somewhat smaller gain on the target task.
  • Does using an adapter eliminate forgetting?
    No — it reduces and contains it. A constrained update drifts less, and the frozen base weights mean the original behaviour is recoverable by disabling the adapter, which matters operationally when a regression is found in production. But an adapter's outputs still route through the whole model and can override prior behaviour, so you must still run the capability and safety comparison.
  • How would you pick a checkpoint when task score and general capability disagree?
    Treat it as a two-metric decision with an explicit floor. Set an acceptable regression budget in advance — say, no measurable increase in the refusal-failure rate and a bounded drop on the general suite — then choose the checkpoint with the best task score inside that budget. Without a pre-agreed floor, teams always pick the highest task score and discover the regression later.

saying these in an interview costs you the question

  • Only measures the target task and never compares against the base model
  • Assumes alignment and refusal behaviour survive fine-tuning untouched
  • Thinks benign-looking narrow data cannot loosen safety behaviour
  • Trains the narrow data to convergence to maximise the task metric
  • Believes using an adapter removes the need for regression testing

context