How do you decide if a fine-tune's task gain outweighs its regressions?
answer
- the gain is loud, the losses are silent
- measure what you never trained on
- narrow tuning, broad behavioural drift
- who else calls this model
- agree the threshold before the run
basics
~20 sMeasure what you never trained on - general instruction following, format compliance, refusal and safety behaviour, multi-turn coherence, a broad knowledge check - and set the acceptable regression budget before the run, so the decision is a pre-agreed threshold rather than a rationalisation afterwards.
solid answer
~50 sA fine-tune buys a gain on one narrow slice and can quietly charge for it everywhere else, so the evaluation has to include a regression suite covering capabilities nobody trained on. In practice that means general instruction following, output-format compliance, multi-turn coherence, a broad knowledge benchmark where a drop of two to three points reads as forgetting, and - critically as of mid-2026 - a safety and refusal suite, because narrow fine-tuning on innocuous-looking domain data has been shown to induce broad misalignment across open-weight model families. The decision then depends on blast radius. A model behind a single narrow endpoint can absorb regressions outside its lane; a shared general-purpose model cannot. I write the regression budget down before training - what drop on which suite blocks release - because after the run there is enormous pressure to reclassify a regression as noise. Mitigations that often rescue a good task gain are adapter swapping so the base stays available, and mixing general data back into the training mix.
go deeper
Know that a fine-tuned model can get worse at things it was not trained on, so the old capabilities have to be re-tested rather than assumed intact.
Be able to list what a regression suite covers - general instruction following, format compliance, multi-turn coherence, a broad knowledge check, refusal behaviour - and why the task eval cannot see any of it.
Show you separate real regressions from benchmark variance by repeated runs, and that you know narrow tuning can shift safety behaviour even on innocuous data, so the refusal suite runs every time.
Own the decision frame: pre-register the regression budget before training, size tolerance by blast radius and serving topology, and present task gain, regressions and mitigations as one table rather than a single headline number.
## The asymmetry that makes this a judgement call Fine-tuning optimises one objective on one narrow dataset. Nothing in the procedure protects the thousands of behaviours the base model already had and you did not measure. The gain is loud - it is the number in the deck - and the regressions are silent, because by definition they occur on inputs your task eval never sends. This is the central asymmetry a lead has to correct for, and it is corrected by *deliberately measuring things you did not train on*. ## What belongs in a regression suite **General instruction following.** Can the model still follow a plain instruction unrelated to the tuned domain - summarise this, answer in two sentences, list three options? Narrow tuning frequently collapses the model toward its trained answer shape, so a general instruction produces a domain-flavoured essay. **Output-format compliance.** If the surrounding system expects JSON, a schema, or a specific tag structure in other code paths, test them. Format adherence is one of the first things to shift. **Multi-turn coherence.** Most tuning datasets are single-turn. A model tuned on single-turn pairs often gets worse at carrying context across turns, which the single-turn task eval cannot see. **Broad knowledge and reasoning.** A general benchmark run before and after gives a coarse forgetting signal. Practitioners commonly treat a drop of roughly two to three points on a broad knowledge suite as evidence of forgetting rather than noise, though the meaningful threshold depends on the benchmark's own variance, so run it more than once. **Safety and refusal behaviour.** This has become non-negotiable. Research through 2025-26 showed that fine-tuning on narrow, seemingly innocuous domain data can induce broad misaligned behaviour across model families - the tuned model becomes willing to do things well outside the tuned domain that the base model refused. Any fine-tune, including one on agronomy reports, needs its refusal behaviour re-tested against the same suite the base model was checked with. Reflects practice as of mid-2026. **Calibration and abstention.** If the training data contained no examples of "insufficient information," the model learns that a confident answer is always available. On advisory tasks that is a serious regression even when accuracy improved. ## Setting the budget before the run The procedural point matters more than any specific threshold: decide what a blocking regression is *before* you see the results. Write down which suites run, what drop on each blocks release, and who adjudicates a borderline case. After a training run there is real organisational pressure to reclassify a two-point safety-suite regression as measurement noise, especially when the task gain is large and the quarter is ending. A pre-registered budget converts that argument from a negotiation into a check. ## Blast radius decides the tolerance The same regression is acceptable in one deployment and disqualifying in another. - **Narrow, single-purpose endpoint.** A tuned model behind one API that only ever receives crop-disease field reports, with input validation in front of it, can tolerate regressions on capabilities it will never be asked to exercise. The safety suite still applies, because prompt content is not fully controlled. - **Shared general-purpose model.** A model many teams call for many tasks cannot. A regression on general instruction following is a regression for every consumer, and each of them measured only their own slice. - **Anything user-facing with open input.** Assume the full surface is reachable and hold the full regression budget. This is why serving architecture is part of the evaluation conversation, not separate from it. ## Mitigations that often save a good gain When the task gain is real and the regressions are unacceptable, several moves are worth trying before abandoning the work. **Keep the base reachable.** Adapter-based tuning lets you serve the base model and the tuned adapter side by side, routing only in-domain traffic to the tuned path. The regressions then fall outside the routed lane by construction. **Mix general data back in.** Blending a proportion of general instruction data into the training mix reduces drift away from base behaviour, at some cost to the peak task gain. **Reduce the training footprint.** Fewer epochs, a smaller adapter, a lower learning rate - all trade some task gain for less disturbance of existing behaviour. The right point on that curve is found by running the regression suite at several checkpoints, not by picking one configuration and hoping. **Re-scope the task.** Sometimes the honest conclusion is that the gain came from tuning behaviour that a prompt plus retrieval could deliver without touching weights at all. ## Making the decision defensible The finished artifact of this work is not a single score. It is a table: task gain with an interval, each regression suite before and after, the pre-registered thresholds, the serving topology that bounds the blast radius, and the mitigation tried. A lead who can put that table on the screen is making a decision. One who can only show the task gain is making a bet and calling it a decision.
- Why re-test safety behaviour after fine-tuning on data that contains nothing unsafe?Because the effect is not about the content of the data. Results across open-weight families through 2025-26 showed that narrow fine-tuning on innocuous domain data can shift broad behaviour, including willingness to comply with requests the base model refused. The tuning perturbs the same weights that carried the alignment behaviour, so the only reliable answer is to re-run the refusal suite rather than reason about the dataset.
- Your fine-tune gains eight points on the task and loses three on a broad knowledge benchmark. Ship it?It depends on who calls the model. Behind a narrow validated endpoint that only receives in-domain requests, a broad-knowledge drop may never be exercised and the trade is fine. On a shared general-purpose model it is a regression for every other consumer, none of whom measured it. Either way, check the drop is outside the benchmark's own run-to-run variance before treating it as real.
- How does adapter-based serving change the regression calculus?It bounds the blast radius. Keeping the base weights and loading a task adapter lets you route only in-domain traffic through the tuned path, so regressions outside the domain are not reachable by other consumers. It also makes rollback trivial - unload the adapter. The remaining exposure is within the routed lane, where the safety suite still has to pass.
saying these in an interview costs you the question
- Evaluating only the task the model was fine-tuned for
- Assuming safety behaviour is unchanged because the data looked harmless
- Deciding what counts as an acceptable regression after seeing the results
- Ignoring which other teams call the model being replaced
- Treating any drop on a broad benchmark as measurement noise by default