What baselines must a fine-tuned model beat before you ship it?
answer
- compared against what, exactly
- the cheapest thing that could have worked
- prompt and few-shot arms, same eval set
- identical decoding across every arm
- paired win rate with an interval
basics
~20 sAt minimum the same base model prompted properly - a strong system prompt and a few-shot variant - plus whatever runs in production today. Score every arm on one held-out set with identical decoding settings, or the improvement is unattributable.
solid answer
~50 sA fine-tune has to beat the cheapest thing that could have worked, so I run a three-arm comparison: the base model with a seriously-written system prompt, the base model with a handful of good few-shot examples, and the fine-tune. All three see the same held-out inputs, the same decoding settings (temperature, max tokens, stop sequences) and the same output parsing - anything that differs between arms gets silently credited to the fine-tune. For a crop-disease advisory model I would sample around 300 held-out field reports, generate all three answers per report, and have agronomists compare them blind and pairwise in randomised order, reporting a win rate with a confidence interval rather than one averaged score. Then I weigh that win against the cost of owning a trained model: if the prompt captures most of the gain, the serving, versioning and retraining burden is usually not worth it.
go deeper
Know that a fine-tuned model must be compared with the same base model given a good prompt, on data it never trained on. Say plainly that scores on training examples prove nothing.
Be ready to lay out the arms - prompted base, few-shot base, fine-tune, incumbent - and to name what must be held constant: same held-out inputs, same decoding settings, same parsing and scoring code.
Show you report paired differences with a confidence interval instead of a single average, and that you slice results by segment to catch a regression hiding under an aggregate win.
Own the framing that the margin has to justify a lifetime cost: retraining on every base-model upgrade, dataset ownership, serving a second artifact. Be willing to kill a statistically real but operationally worthless win.
## Why baselines are the whole question A fine-tune is never evaluated in isolation. The number that matters is not "the fine-tune scores 78%" but "the fine-tune scores 78% where the alternative you could have shipped in an afternoon scores 74%." Fine-tuning costs a training run, a dataset somebody must keep curated, a model artifact you must version and serve, and a standing commitment to retrain whenever the base model or the task drifts. That bill is only justified by a margin over the cheapest thing that could have worked, and the cheapest thing is a well-written prompt. ## The arms of the comparison A defensible evaluation has at least three arms, and often four. 1. **Base model, strong system prompt.** The instructions a competent engineer would write after a day of iteration: role, output format, the domain rules the fine-tune was supposed to teach. 2. **Base model, few-shot.** The same prompt plus a small number of high-quality demonstrations drawn from the training pool - never from the held-out set. 3. **The fine-tuned model**, served with the chat template and system prompt it was trained under. 4. **The incumbent**, if something is already in production. Beating a prompt is academic if you cannot beat the system you are proposing to replace. A retrieval-augmented arm belongs in the list whenever the task needs facts rather than form, because retrieval and fine-tuning solve different problems and are frequently combined rather than chosen between. ## Holding everything else constant The most common way a fine-tune "wins" is that the comparison was not controlled. Every arm must share: - **The same held-out inputs.** One eval set, drawn once, used by all arms. - **The same decoding settings.** Temperature, top-p, max tokens and stop sequences change output quality on their own. Comparing a fine-tune at temperature 0 with a prompted baseline at temperature 1 measures decoding, not training. - **The same parsing and scoring code.** If the fine-tune emits clean JSON and the baseline emits JSON inside prose, and your parser only handles the former, you are measuring your parser. - **The same effort spent on the prompt.** A straw-man baseline prompt is the easiest way to manufacture a win. Have someone who wants the prompt to succeed write it, ideally with the same time budget the fine-tune consumed. ## Reporting a result you can defend Averaged scalar scores hide the interesting structure. Two habits fix that. **Pair the comparisons.** Each held-out input is answered by every arm, and you compare arm-to-arm on the same input. Paired comparison removes item difficulty as a source of noise, so the same sample size buys a much tighter estimate. **Report an interval, not a point.** With 300 items and a small margin, the difference between 76% and 78% is frequently indistinguishable from noise. Bootstrap the paired differences and report the interval. If it straddles zero, the honest statement is "no measurable difference at this sample size," which is a real and useful finding. For open-ended output, blind pairwise preference is the strongest signal available: present two anonymised answers per input, randomise which side each arm appears on, and report a win rate. In the crop-disease example, agronomists judging 300 field reports pairwise gives a directly interpretable number - "the fine-tune was preferred on 61% of reports where the raters expressed a preference." ## Slicing before deciding An aggregate win can hide a regression that matters more than the average. Slice the eval set by the dimensions your users care about - report type, crop, severity, report length, unfamiliar pathogens - and look for slices where the fine-tune is worse. A model that gains four points overall while losing ten on rare-but-severe cases is not a shipping candidate; it is a retraining brief. ## Turning the numbers into a decision The last step is not statistical. A fine-tune that wins by a clear margin still has to justify its lifetime cost against a prompt that is free to change, needs no retraining when the base model is upgraded, and can be debugged by reading it. The cases where a fine-tune reliably clears that bar are consistent behaviour that no prompt reliably enforces, an output format the base model keeps drifting away from, latency and cost gains from moving work onto a smaller tuned model, and tone or style that is easier to demonstrate than to describe. When the prompt gets you most of the way, ship the prompt and revisit later - that decision is much easier to reverse than a training pipeline nobody owns.
- Your fine-tune wins on average but the confidence interval crosses zero. What do you do?Report it as no measurable difference rather than a win. Then either grow the eval set until the interval is informative, or switch to paired blind preference, which is far more sensitive than averaged scalar scores. Also slice the results - a flat aggregate sometimes hides a real gain on one segment and a real loss on another, which is a more useful finding than the average.
- How do you stop the prompted baseline from being a straw man?Give it the same effort budget the fine-tune got, and have it written by someone motivated to make it win. Let it use few-shot examples, an explicit output schema and the domain rules encoded in the training data. If nobody on the team can defend the baseline prompt as their best attempt, the comparison is not evidence.
- Should the fine-tuned model be evaluated with the same prompt as the baselines?No. Each arm gets the prompt it performs best under: the fine-tune is served with the chat template and system prompt it was trained on, and the baselines get their own best prompts. Forcing one shared prompt understates whichever arm it suits less. What must be identical is the inputs, the decoding settings and the scoring.
saying these in an interview costs you the question
- Comparing the fine-tune only to the raw base model with no prompt
- Reporting training-set or in-sample scores as evidence of improvement
- Different temperature or parsing per arm, then crediting the fine-tune
- Calling a two-point gap on forty examples a win
- Never checking whether a prompt alone would have been enough