skip to content

Fine-tune vs Prompting

Fine-tune for form, not facts: style, cost at volume, or behaviour you cannot describe in words. Otherwise prompt, then retrieve — and often the real answer is distilling into a small tuned model.

on this pageshow

questions

5

When is fine-tuning the right call instead of a better prompt or retrieval?

level: juniorimportance: must knowfreq 72%

answer

  1. climb only when the current rung fails
  2. cheapest rung that clears the bar
  3. prompt, retrieve, tune, distill
  4. behaviour gap, not knowledge gap
  5. a tune is a standing commitment

basics

~20 s

Fine-tune only after prompting and retrieval have been tried and still miss. The cases that justify it are behavioural: an output format or house style hard to describe in words, or a stable high-volume task where per-call cost is the binding constraint.

solid answer

~50 s

The working default is an escalation ladder — prompt, then retrieval, then fine-tune, then distill — and you stop at the cheapest rung that clears the quality bar. A better prompt costs an afternoon and is reversible. Retrieval fixes the most common failure, which is that the model was never shown the relevant facts. Fine-tuning is worth reaching for when the gap is **behavioural**: a house output format, a register, or a judgement call you can demonstrate with examples but cannot describe in instructions. It is also worth it when a task is stable and high-volume enough that paying a large model per call dominates your bill, in which case you tune a small model to do the job. What fine-tuning is not is a way to install facts, and it is not free: you now own a dataset, an eval suite, and a retrain every time the base model changes.

go deeper

for a junior

Be able to name the order — prompt, retrieval, fine-tune, distill — and say plainly that you try the cheap, reversible options first. Knowing that tuning changes behaviour rather than knowledge is enough at this level.

for a middle

Explain what evidence moves you up a rung: a fixed eval set, a plateau in prompt iteration, and a look at whether residual errors are missing facts or wrong form. Be ready to say why retrieval survives a fine-tune.

for a senior

Show the production judgement: cost per call at real volume, stability of the requirement, who owns the dataset and eval, and what happens when a new base model ships. Talk about the tune and retrieval as complements in one system.

for a principal

Own the framing that a fine-tune is an organizational commitment, not a technique. Be prepared to argue the portfolio view — which tasks deserve a trained model at all, and what standing budget the team is signing up for in retraining and validation.

## The question behind the question An interviewer asking this wants to know whether "we should fine-tune" is your reflex or your conclusion. Fine-tuning is the most expensive and least reversible option in the toolkit, and in most teams the majority of proposed fine-tunes are solving a problem that a clearer prompt or a retrieval step would have solved in a day. Being able to name the ladder, and to say what evidence moves you up a rung, is the whole answer. ## The ladder The conventional ordering is **prompt → retrieval → fine-tune → distill**. Each rung costs more in engineering time, money and ongoing maintenance than the one before it, so you climb only when the rung you are on has demonstrably failed against a fixed evaluation set. **Rung 1 — prompting (often now called context engineering).** Rewriting instructions, adding a handful of worked examples, splitting one prompt into a short chain, or asking for a structured output shape. It is measured in hours, versioned in git, and rolled back instantly. Long prompts used to be a cost argument for tuning; caching of repeated prompt prefixes has weakened that argument considerably. **Rung 2 — retrieval.** Fetch the relevant documents, records or policy text at query time and put them in the context. This is the right fix whenever the model's failure is "it did not know", including anything that changes on a weekly or monthly cadence. Retrieval is also auditable: you can show which source produced the answer, which is often a hard requirement in regulated work. **Rung 3 — fine-tuning.** Adjust the weights on examples of the behaviour you want. This is the rung for *form*: a memo structure, a tone, a consistent labelling convention, a compressed reasoning style, or an implicit judgement threshold that a hundred examples convey better than three paragraphs of instructions. **Rung 4 — distillation.** Once a large model does the task acceptably, use its outputs as training data for a much smaller model, then serve the small one. This rung is about unit economics, not quality ceiling — the student rarely beats the teacher. ## A worked example A fintech KYC operations team wants to automate adjudication of sanctions alerts. First they improve the prompt: the alert record, the decision options, the escalation criteria, three worked examples. That takes them from unusable to roughly right. Next they add retrieval over the policy manual and the customer file, because most remaining errors are the model not knowing a rule or a prior decision. Now the analysis is correct but the memos do not look like the ones the regulator has seen for five years, and reviewers keep rewriting them. That is a *form* gap, so they fine-tune on past memos. Finally, at three million alerts a month, the frontier model's per-call cost dominates the programme budget, so they distil into a small tuned model for the routine tier and keep the large one for escalations. Notice the ladder never removed retrieval. The sanctions list changes weekly; no amount of tuning keeps that current. ## What actually decides it Ask three questions. *Is the failure missing knowledge or wrong behaviour?* Missing knowledge means retrieval or, at large scale, continued pretraining; wrong behaviour means tuning. *Is the task stable?* A fine-tune amortises over months of unchanged requirements and is wasted on a spec that moves weekly. *Is cost or quality the binding constraint?* If quality is fine and money is the problem, you want a smaller model, which is a distillation project, not a quality project. And one more: can you produce the data? A fine-tune needs examples of the target behaviour, plus a held-out set you did not train on. If the team cannot assemble a few hundred good rows, the decision has already been made for you. ## The costs people forget A fine-tune is a standing commitment. You own the dataset and its provenance, an evaluation suite that proves the tune still beats the prompt, a serving path for an extra model or adapter, and a retraining decision every time a stronger base model ships. Teams routinely discover a year later that simply prompting the newest base model matches their old tune. ## Common mistakes Reaching for training because the prompt got long; expecting a tune to teach facts; treating fine-tuning and retrieval as alternatives rather than complements; and skipping the fixed eval set, so nobody can tell whether the tune helped or whether the prompt was just under-engineered.

  • The prompt is now 6,000 tokens of instructions and examples. Is its length by itself a reason to fine-tune?
    Not on its own. Prompt-prefix caching makes a long, stable preamble much cheaper than it looks, and a long prompt is still versioned, inspectable and instantly editable. Length becomes an argument only when the cached cost still dominates at your call volume, or when the instructions have grown so tangled that the model follows them inconsistently — and that second symptom is a behaviour gap, which is the real justification.
  • How would you tell that prompting has genuinely run out rather than just being under-engineered?
    Fix an evaluation set first, then iterate the prompt against it until improvements plateau — several serious rewrites producing no measurable gain. Look at the residual errors: if they are missing facts, that is a retrieval problem; if they are the model doing the right analysis in the wrong shape, or applying an inconsistent threshold you cannot articulate, that is the behavioural residue only training removes.
  • Does fine-tuning let you drop the retrieval layer?
    Almost never. Tuning fixes how the model behaves, not what it currently knows, so anything volatile — prices, lists, tickets, policy revisions — still has to arrive at query time. The realistic end state is a tuned model plus retrieval: the tune supplies the format and judgement, retrieval supplies the facts. Dropping retrieval after tuning is a classic way to ship confident, stale answers.

saying these in an interview costs you the question

  • Fine-tuning is how you add new knowledge to a model
  • If we have data, fine-tuning always beats prompting
  • A fine-tune replaces the retrieval layer
  • The prompt got long, so we should train instead
  • Treating fine-tuning as a one-off rather than a standing cost

context

open as a page

What does "fine-tuning is for form, not facts" mean for an LLM?

level: middleimportance: must knowfreq 64%

basics

~20 s

Fine-tuning reliably changes how a model behaves — its format, tone and judgement — but not what it knows. Training on facts the base model never learned mostly teaches it to assert unfamiliar claims confidently, which raises the hallucination rate.

open as a page

When does distilling a frontier model into a small tuned model beat prompting it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Distillation pays when quality is already acceptable and cost or latency is the binding constraint: a stable, high-volume task where a frontier model's per-call price dominates. Its outputs become training data for a small model that serves the routine tier.

open as a page

When does reinforcement fine-tuning beat prompting on a task with an automatic grader?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Reinforcement fine-tuning fits when correctness is machine-checkable and the model already succeeds sometimes. You supply problems and a grader instead of gold answers, and training reinforces the model's own successful attempts — useful when you have verifiable outcomes but no written solutions.

open as a page

When does a fine-tune's maintenance bill outweigh the quality it buys?

level: principalimportance: should knowfreq 42%

basics

~20 s

A fine-tune is a standing obligation: dataset, eval suite, serving path, and a retrain decision every time a stronger base model ships. It stops paying when the requirement moves faster than you can retrain, or when prompting the current base already matches it.

open as a page