skip to content

Fine-tuning

Changing the model's weights instead of its prompt: supervised fine-tuning, LoRA and other parameter-efficient methods, the dataset that drives it, and the evaluation that proves it beat prompting. Interviewers usually want the decision and the data story, not the training loop.

on this pageshow

explore

questions

30

How many training examples does a narrow supervised fine-tune actually need?

level: juniorimportance: must knowfreq 60%

answer

  1. quality and coverage over raw count
  2. hundreds for a narrow task
  3. tens of thousands for broad behaviour
  4. inconsistent targets cap the ceiling
  5. duplicates only reweight one example

basics

~20 s

A few hundred to a few thousand consistent examples usually move a narrow, well-defined task; broad behaviour change needs tens of thousands. Coverage of the real input distribution and consistent labelling matter far more than raw row count.

solid answer

~50 s

Volume is the wrong first question — the useful one is how many distinct behaviours you are teaching and whether each is covered. For a narrow task with one output shape, such as turning a pet-owner's description into a triage urgency band plus a short rationale, a few hundred to a couple of thousand carefully written rows will usually get you most of the gain. Broad behaviour change across many task types is where you get into tens of thousands. What kills a fine-tune far more often than a small row count is **inconsistency**: if half the examples hedge and half give a flat verdict, the model learns to do both at random. Duplicating rows to inflate the count adds nothing. Spend the next thousand rows on slices the set does not yet cover, not on more of the type you already have.

go deeper

for a junior

Be ready to say that a narrow task often needs only hundreds to a couple of thousand clean examples, and that consistent, correct target outputs matter more than the row count.

for a middle

Explain why the model imitates the variance in the targets, so mixed formats and label noise put a ceiling on quality, and why duplicated rows only reweight one example instead of adding information.

for a senior

Show how you would allocate a data budget: enumerate the input axes, find the empty slices, run a data-scaling curve, and argue for cleaning versus collecting based on where the errors actually sit.

for a principal

Own the framing that dataset size is a cost decision against alternatives — prompting, retrieval, a bigger base model — and that curation effort, not row count, is the line item that decides whether the programme pays for itself.

## What the row count is really measuring A supervised fine-tuning dataset is a set of pairs: an input the model will see in production, and the exact output you want it to produce. When someone asks "how much data do I need", the honest answer is that the count is a proxy for two things that actually determine the result — how many distinct behaviours you are asking the model to learn, and how well the examples cover the range of inputs it will meet. A thousand rows that all show the same behaviour on the same kind of input is one behaviour taught a thousand times. ## Rough volumes people actually work with - **Narrow, single-format tasks** — classify an inbound veterinary triage message into an urgency band and emit a fixed set of fields; rewrite an internal ticket into a standard summary. A few hundred to roughly two thousand clean rows is typically enough, and the curve often flattens well before the top of that range. - **Broader behaviour** — a whole assistant persona, multi-turn conversational style, many task types under one model. Tens of thousands is the honest order of magnitude, which is why teams reach for synthetic expansion rather than hand-writing at that scale. These are starting points, not laws. The number that matters is rows *per distinct slice*, and a set of 800 rows spread evenly over eight real input types is usually worth more than 5,000 rows concentrated on the two most common ones. ## Why consistency dominates volume A fine-tune imitates the distribution of the outputs it is shown, variance included. If some examples answer in two sentences and some in six, some open with a caveat and some do not, some emit the fields in one order and some in another, the model does not average those into a canonical form — it learns that all of them are acceptable and picks one at random per request. Label noise behaves the same way: a small fraction of wrong targets puts a ceiling on the quality the fine-tune can reach, and no amount of extra rows lifts that ceiling. This is why a written spec for the target output, applied by whoever (or whatever) produces the rows, is worth more effort than another data source. If the rows are produced by several people or several generation prompts, review a sample specifically for cross-row consistency, not just per-row correctness. ## Coverage: matching the real input distribution Before counting rows, enumerate the axes your inputs vary along, then check the set against them. For a triage assistant that might be species, body system, urgency, and how much detail the owner gives. Empty cells in that grid are where the fine-tuned model will behave worst, and they are invisible if you only look at aggregate counts. The cheapest quality win in most projects is filling the sparse cells rather than growing the dense ones. Coverage should reflect production traffic, but not slavishly: rare-but-costly inputs deserve more representation than their natural frequency, because the loss you care about is not uniform across cases. ## Duplication is not data Copying a row to reach a bigger number simply reweights that example — the model sees the same gradient signal repeatedly and drifts toward that row's specifics. Near-duplicates do the same thing more subtly, which is why deduplication is part of preparing the set rather than a nicety. "8,000 rows" that dedupe down to 3,000 distinct behaviours is a 3,000-row dataset with a skew problem. ## Deciding where the next budget goes A practical technique is a data-scaling curve: train on 25%, 50% and 100% of the set and compare quality at each point. If the curve is still climbing steeply at the top, more of the same data is a reasonable buy. If it has flattened, extra volume is wasted and the remaining errors are concentrated in slices the data does not represent — go and write rows for those, or make the existing rows more consistent. ## When more data is not the answer at all If the model's failures are about *knowing* things — facts it was never trained on, information that changes weekly — a bigger fine-tuning set does not fix it, and training on unfamiliar facts tends to increase confident errors. Fine-tuning reliably teaches form, style and task procedure; supplying knowledge is a retrieval problem. Recognising that distinction in an interview is worth more than any specific row count.

  • How would you tell whether more rows are still buying you anything?
    Build a data-scaling curve: train on 25%, 50% and 100% of the set and compare quality at each point. A curve still climbing steeply at the top says more of the same data is worth buying. A flat curve says the remaining errors live in slices the data does not cover, so the next budget should go to new slices or to cleaning inconsistent targets rather than to volume.
  • You only have 40 real examples. What do you do with them?
    Do not train on 40 rows and hope. Use them as seeds — they define the shape and the voice — and expand synthetically against a coverage plan, keeping a handful of the originals untouched as a human-written check set. If the task is narrow enough that prompting already half-works, it is often cheaper to stay with prompting until you have enough real traffic to curate from.
  • Does the same volume guidance hold for adapter-based tuning as for full-parameter training?
    Broadly yes on the data side: the size and coverage of the set is driven by how many behaviours you are teaching, not by which parameters you touch. What changes is the failure mode at the top end — a small adapter has limited capacity, so a very large and diverse dataset can exceed what it can absorb. That is a capacity question, though, not a reason to curate differently.

saying these in an interview costs you the question

  • More rows always make a fine-tune better
  • Scraping 100k noisy rows beats 500 curated ones
  • You need millions of examples, like pretraining
  • Duplicating examples counts as extra data
  • Volume compensates for inconsistent target style
  • Adding rows will teach the model missing facts

context

open as a page

When is fine-tuning the right call instead of a better prompt or retrieval?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Fine-tune only after prompting and retrieval have been tried and still miss. The cases that justify it are behavioural: an output format or house style hard to describe in words, or a stable high-volume task where per-call cost is the binding constraint.

open as a page

How do you decontaminate fine-tuning training data against the evaluation set?

level: middleimportance: must knowfreq 48%

basics

~20 s

Decontamination removes training rows that overlap the evaluation cases. Normalise the text, drop any training row sharing a long n-gram with an eval item, then catch paraphrases with embedding similarity — and split before generating, so synthetic rows never straddle the boundary.

open as a page

What baselines must a fine-tuned model beat before you ship it?

level: middleimportance: must knowfreq 72%

basics

~20 s

At minimum the same base model prompted properly - a strong system prompt and a few-shot variant - plus whatever runs in production today. Score every arm on one held-out set with identical decoding settings, or the improvement is unattributable.

open as a page

Why is held-out loss a poor yardstick for whether a fine-tune helped?

level: middleimportance: must knowfreq 64%

basics

~20 s

Held-out loss scores the token-level likelihood of one reference wording, so it penalises correct answers phrased differently and rewards imitating training style. It tells you the run is healthy, not that the task got better - a task metric or human preference decides that.

open as a page

How does gradient accumulation reach a target effective batch size on one GPU?

level: middleimportance: must knowfreq 52%

basics

~20 s

Gradient accumulation runs several small micro-batches, sums their gradients, and steps the optimizer only once. Effective batch equals micro-batch size times accumulation steps times device count, so a large batch costs time instead of memory.

open as a page

Why does a LoRA fine-tune need a higher learning rate than full fine-tuning?

level: middleimportance: must knowfreq 58%

basics

~20 s

A LoRA run trains only a small adapter, so each step must move far fewer parameters and needs a bigger step size. The working rule is roughly ten times the full fine-tuning rate — about 1e-4 versus 1e-5.

open as a page

In LoRA fine-tuning, what does the low-rank update train while the base stays frozen?

level: middleimportance: must knowfreq 80%

basics

~20 s

LoRA leaves every pretrained weight frozen and trains two small matrices per targeted layer, A and B. Their product BA is added to the frozen weight as a rank-r update, so only those matrices hold gradients and optimizer state.

open as a page

What does QLoRA change about a LoRA run, and why does it fit on one GPU?

level: middleimportance: must knowfreq 62%

basics

~20 s

QLoRA stores the frozen base model in 4-bit instead of 16-bit and trains 16-bit adapters on top of it. Since the base never updates, its low precision is static; cutting weight memory roughly fourfold is what brings a multi-GPU run onto a single card.

open as a page

In supervised fine-tuning, why is the loss computed on completion tokens only?

level: middleimportance: must knowfreq 70%

basics

~20 s

Masking restricts gradients to the assistant's answer, so the model learns to produce completions rather than reproduce prompts. Train on the whole sequence and it also learns to write user turns and system text, wasting capacity and corrupting behaviour.

open as a page

What does "fine-tuning is for form, not facts" mean for an LLM?

level: middleimportance: must knowfreq 64%

basics

~20 s

Fine-tuning reliably changes how a model behaves — its format, tone and judgement — but not what it knows. Training on facts the base model never learned mostly teaches it to assert unfamiliar claims confidently, which raises the hallucination rate.

open as a page

Why do synthetically expanded fine-tuning sets collapse in diversity, and how do you prevent it?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Free-running generation keeps sampling the generator's favourite mode, so thousands of rows re-tell a handful of scenarios. Prevent it by generating against a structured grid of case attributes with a quota per cell, and measure diversity by clustering row embeddings before you train.

open as a page

How do you tell that a fine-tune has overfit its training set?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Two families of signal: curve divergence, where training loss keeps falling while held-out loss turns upward, and output symptoms - answers reciting training phrasings verbatim, and a sharp quality drop when the same question is reworded.

open as a page

After SFT on one narrow task, general answers degrade — what happened and what do you do?

level: seniorimportance: must knowfreq 58%

basics

~20 s

That is catastrophic forgetting: gradients from a narrow dataset overwrite general capability, and often loosen refusal behaviour too. The standard mitigation is mixing general instruction and safety examples back into the training set, plus tracking a capability and safety baseline against the base model.

open as a page

What do warmup and cosine decay do in a fine-tuning learning-rate schedule?

level: juniorimportance: should knowfreq 46%

basics

~20 s

Warmup ramps the learning rate from near zero up to its peak over the first few percent of steps so early updates cannot destabilise the model. Cosine decay then lowers it smoothly back toward zero so late steps refine instead of overwrite.

open as a page

What does the alpha parameter control in a LoRA adapter, relative to the rank?

level: middleimportance: should knowfreq 52%

basics

~20 s

Alpha is a scaling constant: a LoRA adapter's output is multiplied by alpha divided by rank before it is added to the frozen layer. It sets how strongly the learned update is applied, and dividing by rank keeps that strength stable as rank changes.

open as a page

In SFT, why must training data use the model's own chat template?

level: middleimportance: should knowfreq 55%

basics

~20 s

Serving renders every request with the model's specific role markers and end-of-turn tokens. Training on a hand-rolled format teaches a different surface form, so the model meets an unfamiliar prompt shape at inference and quality drops quietly, with no error anywhere.

open as a page

How do you validate an LLM judge used to filter synthetic fine-tuning data?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Treat the filter as a classifier and measure it. Hand-label a few hundred generated rows with a domain expert, then compute the judge's precision and recall against those labels at the exact threshold you plan to ship, and re-measure whenever the generator or the rubric changes.

open as a page

Why does a randomly shuffled held-out split overstate a fine-tune's gain?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A random shuffle scatters near-duplicates and repeated entities across both sides, so the held-out set contains items almost identical to training items. That measures interpolation within the training distribution, not the deployment condition of genuinely new cases.

open as a page

How do you choose the number of epochs for a small fine-tuning dataset?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Start in the 1–3 epoch range: one pass often underfits a new style, three is the usual sweet spot, and more begins memorising a small set. Evaluate on a held-out split at a fixed step interval and keep the best checkpoint rather than the last.

open as a page

How do you choose the LoRA rank for a fine-tuning run on a 7B base model?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Match rank to how much the model must absorb. A small rank around 8 carries a fixed response format or tone; tens to low hundreds are needed to absorb a domain's taxonomy or a post-training-scale dataset. Undersizing plateaus the loss; oversizing mainly costs memory.

open as a page

Which layers should LoRA adapters target, and why is attention-only a weak default?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Target every linear projection, not just attention. Adapting the feed-forward up, gate and down projections alongside query, key, value and output matters because the MLP holds most of a transformer's parameters. Attention-only adapters underperform even at matched trainable-parameter count.

open as a page

When does distilling a frontier model into a small tuned model beat prompting it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Distillation pays when quality is already acceptable and cost or latency is the binding constraint: a stable, high-volume task where a frontier model's per-call price dominates. Its outputs become training data for a small model that serves the routine tier.

open as a page

When does reinforcement fine-tuning beat prompting on a task with an automatic grader?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Reinforcement fine-tuning fits when correctness is machine-checkable and the model already succeeds sometimes. You supply problems and a grader instead of gold answers, and training reinforces the model's own successful attempts — useful when you have verifiable outcomes but no written solutions.

open as a page

What mix of instruction pairs, preference pairs and distilled traces belongs in a fine-tuning set?

level: principalimportance: should knowfreq 34%

basics

~20 s

Instruction pairs teach the task's form and are the bulk of the data. Preference pairs come afterwards, in far smaller volume, for qualities expressible only as a comparison. Distilled teacher traces supply reasoning you cannot afford to write by hand.

open as a page

How do you decide if a fine-tune's task gain outweighs its regressions?

level: principalimportance: should knowfreq 36%

basics

~20 s

Measure what you never trained on - general instruction following, format compliance, refusal and safety behaviour, multi-turn coherence, a broad knowledge check - and set the acceptable regression budget before the run, so the decision is a pre-agreed threshold rather than a rationalisation afterwards.

open as a page

Serving 200 per-tenant LoRA adapters on one base model: merge them or swap them?

level: principalimportance: should knowfreq 34%

basics

~20 s

Swap by default. Keeping adapters separate means one resident base serves every tenant, with small per-tenant files loaded on demand; merging bakes an adapter into the weights and produces a full-size model per tenant, which does not scale past a handful of dedicated deployments.

open as a page

When would you still run full-parameter SFT instead of an adapter-based one?

level: principalimportance: should knowfreq 40%

basics

~20 s

Rarely, as of mid-2026: a well-configured adapter matches full fine-tuning on typical SFT data at a fraction of the compute. Full-parameter updates earn their cost only when the data volume exceeds adapter capacity or the target behaviour is a large distribution shift.

open as a page

When does a fine-tune's maintenance bill outweigh the quality it buys?

level: principalimportance: should knowfreq 42%

basics

~20 s

A fine-tune is a standing obligation: dataset, eval suite, serving path, and a retrain decision every time a stronger base model ships. It stops paying when the requirement moves faster than you can retrain, or when prompting the current base already matches it.

open as a page

Why does packing several SFT examples into one sequence risk cross-example contamination?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Packing concatenates unrelated examples into one full-length sequence to avoid padding waste. Without boundary-aware attention masking and position resets, tokens of the third example attend to the first, so the model learns to condition answers on irrelevant preceding text.

open as a page