skip to content

Fine-Tuning Step Sizes

Fine-tuning uses a far smaller step than pretraining, often lower for early layers than late ones, with a warmed-up random head and gradual unfreezing. Interviewers ask what one large step destroys.

on this pageshow

questions

3

Why is the fine-tuning learning rate for a pretrained network far smaller than its pretraining rate?

level: juniorimportance: must knowfreq 72%

answer

  1. you already start near a solution
  2. step size versus distance still to travel
  3. a random head sends noise backwards
  4. one big step erases the features
  5. compare against the frozen-feature baseline

basics

~20 s

Pretraining already puts the weights in a good region, so fine-tuning only needs small corrections. A large step moves every weight far enough to destroy the learned features, and a small target dataset cannot rebuild them.

solid answer

~50 s

The learning rate sets how far one minibatch may move the weights: `w <- w - lr * g`. Pretraining starts from random weights that carry no information, so it has to travel a long way and can afford big steps. Fine-tuning starts from weights whose *precise values* are the thing of value, so it wants a correction, not a journey. Worse, the new head is randomly initialised, so the first gradients flowing back into the backbone are large and point somewhere unrelated to the useful features. At a pretraining-sized rate like 1e-2 that first step can shift backbone weights by an amount comparable to their own magnitude and the features are gone: validation accuracy then starts *below* what a frozen backbone with a trained head reaches and never catches up, while the same run at 1e-4 passes that baseline in the first epoch. Typical backbone rates sit one to two orders of magnitude below the pretraining rate.

go deeper

for a junior

Be ready to state the rule and the reason in one breath: fine-tuning starts near a good solution, so it takes small steps, typically one to two orders of magnitude below the pretraining rate.

for a middle

Explain the mechanics: the rate multiplies the gradient in the weight update, the fresh head emits large early gradients, and a pretraining-sized step moves backbone weights by an amount comparable to their own scale.

for a senior

Show the diagnosis. Keep a frozen-backbone reference number, read a run that lands below it as feature destruction rather than slow learning, and recognise the opposite failure where the rate is so small the model never leaves the baseline.

for a principal

Own the policy question: what the default rate is for teams fine-tuning off your checkpoints, whether the frozen-feature baseline is a mandatory reference run, and how much tuning budget is worth spending per downstream task before a fixed recipe is good enough.

## What the learning rate actually controls A gradient step is `w <- w - lr * g`. The gradient direction comes from the data and the loss; the learning rate alone decides *how far* along it a single minibatch is allowed to drag the model. Two runs with identical gradients but rates differing by 100x travel 100 times as far per step. So the question "what rate?" is really "how far should one batch be allowed to move this model?", and the answer depends entirely on where the model already is. ## Pretraining and fine-tuning start in different places Pretraining begins at a random initialisation. No weight carries information; the model must cross a long distance in parameter space, and the gradients early on are large and broadly consistent across the dataset. Big steps are what make that trip finish in a feasible number of updates, so pretraining rates are comparatively large. Fine-tuning begins from a point that already computes useful features. Edge and texture detectors, or general token statistics, live in the *specific numeric values* of those weights and in the way each layer's output distribution matches what the next layer expects. You are not looking for a new solution; you are looking for a nearby one. The distance you need to travel is small, so the step size should be small. ## Why one oversized early step is destructive At step zero the classification head is random, so its predictions are essentially arbitrary and the loss is high. For softmax with cross-entropy the gradient at the logits is `dL/dz = p - y`, which is large when `p` is far from the label, and that large signal is backpropagated into every backbone layer through the head's random weights. The direction it induces in the backbone is close to noise with respect to the pretrained representation. Multiply that noisy direction by a pretraining-sized rate and each backbone weight moves by an amount comparable to its own scale. Every layer's output statistics shift, every downstream layer now receives inputs outside the distribution it was tuned for, and the composition that made the representation useful stops holding. The concrete signature: a fine-tune at 1e-2 on a pretrained backbone with a fresh head shows first-epoch validation accuracy *below* the frozen-feature baseline (backbone frozen, head trained) and it never recovers inside the epoch budget, while the identical run at 1e-4 clears that baseline in one epoch. ## Why the damage does not heal Generic features were built by a large, diverse corpus and an enormous number of updates. Rebuilding them needs roughly the same resources. A target set of a few thousand labelled examples has neither the diversity nor the step count, so the wrecked run is effectively training from scratch on far too little data, from an initialisation that is now arbitrary. More epochs on the same small set fit the training data and do not restore the representation. ## The baseline that tells you which failure you have Always keep the frozen-backbone number as a reference line. It is the score the pretrained features already deliver with no risk. From there, two diagnoses: - **Rate too large.** Validation dives below the baseline early, the training loss may be erratic or spike, and the gap never closes. Cut the backbone rate by 10x and re-run. - **Rate too small.** Validation sits *at* the baseline and barely moves, training loss is still descending steadily when the budget ends. The model is under-adapting: you are paying for a fine-tune and getting a frozen-feature model. Raise the rate or extend the budget. The second failure is the one people forget exists. "Smaller is safer" is only half true — a rate small enough to be harmless can also be small enough to be pointless, especially when the target domain is far from the source and the model genuinely needs to move. ## How much smaller, in practice There is no universal constant. A useful starting posture is one to two orders of magnitude below the rate the model was pretrained at, then adjust by watching the two diagnoses above. The right value also depends on the optimizer in use and on how far the target distribution sits from the source, so treat any quoted number as a starting point and let the frozen baseline and the training-loss curve arbitrate. ## The other levers on the same problem A small global rate is the blunt instrument. Two sharper ones attack the same danger: warming up the head with the backbone frozen so the random head never gets to push the backbone around, and giving different depths different rates so the generic bottom layers barely move while the task-specific top moves more. Both reduce how much you have to rely on a single tiny number.

  • A fine-tune ends below the frozen-feature baseline. What does that tell you?
    That the run destroyed more representation than it added, which almost always means the backbone step size was too large. The frozen baseline is the score the pretrained features give for free, so falling under it is not slow learning, it is damage. Drop the backbone rate by an order of magnitude and re-run; if the collapse happens in the first few hundred steps, suspect the untrained head driving those steps as well.
  • Why can't more epochs repair a backbone wrecked by an early oversized step?
    Because the generic features were produced by a large, diverse corpus and a very long schedule, and the target set offers neither. After the damage the run is effectively training from scratch on a few thousand examples from an arbitrary initialisation, so extra epochs mostly fit the training data. The cheap fix is to reload the pretrained weights and restart with a smaller rate, not to train longer.
  • Can the fine-tuning rate be too small, and how would you spot it?
    Yes. The symptom is validation accuracy stuck at roughly the frozen-feature baseline with the training loss still falling steadily when the epoch budget ends: the backbone never moved enough to adapt. That is common when the target domain is far from the source. Raise the backbone rate, or extend the schedule, and check whether the gap over the baseline opens up.

Pretraining is driving across the country, where big highway distances are the point. Fine-tuning is parking: the same speed that got you there will put you through the garage wall.

saying these in an interview costs you the question

  • A high rate only slows convergence, it cannot lose accuracy
  • Fine-tuning can never do worse than the frozen features
  • Any damage from a big step is repaired by training longer
  • Use the pretraining rate, the optimizer adapts it anyway
  • Smaller is always safer, there is no downside

context

open as a page

Why train a fine-tune's randomly initialised head with the pretrained backbone frozen first?

level: middleimportance: should knowfreq 52%

basics

~20 s

A fresh head predicts almost randomly, so its loss gradient is large and, seen from the backbone, close to noise. Freezing the backbone for the first epochs lets the head reach sensible outputs before any pretrained weight moves.

open as a page

In layer-wise discriminative fine-tuning, why does the bottom block get a much smaller rate than the head?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Depth decides how much a layer must change. Bottom blocks hold generic features that transfer almost unchanged and should barely move, while upper blocks are task-specific and the head is random, so rates rise geometrically from bottom to top.

open as a page