skip to content

In layer-wise discriminative fine-tuning, why does the bottom block get a much smaller rate than the head?

level: seniorimportance: nice to knowfreq 34%

answer

  1. not every depth needs the same step
  2. generic at the bottom, specific on top
  3. rates rise geometrically toward the head
  4. one ratio hyperparameter, not one per layer
  5. freezing is the same idea at rate zero

basics

~20 s

Depth decides how much a layer must change. Bottom blocks hold generic features that transfer almost unchanged and should barely move, while upper blocks are task-specific and the head is random, so rates rise geometrically from bottom to top.

solid answer

~50 s

Discriminative fine-tuning gives each depth its own step size instead of one rate for the whole network, because the layers do not need equal amounts of change. Bottom blocks encode generic structure — edges and textures, or general token statistics — that the target task reuses nearly as is, so a large step there only risks damage. Upper blocks encode task-specific abstractions that genuinely have to shift, and the head is randomly initialised and must be learned outright. A common layout spans roughly 1e-5 at the bottom block up to about 1e-3 at the head, with a constant multiplicative ratio between adjacent groups rather than equal spacing. Keeping the ratio constant also keeps the search cheap: you tune a top rate and one ratio, not one rate per layer. Freezing is the limiting case of the same idea, a rate of zero on blocks you never want to move.

go deeper

for a junior

Recall the shape: in a fine-tune the layers nearest the output get the largest step sizes and the earliest layers get the smallest, because early layers already compute features the new task can reuse.

for a middle

Explain why the ladder is multiplicative rather than linear, and what each group's rate is doing: near-zero movement for generic features, real movement for task-specific blocks, largest of all for a head starting from random.

for a senior

Demonstrate that you can set and defend one. Anchor the top rate, derive the ratio from the span and group count, and say plainly when a single global rate is the better engineering call.

for a principal

Own the cost side. Every extra hyperparameter is validation-set budget spent, so decide whether the ladder belongs in your organisation's default fine-tuning recipe or stays a tool for the small-data, distant-domain cases where it actually earns out.

## The premise: layers differ in how much they must change A single fine-tuning rate assumes every depth needs the same amount of movement. Transfer learning makes that assumption obviously false. Representations in a pretrained network get more general the lower you go and more task-specific the higher you go. The earliest layers of a vision backbone compute edges, colours and simple textures; the earliest layers of a text encoder compute broad token and morphology statistics. Those things are true of the target task too, so the correct amount of change there is close to zero. The upper blocks compose those primitives into abstractions shaped by the *source* task, and those must be re-shaped. Above them sits a head that is random and has to be learned from scratch. Discriminative fine-tuning encodes that gradient of need directly into the step sizes: one rate per layer group, increasing with depth toward the output. ## The ladder and why it is geometric The usual arrangement spans roughly two orders of magnitude, something like 1e-5 on the bottom block rising to about 1e-3 at the head, with a constant *ratio* between adjacent groups rather than a constant difference. The reason to make it multiplicative is that what matters at each depth is the update size relative to that layer's own weight scale, not an absolute number of units in parameter space. A constant ratio means each group moves in proportion to the one below it, so the ladder is scale-free and keeps its shape whether you span four groups or twelve. The arithmetic follows from the span and the number of groups. Spreading 1e-5 to 1e-3 — a factor of 100 — across four groups means three gaps, so each adjacent pair differs by about 4.6x. Across more groups the per-step ratio shrinks: five groups gives four gaps and a ratio near 3.2. The published discriminative-fine-tuning recipe from ULMFiT used a fixed divisor of about 2.6 per group going downward, which lands in the same family. The precise number matters much less than the shape: strictly increasing toward the output, multiplicative, and anchored by a top rate you actually tuned. ## Why not just tune every layer separately Because the search cost explodes and the payoff does not. Parameterising the ladder as (top rate, ratio) gives you two knobs to tune and a monotone structure you can reason about; per-layer rates give you as many knobs as blocks, no structure, and enormous room to overfit the validation set with hyperparameters. The two-knob version also degrades gracefully: set the ratio to one and you are back to a single global rate, which is a sensible fallback rather than a broken configuration. ## Relationship to freezing Freezing a block is the same idea with the rate set to zero. That makes the two techniques points on one axis: freeze the bottom (hard zero), give it a tiny rate (soft, still nearly fixed), or give the whole network one rate (no discrimination at all). The soft version is often preferable when the target domain has drifted a little from the source, because even the generic layers benefit from a small correction that a hard freeze forbids. ## When the ladder is worth it, and when it is not The ladder pays off most when the target dataset is small, when the domain is somewhat distant, or when a single global rate has already shown you both failure modes — collapse at the high end of your sweep and no movement off the frozen baseline at the low end. That combination is exactly the sign that different depths want different step sizes and no single number satisfies both. It pays off least in three situations. First, when the dataset is large and the domain is close: everything can move at one modest rate and the extra structure buys little. Second, when the tuning budget is tight: two extra hyperparameters on a small validation set is a real risk of tuning noise. Third, when you are training with an adaptive optimizer, which already rescales updates per parameter and so partially flattens the differences the ladder is trying to create — the effect is real but smaller than it is with plain gradient descent. ## What to say when asked Good answers name the mechanism (rates increasing with depth), the reason (generality at the bottom, specificity and randomness at the top), the parameterisation (constant ratio, two hyperparameters), and the honest cost (more knobs, diminishing returns near domains, softened by adaptive optimizers). Weak answers describe the ladder as a universal best practice and cannot say what it would cost or when a single rate is fine.

  • What ratio between adjacent blocks would you start with, and how do you get it?
    Work backwards from the span and the group count. A 100x span from about 1e-5 at the bottom to 1e-3 at the head across four groups is three gaps, so roughly 4.6x per step; across five groups it drops to about 3.2x. Ratios in the two-to-five range are the normal territory. Tune the top rate first with the ratio fixed, since the top rate is what most affects the result.
  • How does a per-layer ladder relate to simply freezing the lower blocks?
    Freezing is the ladder with a rate of zero on those blocks, so they are the same technique at different settings. The soft version is usually better when the target domain has drifted slightly from the source, because generic layers can still take a small correction that a hard freeze forbids. Freezing wins on cost, since frozen blocks need no gradient computation.
  • When would you not bother with discriminative rates at all?
    When the target domain is close to the source and the dataset is large enough that one modest rate adapts everything safely, or when the tuning budget is small enough that two extra hyperparameters would mostly fit validation noise. Adaptive optimizers also rescale updates per parameter, which partly flattens the differences the ladder creates, so the marginal gain there is smaller than with plain gradient descent.

saying these in an interview costs you the question

  • Give the bottom layers the largest rate, they are furthest from the loss
  • Space the rates linearly across the blocks
  • Discriminative rates always beat a single global rate
  • Tune a separate independent rate for every layer
  • It is unrelated to freezing, a different technique entirely

context