A 40-layer network has gradient norms of 1e-2 at the head and 1e-9 at the input — what do you conclude?
answer
- convert the ratio to per-layer gain
- seven decades over thirty-six layers
- effective depth, not real depth
- frozen prefix, trainable head
- smooth falloff versus a single cliff
basics
~20 sSeven orders of magnitude across the depth means the gradient is decaying geometrically layer by layer. Only the top few layers learn; the rest still hold their initial weights, acting as a fixed random feature map. Effective depth is about four.
solid answer
~50 sSpread the seven orders of magnitude over roughly 36 layers and you get a per-layer factor of about 0.6 — a compounding decay, not one broken layer. The consequence is that only the top handful of layers receive a gradient large enough to move them at any learning rate the head tolerates; the rest sit at their initial values, so the model is really a shallow network stacked on a fixed random projection of the input. That explains why it trains at all — a random projection preserves enough signal for the top layers to fit something — and why it plateaus below what a properly trained shallower model reaches. I would confirm it by diffing early-layer weights against the initial checkpoint, relative to their own scale, and by checking that the norm profile falls smoothly rather than dropping at one layer. A smooth geometric profile means compounding depth; a cliff at one layer means a single mis-scaled or saturated block.
go deeper
Be ready to say that a gradient many orders of magnitude smaller at the input than at the output means the early layers are not learning, and that this is the vanishing gradient problem.
Convert the end-to-end ratio into a per-layer factor and explain why roughly 0.6 per layer is harmless once and fatal thirty-six times over, and why the model still trains on the frozen layers' output.
Demonstrate the triage: smooth falloff versus a single-layer cliff, weight deltas against the initial checkpoint, and the contrast with a global scale problem that shows no depth trend at all.
Own the framing that effective depth, not nominal depth, is what a model costs and delivers, and decide what evidence a team must show before a depth increase is approved.
## Reading the number first The gradient norm at the head is `1e-2` and at the input `1e-9`: a ratio of `1e-7` accumulated over about 36 layers. Convert that to a per-layer factor: `10^(-7/36)` is about `0.64`. So every layer is passing back roughly two-thirds of what it received. No layer is pathological on its own — two-thirds is unremarkable — but composed 36 times it is fatal. This is the central diagnostic move: turn the end-to-end ratio into a per-layer gain, because the per-layer gain is the quantity a fix has to change. This pattern shows up in exactly the situation the numbers suggest: a 40-layer plain stack over raw, unnormalised inputs — sensor windows, say — where nothing in the design pins the per-layer gain near one. ## What the network actually is A gradient of `1e-9` at the input side does not mean the early layers learn slowly; it means they do not learn. To move them meaningfully you would need a learning rate seven orders of magnitude larger than the one the head can tolerate, and that learning rate would destroy the head on the first step. Those layers stay at their initial random values for the whole run. So the model in production is: a fixed random transformation of the input, followed by a four-layer trainable network. That is why it trains at all — a random projection of the input keeps a good deal of usable structure, and four trainable layers on top of it can fit a real function. It is also why it plateaus: the representation feeding the trainable part was never adapted to the task, and it was not chosen either, it is whatever the initialisation happened to produce. Two predictions follow that you can test. Adding more depth at the bottom will make things slightly worse, not better, because it lengthens the untrained prefix. And a deliberately shallow model of comparable trainable capacity, trained end to end on the same data, will usually beat it. ## Confirming the diagnosis Three checks separate this from the things it can be mistaken for. **Is the profile smooth or a cliff?** A gradient norm that falls by a near-constant ratio per layer is compounding decay: the per-layer gain is systematically below one across the whole stack. A profile that is flat and then drops by several orders of magnitude at a single layer is a different problem — one mis-scaled weight matrix, or one block whose units are saturated — and it is fixed locally rather than structurally. **Are the early weights actually frozen?** Compare early-layer weights against the initial checkpoint, measuring the change relative to the weights' own scale rather than in absolute terms. A relative change near zero after thousands of steps confirms the layers are inert. This is stronger evidence than the gradient snapshot, because it integrates over the whole run. **Is it depth, or a global scale problem?** A learning rate that is simply too small shrinks every layer's step by the same amount; the *depth trend* in the gradient norms disappears. If all layers show small gradients with no systematic falloff from head to input, the problem is at the head — a tiny loss gradient, a nearly converged model, a scale issue in the loss — not compounding through depth. The signature of this leaf's failure is specifically the systematic ordering by depth. A fourth thing worth ruling out is that the run has simply converged. Convergence gives small gradients *everywhere*, including at the head, and comes with a loss that has stopped improving from a good value. Here the head gradient is `1e-2`, seven orders of magnitude larger, so the top of the network is still learning hard. ## What not to do Raising the global learning rate to force the early layers to move is the reflex answer and it is wrong: it multiplies the head's step by the same amount, and the head is the part already at the edge of stability. Adding capacity by adding layers is also wrong for the same reason — capacity is not the binding constraint when most of the network is not being trained. And a small input-side gradient is not by itself evidence of an implementation bug; it is the expected arithmetic of a long product of sub-unit factors. ## Where the fixes live The per-layer gain is the product of the weight matrix scale and the activation derivative, so every real remedy attacks one of those factors: the scale the weights are initialised at, layers that renormalise activations, activation functions that do not shrink the derivative, and connection patterns that give the gradient a shorter path. Each is a substantial topic of its own, and the useful thing this diagnosis buys you is knowing *which* factor is out of range, so you are choosing among them on evidence rather than applying all of them at once.
- How do you distinguish this from a learning rate that is simply too small?By the depth trend. A too-small learning rate shrinks every layer's step equally, so gradient norms show no systematic ordering from head to input. Here they fall monotonically with depth by a near-constant ratio, which only a compounding per-layer factor produces. Raising the learning rate would also push the head into instability long before the input layers moved.
- Why does the model train and improve at all if most of it is frozen?Because the frozen layers still compute a random but information-preserving transformation of the input, and the trainable top layers fit on that representation. The loss genuinely falls, which is what makes this failure deceptive — it looks like a working model until you compare it against a shallower one trained end to end.
- What would a single sharp drop at one layer mean instead?A local problem rather than compounding depth: one weight matrix initialised or scaled far too small, or one block whose units are saturated so their derivatives are near zero. The rest of the profile would be flat. That is fixed at that layer, and it will not be helped by anything aimed at the stack as a whole.
- What would you expect if you added ten more layers at the bottom?Slightly worse results. The extra layers lengthen the untrained prefix and add more sub-unit factors to the product, pushing the input-side gradient lower still. Trainable capacity does not increase, because the new layers will not receive a usable gradient either.
It is a bucket brigade forty people long where each person passes on two-thirds of what they receive. Nobody is doing anything obviously wrong, and almost nothing arrives at the far end.
saying these in an interview costs you the question
- Raises the learning rate to force the early layers to move
- Concludes the model converged because the loss fell
- Adds more layers to increase capacity
- Reads a tiny input-side gradient as proof of a code bug
- Ignores whether the falloff is smooth or a single cliff
- Never checks whether the early weights actually changed