skip to content

Where do you place stop-gradients in a model's graph, and what does each cut change?

level: principalimportance: nice to knowfreq 30%

answer

  1. identity forward, zero backward
  2. cuts a path, does not freeze weights
  3. which shortcut does the cut remove
  4. the cut redefines the objective
  5. loss curve will not reveal a misplaced one

basics

~20 s

A stop-gradient passes its input forward unchanged and returns zero gradient, so the optimizer treats that value as a constant. Put one wherever a model-computed quantity should act as data, not as something the loss can lower by changing it.

solid answer

~50 s

Mechanically a stop-gradient is the identity forward and zero backward: the value still shapes the loss, but no gradient flows into the branch that produced it. The decision rule I use is to ask, for each path, "if the optimizer could move this quantity, would lowering the loss that way be learning or a shortcut?" A self-paced curriculum makes this concrete: each sample's weight is computed from the model's own loss on that sample. Without a cut, the cheapest way to reduce the weighted loss is to shrink the weights on the hard samples rather than fit them — the weighting stops being a curriculum and becomes part of the objective. Cutting the weight's path fixes it. The thing to be honest about is that a cut *redefines the objective*: it is not an optimisation detail, and a misplaced one trains happily while descending a gradient you never intended.

go deeper

for a junior

Know the basic behaviour: the number passes forward unchanged, and no gradient travels back through that point.

for a middle

Be able to separate cutting a path from freezing parameters, and to give one case where a quantity the model itself computed must be treated as a constant inside the step.

for a senior

Expect to be handed a training setup and asked where the cut belongs. Justify each cut by naming the shortcut it removes, and describe the one-backward-pass test that pins which parameters receive gradient.

for a principal

Own the framing that every cut redefines the objective. Argue when a cut is a principled statement about the model and when it is an undocumented hack that leaves the team unable to say what is being minimised.

## What the operation is A stop-gradient (also called a detach or a gradient barrier) is a node with two faces. **Forward:** it is the identity — the value passes through untouched, and the loss you compute is exactly the loss you would have computed without it. **Backward:** it returns zero to its input regardless of the incoming adjoint. Everything upstream of the cut therefore sees that value as a constant, as if it had been supplied by the dataset rather than produced by the model. A useful side effect is that the branch feeding the cut needs no stored intermediates for the backward pass. ## Cutting a path is not freezing a parameter These are routinely confused and they are different operations. - **Freezing** removes parameters from the update. The frozen layer still *transmits* gradient: a frozen block in the middle of a network happily passes gradient to everything before it. - **A stop-gradient** severs a *path*. The parameters on the other side of the cut are not frozen — if another live path reaches them, they still learn from that path. A shared trunk with one head's path cut still trains from the other heads. So the question "which parameters should stop moving?" is answered by freezing; the question "which route may the loss travel?" is answered by a stop-gradient. ## A decision framework For each edge you are considering cutting, ask four questions. 1. **Is this quantity conceptually a constant?** Weights, temperatures, thresholds and statistics that are *computed from* the model but intended as coefficients, not as things to be optimised, belong behind a cut. Their role in the loss is to modulate it, not to be tuned by it. 2. **Does the open path admit a degenerate solution?** This is the strongest reason to cut. If there is a way to lower the loss by manipulating the quantity instead of improving the prediction, gradient descent will find it — reliably and quickly. Cutting removes that route from the search space. 3. **Is the path's gradient meaningful?** Some routes carry gradient that is real but noisy, expensive to compute, or of the wrong sign for the behaviour you want. Cutting is then a bias-for-variance trade you should be able to state out loud. 4. **Can I write down what the objective becomes?** If you cannot say what function is now being minimised, the cut is a hack, not a design. ## The worked case: a self-paced curriculum Suppose each sample's contribution is scaled by a weight derived from the model's own per-sample loss — easy samples get more weight early, hard ones later. The weighted loss is `sum_i w_i * l_i`, and `w_i` is a function of `l_i`. Leave the path open, and `dL/dtheta` picks up a term through `w_i` as well as through `l_i`. Since making `w_i` small on a high-loss sample reduces the product immediately, the optimizer learns to *downweight what it cannot fit* instead of fitting it. The curriculum has silently become part of the objective, and the failure mode looks like plausible training: the reported weighted loss falls beautifully while unweighted validation loss stalls. Wrapping `w_i` in a stop-gradient restores the intended semantics — the weights are recomputed each step from the current model, but within a step they are constants. ## Asymmetry and role assignment Cutting one branch of an otherwise symmetric pair is how you give the two branches different jobs: one becomes the thing being fitted, the other a reference that only moves for reasons outside the loss. Whenever you rely on that, be explicit that the reference's update rule is now a separate design decision, because the loss no longer defines it. ## The cost of getting it wrong — in both directions **Too few cuts** and you get an objective with a shortcut in it: it trains, the loss curve is smooth, and the model solves the wrong problem. **Too many cuts** and you starve a branch that should have learned; it also trains, just worse, and you will blame the architecture. Neither failure shows up as an error, which is what makes this a judgment topic rather than a mechanics one. Because the loss curve will not tell you, test it directly: construct a loss, run one backward pass, and assert the *set* of parameters with nonzero gradient is exactly what you intended. That assertion is cheap, deterministic, and catches both a cut that migrated during a refactor and one that was never there. ## What to say in an interview Name the mechanism in one line, then talk about objectives rather than implementation: every stop-gradient is a claim about which quantities the optimizer is permitted to manipulate, and a model with undocumented cuts is a model whose objective nobody can write down.

  • How does a stop-gradient differ from freezing a layer's parameters?
    Freezing excludes parameters from the update but leaves the path intact, so gradient still flows through that layer to everything upstream. A stop-gradient does the opposite: it severs the path, while the parameters behind it stay trainable and will still be updated by any other live path that reaches them. One controls updates, the other controls routes.
  • How would you test that a stop-gradient sits where you intended?
    Build a minimal loss, run a single backward pass, and assert on the set of parameters with nonzero gradient — not on the loss value. That assertion pins the graph topology, so a cut that disappears or migrates in a refactor fails the test immediately. Finite differences on a parameter behind the cut are a second check: the numerical derivative includes the path the cut removes, so the two should disagree in a known way.
  • What is the risk of scattering stop-gradients liberally through a model?
    Each cut removes a real learning signal. Over-cutting starves branches that should adapt, and the symptom is a model that trains smoothly to a mediocre optimum, which gets misdiagnosed as an architecture or capacity problem. It also makes the objective impossible to state, so the next person cannot reason about the model at all.

saying these in an interview costs you the question

  • Calls a stop-gradient purely a memory or speed optimisation
  • Says it freezes the parameters that produced the value
  • Thinks it changes the forward value or zeroes it
  • Assumes a misplaced cut will show up in the loss curve
  • Cannot state what objective is being minimised after the cut

context