skip to content

How does clipping gradients by global norm differ from clipping each gradient element by value?

level: middleimportance: must knowfreq 62%

answer

  1. one scalar versus many independent clamps
  2. which one is a pure rescale
  3. direction preserved or direction rotated
  4. (10, 1) clamped at one becomes (1, 1)

basics

~20 s

Global-norm clipping rescales the whole gradient by a single factor once its norm passes a threshold, so the update direction is unchanged. Elementwise value clipping clamps each component on its own, which rotates the update toward the diagonal.

solid answer

~50 s

Global-norm clipping first computes one number: the Euclidean norm of the gradient with every parameter tensor concatenated into a single vector. If that norm exceeds the threshold `c`, every component is multiplied by the same factor `c / norm`; otherwise nothing happens. Because it is one positive scalar applied to all coordinates, the direction of the update is preserved exactly and only its length is capped. Elementwise value clipping instead clamps each component independently into `[-v, v]`. Only the components that were already large get touched, so the large ones shrink while the small ones stay put, and the resulting vector points somewhere else. On a gradient of `(10, 1)`, value clipping at 1 gives `(1, 1)` — a step 45 degrees off the true descent direction — while global-norm clipping at 1 gives roughly `(0.995, 0.0995)`, the same heading at a safe length. That is why global-norm clipping is the default choice.

code

python · 17 lines
python
import math

g = [10.0, 1.0]        # gradient of a two-parameter model
clip = 1.0

norm = math.sqrt(sum(x * x for x in g))
by_norm = [x * (clip / norm) for x in g] if norm > clip else list(g)
by_value = [max(-clip, min(clip, x)) for x in g]

def heading(v):        # degrees away from the first axis
    return round(math.degrees(math.atan2(v[1], v[0])), 2)

print(round(norm, 3))                    # 10.05
print([round(x, 4) for x in by_norm])    # [0.995, 0.0995]
print([round(x, 4) for x in by_value])   # [1.0, 1.0]
print(heading(g), heading(by_norm), heading(by_value))
# 5.71 5.71 45.0

go deeper

for a junior

Know that gradient clipping caps how large a single training step can be, and that the common form compares one global gradient norm against a threshold rather than looking at entries one by one.

for a middle

Be ready to write both rules out: a shared factor c/norm applied to everything versus an independent clamp into [-v, v], and to explain why only the first leaves the update direction untouched.

for a senior

Show that you know the norm has to be reduced across all tensors before anything is scaled, that the threshold does not transfer between models, and when a per-group clip beats one global norm.

for a principal

Own the argument that clipping trades a small bias in the average update for tail-event stability, and be able to say when that trade stops being worth it for a team's training stack.

## What clipping is trying to do During training the gradient of the loss with respect to the parameters occasionally comes back enormous — orders of magnitude bigger than a typical step. Taking that step at the usual learning rate throws the parameters far outside the region where the current loss surface is anything like linear, and the run either diverges or lands somewhere much worse. Clipping is a guard placed on the assembled gradient before it is handed to the update rule: it puts a ceiling on how big a single step may be, while leaving ordinary steps untouched. There are two families, and they differ in what exactly they bound. ## Clipping by global norm Treat every parameter tensor in the model as one long vector `g`. Compute one scalar, the Euclidean norm ``` norm = sqrt(sum over all parameters of g_i^2) ``` and compare it to a threshold `c`. If `norm <= c` the gradient passes through untouched. If `norm > c`, every component is multiplied by the same factor: ``` g <- g * (c / norm) ``` After the rescale the norm is exactly `c`. The crucial property is that `c / norm` is a single positive scalar applied to all coordinates, so `g` after clipping is a positive multiple of `g` before it. In the full parameter space, the clipped vector points in exactly the same direction; only its length changed. The step is shortened, never redirected. That matches the intent: the direction computed by backpropagation is the information you want, and it is the magnitude that is untrustworthy on a rare outlier batch. Note also that the operation is one-sided. It only ever shrinks. A tiny gradient is never inflated up to `c`, so clipping is silent on the vanishing side of the problem. ## Clipping elementwise by value The alternative clamps each coordinate independently against a value `v`: ``` g_i <- max(-v, min(v, g_i)) ``` Every coordinate is now bounded, which does cap the step, but the components are not scaled by a common factor. Components already inside `[-v, v]` are untouched; only the ones sticking out are pulled in. The ratios between coordinates change, and so does the direction. The two-parameter case makes it concrete. Take `g = (10, 1)`, a gradient that wants to move mostly along the first parameter — about 5.7 degrees off that axis. Clamp elementwise at 1 and you get `(1, 1)`: a step at 45 degrees, giving the second parameter exactly as much movement as the first even though the loss said it mattered a tenth as much. Rescale the same gradient by global norm to `c = 1` and you get `(0.995, 0.0995)`, still 5.7 degrees off the axis. In the limit where many coordinates exceed `v`, value clipping degenerates toward a sign-like update: every saturated coordinate moves by the same amount regardless of how strongly the loss wanted it to. ## Which to use Global-norm clipping is the default in practice, for exactly the reason above: it is the version that respects what backpropagation computed. It has one cost — computing the global norm requires reducing over every parameter tensor before any of them can be scaled, which means a full pass over the gradients and, on a multi-device run, a synchronization point. Value clipping is cheaper and purely local: each tensor can be clamped without knowing anything about the others. It is a reasonable choice when you specifically want a hard per-coordinate bound — for instance to keep any single entry from reaching a magnitude that overflows a low-precision number format — and you accept the directional distortion as the price. It is a poor choice as a general stability knob. ## Two caveats worth stating First, the threshold is not transferable. The norm's typical scale depends on the model size, the loss reduction (sum versus mean), the batch size and the parameterization; a value that clips the tail on one model may clip everything or nothing on another. Second, even global-norm clipping is not free. It is a nonlinear function of the minibatch gradient, so the average clipped gradient is not simply a rescaled copy of the true full-data gradient. When clipping fires on a small tail of steps this bias is negligible and the stability is worth it. When it fires on most steps you are no longer following an unbiased descent direction of your stated loss, and the threshold is the thing to fix.

  • Does the norm have to be computed over every parameter tensor at once, or can it be per group?
    It can be per group, and sometimes should be. With one global norm, a single large tensor — say an embedding table whose gradient spikes — dominates the sum of squares, so when it spikes the rescale factor shrinks every other tensor's gradient too, quietly cutting the effective learning rate for the rest of the network. Clipping each group against its own threshold isolates that, but you lose direction preservation across groups: the relative scaling between them now changes.
  • If elementwise value clipping distorts the direction, why does anyone use it?
    Because it is local and cheap. Each tensor can be clamped on its own with no reduction over the whole model and no synchronization point, and it gives a hard guarantee that no single entry exceeds a chosen magnitude — useful when a low-precision number format would overflow. It is a bound on individual coordinates, not a stability policy, and you accept the rotated update as its cost.
  • Does global-norm clipping help with vanishing gradients as well?
    No. The rule only fires when the norm exceeds the threshold and only ever multiplies by a factor below one, so it can shorten a step but never lengthen one. A gradient that has shrunk toward zero passes through untouched. Vanishing needs a structural fix — initialization scale, activation choice, normalization or skip connections — not a clip.

Norm clipping is turning down one master volume knob; value clipping is capping each channel separately, which changes the mix, not just the loudness.

saying these in an interview costs you the question

  • Says both rules shrink the gradient in the same way
  • Claims global-norm clipping changes the update direction
  • Says value clipping preserves direction because it only touches big entries
  • Confuses clipping the gradient with clipping the weights
  • Thinks clipping also scales small gradients up to the threshold

context