skip to content

Gradient Pathologies

What goes wrong once the gradient is a long product of layer Jacobians: it shrinks to nothing or blows up, and clipping or a finite-difference check is how you contain and verify it.

on this pageshow

explore

questions

10

How does a central finite-difference check verify a hand-derived analytic gradient?

level: middleimportance: must knowfreq 55%

answer

  1. measure the slope, do not trust it
  2. perturb one coordinate in both directions
  3. difference of two losses over 2h
  4. score by relative, not absolute, gap
  5. double precision, h near 1e-5

basics

~20 s

Perturb one parameter by plus and minus a small h, recompute the scalar loss each time, and compare (L_plus - L_minus)/(2h) with the analytic gradient for that coordinate. Score the agreement by relative error, not by raw difference.

solid answer

~50 s

Treat the loss as a scalar function of one parameter at a time. For coordinate i, set the parameter to `w_i + h`, evaluate the loss, set it to `w_i - h`, evaluate again, and take the central difference `(L_plus - L_minus) / (2h)` as a numerical estimate of `dL/dw_i`. Compare it with what the backward pass reports using relative error, `|a - n| / max(1e-12, |a| + |n|)`, so one threshold works whether the gradient is 1e-4 or 1e4. With `h` around 1e-5 and the loss evaluated in double precision, a correctly derived layer -- say a hand-written attention-scoring layer -- lands below 1e-7, while 1e-3 or worse means the algebra is wrong. Everything stochastic has to be frozen so both evaluations see the same function, and because each coordinate costs two full forward passes you check a random handful on a deliberately tiny model, never every parameter of a real one.

code

python · 18 lines
python
import math

def loss(w):                 # toy scalar loss: sum of tanh(x)**2
    return sum(math.tanh(x) ** 2 for x in w)

def analytic_grad(w):        # d/dx tanh(x)**2 = 2*tanh(x)*(1 - tanh(x)**2)
    return [2 * math.tanh(x) * (1 - math.tanh(x) ** 2) for x in w]

w = [0.7, -1.3, 0.05]
h = 1e-5
g = analytic_grad(w)

for i in range(len(w)):
    plus = list(w); plus[i] += h
    minus = list(w); minus[i] -= h
    num = (loss(plus) - loss(minus)) / (2 * h)
    rel = abs(num - g[i]) / max(1e-12, abs(num) + abs(g[i]))
    print(i, "analytic=%.9f" % g[i], "numeric=%.9f" % num, "rel=%.2e" % rel)

go deeper

for a junior

Be ready to state what the check compares: a slope estimated from two extra loss evaluations, against the gradient the backward pass reported for that same parameter.

for a middle

An interviewer expects the formula (L_plus - L_minus)/(2h), why the divisor is 2h rather than h, and why the pass criterion is a relative error rather than an absolute one.

for a senior

Show you could run this on real code: freeze every random draw, use double precision, shrink the model to a few dozen parameters, and sample coordinates instead of sweeping all of them.

for a principal

Own what the check actually certifies -- agreement between forward and backward code, not correctness of the maths the forward pass implements -- and say what else your verification strategy needs to cover the rest.

## Why anyone runs this check A backward pass built by composing primitives that are already trusted rarely needs verification. The moment a human writes a derivative by hand -- a custom layer, a custom loss, a fused operation written for speed -- the gradient becomes a piece of algebra that can be quietly wrong. Quietly is the important word. A wrong gradient almost never raises an error. It produces a model that trains more slowly than it should, plateaus early, or converges somewhere subtly worse, and those symptoms are indistinguishable from a dozen other causes. A finite-difference check converts the question into an empirical measurement. The gradient is the slope of the loss with respect to a parameter, so measure the slope directly and see whether the analytic value agrees. ## The procedure Freeze everything except a single scalar parameter `w_i`, so the loss becomes a one-dimensional function of that number. Then: 1. Set `w_i <- w_i + h`, run the forward pass, record `L_plus`. 2. Set `w_i <- w_i - h`, run the forward pass, record `L_minus`. 3. Restore `w_i` and compute `n_i = (L_plus - L_minus) / (2h)`. 4. Compare `n_i` against `a_i`, the value the backward pass produced for that same coordinate. The divisor is `2h`, not `h`, because the two probe points are `2h` apart. Getting that factor wrong produces an estimate exactly twice the truth, which reads as a broken derivation and sends you hunting in the wrong file. ## Why the two-sided form Expanding the loss around the current point gives `L(w + h) = L + h*g + (h^2/2)*L'' + (h^3/6)*L''' + ...` and `L(w - h) = L - h*g + (h^2/2)*L'' - (h^3/6)*L''' + ...`. Subtracting cancels every even-order term, including the second-derivative term, so `(L_plus - L_minus)/(2h) = g + O(h^2)`. The one-sided form `(L(w + h) - L(w))/h` keeps the second-derivative term and is only `O(h)` accurate: at the same step size it is typically several orders of magnitude further from the truth, and it costs only one evaluation less. Two-sided is the default for a reason. ## Relative error, not absolute Gradients in one model span many orders of magnitude -- an output-layer bias may see 1e1 while a deep weight sees 1e-6. An absolute gap of 1e-4 is catastrophic for the second and irrelevant for the first, so an absolute criterion needs a different threshold per coordinate. The scale-free version is `rel = |a - n| / max(1e-12, |a| + |n|)` The guard in the denominator matters: when both values are essentially zero, dividing two tiny noisy numbers by each other manufactures a huge ratio out of nothing. With a guard, near-zero coordinates simply pass. As a rule of thumb with a central difference at `h = 1e-5` in double precision: below 1e-7 the derivation is right; 1e-5 to 1e-7 deserves a look and is often explained by non-smooth operations; 1e-3 and above is a bug. These are conventions rather than laws -- deeper models accumulate more arithmetic and float the achievable floor upward -- so calibrate them on a case you know is correct before wiring them into a test. ## What must be held fixed The check is only meaningful if `L_plus` and `L_minus` are two evaluations of the *same* function. - **Randomness.** Any sampling inside the loss -- a dropout mask, a random augmentation, a sampled latent -- must be identical across both evaluations, or the difference reflects the change in randomness rather than the change in the parameter. Reuse a single fixed draw, or evaluate the deterministic path. - **The data.** The same examples, in the same order, for both evaluations. - **Extra loss terms.** If a weight penalty contributes to the loss, it must also contribute to the analytic gradient you compare against. Include both sides, or check them separately. - **Precision.** Run the loss in double precision. Single precision raises the noise floor so far that a correct gradient cannot reach a 1e-7 relative error. ## Cost, and what that forces Every coordinate costs two full forward passes. A million-parameter model would need two million of them, which is why the check is never run on the production model. Build a miniature version -- a handful of channels, a batch of two, a few dozen parameters -- so the whole sweep runs in a second, and on anything larger sample a random subset of coordinates. Systematic algebra errors corrupt many coordinates at once, so a sample finds them; that is what makes sampling defensible rather than lazy. ## What a pass actually certifies It certifies that the backward code correctly differentiates the forward code as written. It says nothing about whether the forward code computes the objective you intended. A forward pass implementing the wrong formula will be differentiated faithfully and pass with a relative error of 1e-11. Closing that gap takes a separate test that compares the forward output against a value worked out independently.

  • What do you do when both the analytic and the numerical value for a coordinate are essentially zero?
    Relative error divides by their magnitudes, so two near-zero numbers can produce an enormous ratio out of pure noise. Guard the denominator with a small constant -- `max(1e-12, |a| + |n|)` -- and treat the coordinate as passing when both values sit far below the scale of the gradients elsewhere in the layer. If an entire layer reports near-zero gradients, that is a separate finding worth chasing on its own.
  • Your network uses dropout. What breaks in the check, and how do you fix it?
    The two perturbed evaluations draw different masks, so you are differencing two different functions; that mask noise then gets divided by a tiny `2h` and swamps the signal. Make the loss deterministic for the check: reuse one fixed mask across both evaluations, or verify the deterministic path with dropout inactive. The same applies to random augmentation and to any sampling inside the loss.
  • Why check the gradient of the scalar loss rather than of a layer's raw output?
    Finite differences approximate the derivative of a scalar. A layer's output is a vector, so its derivative is a Jacobian and probing it would cost passes proportional to the number of output elements as well as inputs. Reducing the output to a scalar -- even with an arbitrary fixed random projection -- gives one number to difference while still exercising the whole backward path through the layer.

You are told a road climbs at a certain gradient. Instead of trusting the sign, you walk one metre forward and one metre back, measure the height at both points, and divide the height difference by the two metres you covered.

saying these in an interview costs you the question

  • Comparing absolute difference, so large gradients always look broken
  • Using a one-sided quotient and expecting 1e-7 agreement
  • Perturbing every parameter at once instead of one coordinate
  • Letting a fresh random mask be drawn for each loss evaluation
  • Treating a green check as proof the forward maths is right
  • Dividing by h instead of 2h and chasing a factor-of-two bug

context

open as a page

How does clipping gradients by global norm differ from clipping each gradient element by value?

level: middleimportance: must knowfreq 62%

basics

~20 s

Global-norm clipping rescales the whole gradient by a single factor once its norm passes a threshold, so the update direction is unchanged. Elementwise value clipping clamps each component on its own, which rotates the update toward the diagonal.

open as a page

Why can gradients vanish or explode as they backpropagate through 50 layers?

level: middleimportance: must knowfreq 82%

basics

~20 s

Backpropagation multiplies one Jacobian per layer, so the gradient reaching layer 1 is a product of 50 factors. If those factors average below one the gradient decays geometrically with depth; above one it grows the same way.

open as a page

What does a training loss that spikes and then becomes NaN say about gradients?

level: juniorimportance: should knowfreq 60%

basics

~20 s

It is the classic signature of exploding gradients: one update was large enough to push weights into a range where activations overflow to infinity, and infinity arithmetic then produces NaN. Once NaN reaches the weights, the run never recovers.

open as a page

Why does shrinking h in a finite-difference gradient check eventually make it worse?

level: middleimportance: should knowfreq 43%

basics

~20 s

Two errors pull in opposite directions. A central difference's truncation error falls like h squared, but subtracting two nearly identical losses and dividing by 2h amplifies round-off like 1/h. In double precision the total is smallest near h = 1e-5.

open as a page

Why does a gradient check fail on ReLU units whose pre-activation sits near zero?

level: seniorimportance: should knowfreq 34%

basics

~20 s

The loss is piecewise linear in that parameter with a corner between the two probe points, so the numerical secant blends both slopes while the backward pass reports one. The check fails though the code is correct.

open as a page

How do you pick a gradient-clipping threshold instead of inheriting a default of 1.0?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Measure before you choose. Record the global gradient norm for the first several hundred steps with the clip effectively off, look at the distribution, and set the threshold near its upper tail — around the 90th percentile — so ordinary steps pass and only outliers get rescaled.

open as a page

A 40-layer network has gradient norms of 1e-2 at the head and 1e-9 at the input — what do you conclude?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Seven orders of magnitude across the depth means the gradient is decaying geometrically layer by layer. Only the top few layers learn; the rest still hold their initial weights, acting as a fixed random feature map. Effective depth is about four.

open as a page

When is gradient clipping the wrong fix for a run whose gradients keep exploding?

level: principalimportance: should knowfreq 37%

basics

~20 s

Whenever the huge gradient has a findable cause. Clipping bounds the step but leaves the cause running, and because the run now survives, nobody investigates. It is a seatbelt for rare tail events, not a substitute for finding what produces them.

open as a page

When does a finite-difference gradient check earn a permanent place in a test suite?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Wherever a derivative was written by a human rather than composed from already-trusted primitives: a custom layer, a custom loss, a hand-optimised operation. Keep the test tiny, deterministic and in double precision so it runs in seconds and never flakes.

open as a page