How does a central finite-difference check verify a hand-derived analytic gradient?
answer
- measure the slope, do not trust it
- perturb one coordinate in both directions
- difference of two losses over 2h
- score by relative, not absolute, gap
- double precision, h near 1e-5
basics
~20 sPerturb one parameter by plus and minus a small h, recompute the scalar loss each time, and compare (L_plus - L_minus)/(2h) with the analytic gradient for that coordinate. Score the agreement by relative error, not by raw difference.
solid answer
~50 sTreat the loss as a scalar function of one parameter at a time. For coordinate i, set the parameter to `w_i + h`, evaluate the loss, set it to `w_i - h`, evaluate again, and take the central difference `(L_plus - L_minus) / (2h)` as a numerical estimate of `dL/dw_i`. Compare it with what the backward pass reports using relative error, `|a - n| / max(1e-12, |a| + |n|)`, so one threshold works whether the gradient is 1e-4 or 1e4. With `h` around 1e-5 and the loss evaluated in double precision, a correctly derived layer -- say a hand-written attention-scoring layer -- lands below 1e-7, while 1e-3 or worse means the algebra is wrong. Everything stochastic has to be frozen so both evaluations see the same function, and because each coordinate costs two full forward passes you check a random handful on a deliberately tiny model, never every parameter of a real one.
code
python · 18 linesimport math
def loss(w): # toy scalar loss: sum of tanh(x)**2
return sum(math.tanh(x) ** 2 for x in w)
def analytic_grad(w): # d/dx tanh(x)**2 = 2*tanh(x)*(1 - tanh(x)**2)
return [2 * math.tanh(x) * (1 - math.tanh(x) ** 2) for x in w]
w = [0.7, -1.3, 0.05]
h = 1e-5
g = analytic_grad(w)
for i in range(len(w)):
plus = list(w); plus[i] += h
minus = list(w); minus[i] -= h
num = (loss(plus) - loss(minus)) / (2 * h)
rel = abs(num - g[i]) / max(1e-12, abs(num) + abs(g[i]))
print(i, "analytic=%.9f" % g[i], "numeric=%.9f" % num, "rel=%.2e" % rel)go deeper
Be ready to state what the check compares: a slope estimated from two extra loss evaluations, against the gradient the backward pass reported for that same parameter.
An interviewer expects the formula (L_plus - L_minus)/(2h), why the divisor is 2h rather than h, and why the pass criterion is a relative error rather than an absolute one.
Show you could run this on real code: freeze every random draw, use double precision, shrink the model to a few dozen parameters, and sample coordinates instead of sweeping all of them.
Own what the check actually certifies -- agreement between forward and backward code, not correctness of the maths the forward pass implements -- and say what else your verification strategy needs to cover the rest.
## Why anyone runs this check A backward pass built by composing primitives that are already trusted rarely needs verification. The moment a human writes a derivative by hand -- a custom layer, a custom loss, a fused operation written for speed -- the gradient becomes a piece of algebra that can be quietly wrong. Quietly is the important word. A wrong gradient almost never raises an error. It produces a model that trains more slowly than it should, plateaus early, or converges somewhere subtly worse, and those symptoms are indistinguishable from a dozen other causes. A finite-difference check converts the question into an empirical measurement. The gradient is the slope of the loss with respect to a parameter, so measure the slope directly and see whether the analytic value agrees. ## The procedure Freeze everything except a single scalar parameter `w_i`, so the loss becomes a one-dimensional function of that number. Then: 1. Set `w_i <- w_i + h`, run the forward pass, record `L_plus`. 2. Set `w_i <- w_i - h`, run the forward pass, record `L_minus`. 3. Restore `w_i` and compute `n_i = (L_plus - L_minus) / (2h)`. 4. Compare `n_i` against `a_i`, the value the backward pass produced for that same coordinate. The divisor is `2h`, not `h`, because the two probe points are `2h` apart. Getting that factor wrong produces an estimate exactly twice the truth, which reads as a broken derivation and sends you hunting in the wrong file. ## Why the two-sided form Expanding the loss around the current point gives `L(w + h) = L + h*g + (h^2/2)*L'' + (h^3/6)*L''' + ...` and `L(w - h) = L - h*g + (h^2/2)*L'' - (h^3/6)*L''' + ...`. Subtracting cancels every even-order term, including the second-derivative term, so `(L_plus - L_minus)/(2h) = g + O(h^2)`. The one-sided form `(L(w + h) - L(w))/h` keeps the second-derivative term and is only `O(h)` accurate: at the same step size it is typically several orders of magnitude further from the truth, and it costs only one evaluation less. Two-sided is the default for a reason. ## Relative error, not absolute Gradients in one model span many orders of magnitude -- an output-layer bias may see 1e1 while a deep weight sees 1e-6. An absolute gap of 1e-4 is catastrophic for the second and irrelevant for the first, so an absolute criterion needs a different threshold per coordinate. The scale-free version is `rel = |a - n| / max(1e-12, |a| + |n|)` The guard in the denominator matters: when both values are essentially zero, dividing two tiny noisy numbers by each other manufactures a huge ratio out of nothing. With a guard, near-zero coordinates simply pass. As a rule of thumb with a central difference at `h = 1e-5` in double precision: below 1e-7 the derivation is right; 1e-5 to 1e-7 deserves a look and is often explained by non-smooth operations; 1e-3 and above is a bug. These are conventions rather than laws -- deeper models accumulate more arithmetic and float the achievable floor upward -- so calibrate them on a case you know is correct before wiring them into a test. ## What must be held fixed The check is only meaningful if `L_plus` and `L_minus` are two evaluations of the *same* function. - **Randomness.** Any sampling inside the loss -- a dropout mask, a random augmentation, a sampled latent -- must be identical across both evaluations, or the difference reflects the change in randomness rather than the change in the parameter. Reuse a single fixed draw, or evaluate the deterministic path. - **The data.** The same examples, in the same order, for both evaluations. - **Extra loss terms.** If a weight penalty contributes to the loss, it must also contribute to the analytic gradient you compare against. Include both sides, or check them separately. - **Precision.** Run the loss in double precision. Single precision raises the noise floor so far that a correct gradient cannot reach a 1e-7 relative error. ## Cost, and what that forces Every coordinate costs two full forward passes. A million-parameter model would need two million of them, which is why the check is never run on the production model. Build a miniature version -- a handful of channels, a batch of two, a few dozen parameters -- so the whole sweep runs in a second, and on anything larger sample a random subset of coordinates. Systematic algebra errors corrupt many coordinates at once, so a sample finds them; that is what makes sampling defensible rather than lazy. ## What a pass actually certifies It certifies that the backward code correctly differentiates the forward code as written. It says nothing about whether the forward code computes the objective you intended. A forward pass implementing the wrong formula will be differentiated faithfully and pass with a relative error of 1e-11. Closing that gap takes a separate test that compares the forward output against a value worked out independently.
- What do you do when both the analytic and the numerical value for a coordinate are essentially zero?Relative error divides by their magnitudes, so two near-zero numbers can produce an enormous ratio out of pure noise. Guard the denominator with a small constant -- `max(1e-12, |a| + |n|)` -- and treat the coordinate as passing when both values sit far below the scale of the gradients elsewhere in the layer. If an entire layer reports near-zero gradients, that is a separate finding worth chasing on its own.
- Your network uses dropout. What breaks in the check, and how do you fix it?The two perturbed evaluations draw different masks, so you are differencing two different functions; that mask noise then gets divided by a tiny `2h` and swamps the signal. Make the loss deterministic for the check: reuse one fixed mask across both evaluations, or verify the deterministic path with dropout inactive. The same applies to random augmentation and to any sampling inside the loss.
- Why check the gradient of the scalar loss rather than of a layer's raw output?Finite differences approximate the derivative of a scalar. A layer's output is a vector, so its derivative is a Jacobian and probing it would cost passes proportional to the number of output elements as well as inputs. Reducing the output to a scalar -- even with an arbitrary fixed random projection -- gives one number to difference while still exercising the whole backward path through the layer.
You are told a road climbs at a certain gradient. Instead of trusting the sign, you walk one metre forward and one metre back, measure the height at both points, and divide the height difference by the two metres you covered.
saying these in an interview costs you the question
- Comparing absolute difference, so large gradients always look broken
- Using a one-sided quotient and expecting 1e-7 agreement
- Perturbing every parameter at once instead of one coordinate
- Letting a fresh random mask be drawn for each loss evaluation
- Treating a green check as proof the forward maths is right
- Dividing by h instead of 2h and chasing a factor-of-two bug