skip to content

Why does shrinking h in a finite-difference gradient check eventually make it worse?

level: middleimportance: should knowfreq 43%

answer

  1. two competing error sources
  2. Taylor remainder versus arithmetic noise
  3. nearly equal numbers get subtracted
  4. cancellation error grows like 1/h
  5. U-shaped curve, floor near 1e-5

basics

~20 s

Two errors pull in opposite directions. A central difference's truncation error falls like h squared, but subtracting two nearly identical losses and dividing by 2h amplifies round-off like 1/h. In double precision the total is smallest near h = 1e-5.

solid answer

~50 s

The total error in a numerical gradient is a sum of two terms. Truncation error comes from the Taylor remainder the difference quotient throws away; for the central form `(L_plus - L_minus)/(2h)` it scales like `h^2`, so it shrinks fast as `h` gets smaller. Round-off error comes from the arithmetic: `L_plus` and `L_minus` agree in more and more leading digits as `h` shrinks, so their difference keeps fewer and fewer significant digits, and dividing that eroded number by `2h` magnifies the damage like `1/h`. Total error therefore traces a U: `h = 1e-2` is wrong from curvature, `h = 1e-10` is wrong because the subtraction cancelled the signal away, and the floor sits near the cube root of double-precision machine epsilon, around 1e-5. That U-curve is also a diagnostic -- sweep `h` over several decades, and a correct derivation shows a clear dip while a wrong one stays flat and high.

go deeper

for a junior

Know that the step size is a real choice with a sweet spot, and that making it too small degrades the numerical estimate instead of sharpening it.

for a middle

Be able to name both error terms -- truncation shrinking like h squared, round-off growing like 1/h -- and say roughly where they balance in double precision.

for a senior

Demonstrate the diagnostic: sweep h across several decades and read the shape. A U-curve with a low floor exonerates the derivation; a flat, high curve indicts it.

for a principal

Treat this as the error budget the whole verification strategy rests on, and decide what precision and what tolerance make a green check trustworthy rather than merely green.

## Two errors, not one A numerical gradient is never exact, and the reason it is inexact changes completely depending on the step size. Understanding both halves is what separates a candidate who picked `1e-5` off a tutorial from one who can debug a check that will not pass. ### Truncation error: the maths you dropped Expand the loss about the current parameter value: `L(w + h) = L + h*g + (h^2/2)*L'' + (h^3/6)*L''' + ...` `L(w - h) = L - h*g + (h^2/2)*L'' - (h^3/6)*L''' + ...` Subtract and divide by `2h`. The constant and the second-derivative term cancel because they are even in `h`, leaving `(L_plus - L_minus)/(2h) = g + (h^2/6)*L''' + ...` So the central difference is second-order accurate: the error it inherits from ignoring curvature falls like `h^2`. Cut `h` by ten and this term drops by a hundred. A one-sided quotient `(L(w + h) - L(w))/h` cannot cancel the even terms and carries an error of `(h/2)*L''` -- first order, so cutting `h` by ten only cuts the error by ten. That analysis alone says: make `h` as small as you can. Which is exactly the wrong conclusion. ### Round-off error: the arithmetic you cannot escape Loss values are stored with finite precision. Each carries a relative error of roughly machine epsilon -- about 2.2e-16 for a double, about 1.2e-7 for a single. So `L_plus` and `L_minus` each come with an absolute uncertainty of about `eps * |L|`. As `h` shrinks the two values converge toward each other, but their uncertainties do not shrink at all. Their difference is a small number contaminated by a fixed-size error -- subtractive cancellation -- and then you divide by `2h`, inflating the contamination to roughly `eps * |L| / h` This term grows without bound as `h` goes to zero. At `h = 1e-10` in double precision the two losses may agree in fourteen of their sixteen significant digits, so the difference retains only a couple of meaningful digits before the division; at `h` near machine epsilon itself, adding `h` to the parameter may not change the stored loss at all and the estimate is exactly zero. ### The U-curve and its minimum Total error behaves like `E(h) = C1 * h^2 + C2 * eps / h` with `C1` set by the third derivative of the loss and `C2` by the loss magnitude. Setting the derivative to zero gives `h_opt` proportional to the cube root of `eps`. For double precision that is about `6e-6`, which is why `1e-5` is the conventional default, and the best achievable error is roughly `eps^(2/3)`, around 1e-10 to 1e-11 in relative terms -- comfortably under a 1e-7 pass threshold. For a one-sided difference the trade is `C1*h + C2*eps/h`, minimised at the square root of `eps`, about 1e-8, with a best achievable error only around `eps^(1/2)` -- roughly 1e-8. Note what that means: the one-sided form's *best possible* accuracy is far worse than the two-sided form's, no matter how carefully you tune the step. Precision moves the whole curve. In single precision, `eps^(1/3)` is around 5e-3, and the error floor rises to roughly 1e-5 or worse -- so a 1e-7 threshold is unreachable there even for a perfectly derived gradient, and any check should be run in double. ### Absolute or relative step An `h` of 1e-5 makes sense for a parameter of order one. For a parameter of magnitude 1e6, adding 1e-5 may be lost entirely in the stored value; for one of magnitude 1e-8, adding 1e-5 is a huge relative move that leaves the local region. Scaling the step -- `h_i = h * max(1, |w_i|)` -- keeps the perturbation meaningful in both regimes. ### Using the curve as a diagnostic The strongest practical use of all this is a sweep. Evaluate the relative error at `h` = 1e-2, 1e-3, ..., 1e-9 and plot the shape. - A clear U with a floor below 1e-7 means the analytic gradient is right and the earlier failure was a step-size choice. - A flat, high curve at every `h` means no step size will save it: the derivation itself disagrees with the loss. - A curve that improves as `h` grows is a sign the loss is not smooth at that point, so the small-`h` probes are seeing something other than curvature. That interpretation turns a red check from a mystery into a two-minute experiment.

  • How does the best step size differ for a one-sided forward difference?
    Its truncation term is first order, so the total error looks like `C1*h + C2*eps/h`, minimised near the square root of machine epsilon -- about 1e-8 in double precision rather than 1e-5. More importantly the achievable floor is far higher: roughly `eps^(1/2)` instead of `eps^(2/3)`. Even perfectly tuned, the one-sided estimate cannot reach the accuracy a central difference reaches at its own optimum.
  • Should h be the same absolute size for every parameter?
    Not when parameter magnitudes vary widely. A step of 1e-5 is sensible for a parameter near one, but for a parameter of magnitude 1e6 it may barely change the stored value, and for one of magnitude 1e-8 it is an enormous relative move that leaves the local neighbourhood entirely. Scaling the step as `h * max(1, |w_i|)` keeps the perturbation proportionate in both directions.
  • The relative error is large and flat across every h you try. What does that tell you?
    That the problem is not the step size. A correct analytic gradient produces a visible U: high at large `h` from curvature, high at tiny `h` from cancellation, with a deep floor in between. If the curve never dips, no step size reconciles the two quantities, which points at the derivation or at the code path that produced the analytic value rather than at the numerics.

Measuring a hill's slope with a plank. A long plank bridges the curvature and reports a slope the ground never had; a very short plank measures a height drop so tiny that the wobble in your ruler is bigger than the drop itself.

saying these in an interview costs you the question

  • Just make h as small as possible, like 1e-12
  • Round-off is negligible in double precision
  • One-sided and central differences are equally accurate
  • Blaming a failed check on the derivation without sweeping h
  • Using one absolute step size regardless of parameter magnitude

context