skip to content

Why does a gradient check fail on ReLU units whose pre-activation sits near zero?

level: seniorimportance: should knowfreq 34%

answer

  1. the function is not smooth there
  2. piecewise linear with a corner
  3. the two probes land on opposite sides
  4. secant blends two different slopes
  5. few isolated coordinates fail, not the layer

basics

~20 s

The loss is piecewise linear in that parameter with a corner between the two probe points, so the numerical secant blends both slopes while the backward pass reports one. The check fails though the code is correct.

solid answer

~50 s

ReLU has a kink at zero: no single derivative exists there, and just either side of it the slope jumps between 0 and 1. If a unit's pre-activation is within the perturbation's reach of zero, the `+h` probe lands in the active regime and the `-h` probe in the dead one. The difference quotient then measures a secant across the corner -- effectively a blend of the two one-sided slopes -- while the backward pass consistently uses the slope of the side the unperturbed point is on. The relative error can be of order one on those coordinates even though nothing is wrong. The signature is diagnostic: a handful of coordinates fail badly while the rest sit at 1e-9, and the failures move when you re-probe from a shifted point. Fix it by choosing probe points where no pre-activation is near zero, not by loosening the tolerance until everything passes.

go deeper

for a junior

Remember that ReLU has a corner at zero where no single slope exists, and that differencing across that corner cannot match either side.

for a middle

Explain the mechanism concretely: with a pre-activation within the perturbation's reach of zero, one probe lands in the active regime and the other in the dead one, so the quotient blends two slopes.

for a senior

Show the triage you would actually run -- count how many coordinates fail and by how much, re-probe from a shifted point, and check whether the failures track the units sitting closest to zero.

for a principal

Set the team's policy for non-smooth models: which operations are checked at deliberately shaped inputs, which are exempt and verified another way, and how you keep those exemptions from hiding real bugs.

## The mechanism A finite-difference check silently assumes the loss is smooth in the neighbourhood it probes. Taylor's theorem, the `O(h^2)` truncation term, the whole error analysis -- all of it needs derivatives to exist between the two probe points. ReLU, defined as `r(z) = max(0, z)`, breaks that assumption at exactly one point: `z = 0`, where the left slope is 0 and the right slope is 1 and no single derivative exists. Now trace one coordinate. Suppose parameter `w` feeds a pre-activation `z = w*x + b` with `x` positive, and at the current point `z = 3e-6` -- just barely active. Perturbing `w` by `h` moves `z` by `h*x`; say that is 1e-5. Then: - the `+h` probe has `z_plus = 1.3e-5`, still active, contributing `z_plus` downstream; - the `-h` probe has `z_minus = -7e-6`, now dead, contributing 0. The numerical estimate is a secant drawn across the corner. It reflects the change over an interval that spent part of its length on the flat side and part on the sloped side, so it lands somewhere between the two one-sided derivatives. The backward pass, meanwhile, evaluated the unit once at `z = 3e-6`, found it active, and consistently propagated the active-side slope. The two numbers disagree by a large fraction, and the relative error can approach one. Nothing is broken. The function genuinely has no derivative in that interval, so there is no true value for the two methods to agree on. ## Recognising it rather than guessing The failure has a distinct fingerprint, and reading it correctly is the actual senior skill here. - **A wrong derivation fails systematically.** A missing factor, a transposed term, a dropped chain-rule step corrupts every coordinate flowing through that expression, so you see a whole layer failing, often by a consistent ratio like exactly 2 or exactly the batch size. - **A kink fails sparsely and erratically.** Three coordinates out of forty blow past the threshold while the remaining thirty-seven agree to 1e-9. Sparse failures with an otherwise pristine floor are almost always non-smoothness. - **A kink failure is location-dependent.** Re-run the check from a slightly different starting point or with a different input draw, and a kink failure moves to different coordinates or vanishes. A derivation bug fails from everywhere. - **A kink failure gets worse as h grows.** The chance that the interval `[z - h*x, z + h*x]` contains the corner is proportional to `h`, so shrinking the step reduces how many coordinates are affected -- until round-off takes over. ## Fixing it properly The wrong fix is to raise the tolerance until the test is green, because that tolerance now also admits genuine derivation errors of comparable size. The right fixes make the probed region smooth: - **Shape the probe point.** Draw inputs and initial parameters so that no unit's pre-activation lies within the perturbation's reach of zero, and nudge the offending coordinate if one does. This is the standard move and it costs nothing. - **Substitute a smooth activation for the check.** Verifying the layer's algebra with tanh or softplus in place of ReLU exercises the same chain-rule structure without corners, then check the ReLU path separately at points well away from zero. - **Shrink the step**, accepting the round-off trade, so fewer intervals straddle a corner. - **Report, do not average.** Log which coordinates fail and by how much rather than reducing the whole layer to a single pass or fail; the pattern is the information. ## The same trap in other operations ReLU is the famous case but far from the only one. Max-pooling has a kink wherever two inputs tie, because an infinitesimal perturbation can change which input wins and route the gradient elsewhere. An absolute-value term in a loss has a corner at zero. Hard clipping at a bound is flat on one side and sloped on the other. Any hard selection -- a top-k, a discrete branch, an argmax -- produces the same discontinuous routing. A related but distinct case is the straight-through estimator, where a non-differentiable step-like operation is deliberately given a surrogate backward rule. There the analytic gradient is not meant to equal the true derivative, which is zero almost everywhere; a finite-difference check will disagree by design and disagreeing is not evidence of a bug. Such code has to be verified against the intended surrogate formula instead, not against numerical differentiation. ## Why this matters beyond the test The same non-smoothness that breaks the check is present during training, where the optimiser is happily using one-sided slopes at corners. That is fine in practice -- the set of exact kinks has measure zero and the subgradient chosen is a legitimate descent direction -- but it explains why a numerical check is a statement about a *smooth* neighbourhood, and why a rectified network can never be checked as cleanly as a smooth one.

  • How do you tell a kink failure apart from a genuine derivation bug?
    Look at the pattern, not the single worst number. A derivation error corrupts every coordinate that flows through the faulty expression, so a whole layer fails, often by a consistent ratio. A kink failure is sparse -- a few coordinates far off while the rest sit at 1e-9 -- and it moves when you re-probe from a shifted point or a different input draw. Shrinking the step also thins kink failures while leaving a derivation error untouched.
  • What happens if you gradient-check a layer that uses a straight-through estimator?
    It fails, and correctly so. A straight-through estimator deliberately substitutes a surrogate backward rule for an operation whose true derivative is zero almost everywhere, so the analytic value is not trying to equal the numerical one. Verifying that code means asserting it implements the intended surrogate formula, and checking the surrounding smooth parts of the graph separately. Running a finite-difference check across it proves nothing either way.
  • Which other operations produce the same false failures?
    Anything with a corner or a hard selection. Max-pooling ties, where a tiny perturbation flips which input wins and reroutes the gradient; absolute-value terms in a loss; hard clipping at a bound, flat on one side and sloped on the other; top-k selection and discrete branches. The common thread is a point where the function is continuous but its slope jumps, or where routing changes discontinuously.

Measuring the pitch of a roof by standing with one foot either side of the ridge and comparing heights. Your two readings give you the average of two slopes, and neither side of the roof actually has that pitch.

saying these in an interview costs you the question

  • Any failed coordinate proves the backward pass is wrong
  • ReLU is differentiable everywhere, so kinks cannot matter
  • Fix it by loosening the tolerance until everything passes
  • Make h much larger so the kink averages out
  • Reporting one pass or fail for the layer instead of per-coordinate errors

context