When does a finite-difference gradient check earn a permanent place in a test suite?
answer
- only where a human wrote the derivative
- tiny model, fixed inputs, seconds to run
- sample coordinates, do not sweep them
- double precision or the threshold is unreachable
- consistency, not correctness of the objective
basics
~20 sWherever a derivative was written by a human rather than composed from already-trusted primitives: a custom layer, a custom loss, a hand-optimised operation. Keep the test tiny, deterministic and in double precision so it runs in seconds and never flakes.
solid answer
~50 sThe check earns its keep exactly where a human wrote algebra -- a hand-derived layer, a custom loss, a fused operation -- and nowhere else; on a model assembled from trusted primitives it burns runtime without exercising any new derivative code. To be permanent it must be cheap and stable: a miniature model with a few dozen parameters, fixed inputs, every random draw pinned, double precision, and a tolerance calibrated from what a known-good version actually produces. Probe a sample -- say twenty random coordinates plus deliberately chosen ones like biases and each distinct code path -- since a systematic algebra error corrupts many coordinates at once and a sample finds it. The limit is the part worth stating out loud: the check only proves the backward code differentiates the forward code as written. If the forward pass implements the wrong objective, it passes at 1e-11, so the forward pass needs its own reference test.
go deeper
Know that these checks exist for gradients someone derived by hand, and that they run on a deliberately tiny model rather than on the real one.
Be able to describe a check that is fast and repeatable: few parameters, fixed inputs, no live randomness, double precision, and a tolerance taken from observed error rather than invented.
Show judgment about coverage and durability -- which custom operations get a check, how coordinates are sampled, and how you stop the test from becoming flaky and quietly disabled.
Own the boundary of what this verification buys: agreement between forward and backward code only, which means the forward pass needs its own independent reference test to close the remaining gap.
## The decision, framed properly The question is not whether finite-difference checking is a good technique -- it is -- but where it belongs in a codebase that a team maintains for years. A permanent test costs runtime on every run, attention every time it goes red, and credibility every time it goes red for a reason nobody can explain. So the standard is: does this test catch a class of bug that nothing else in the pipeline catches, and can it do so without flaking? ## Where it pays **Hand-written derivatives.** A custom loss whose gradient someone worked out on paper, a layer implemented for speed with its own backward rule, a specialised operation with an analytic form. This is code where an algebra slip is genuinely plausible and where no other test would notice: the model still trains, just worse. This is the whole justification for the technique. **Custom operations at their boundaries.** Beyond a happy-path check, probe the shapes and regimes that the derivation treated specially -- an edge case in a padding rule, a branch that triggers only for a particular input size. ## Where it does not pay **Standard compositions.** A model built from primitives whose derivatives are already trusted has no new derivative code, so the check verifies machinery that is verified elsewhere and merely slows the suite down. **Surrogate gradients.** Where a backward rule is deliberately not the true derivative -- a straight-through estimator, for instance -- there is no agreement to expect and the check has nothing to say. **The production model.** Every coordinate costs two forward passes, so checking a large model is arithmetically hopeless. The test model should be a miniature. ## Designing a test that will still be green in a year - **Make it small.** A few channels, a batch of two, a few dozen parameters. The whole check should finish in about a second, which also makes it cheap enough to run on every commit. - **Make it deterministic.** Fixed inputs, fixed initial parameters, every random draw inside the loss pinned to a single value reused across both perturbed evaluations. A test that is stochastic will eventually fail for no reason and then be disabled by someone under deadline pressure. - **Run it in double precision.** In single precision, subtracting two nearly identical losses leaves so few significant digits that the achievable relative error floor sits far above a 1e-7 threshold; a perfectly correct gradient would fail. This is also the diagnosis when a custom geometric or physics loss passes in double and fails in single: the derivation and the backward code are right, and the failure indicts the arithmetic of the check, not the gradient. - **Calibrate the tolerance, do not guess it.** Run the check on a version you are confident about, look at the distribution of relative errors, and set the threshold with margin above the observed worst case. Too tight and the suite becomes flaky; too loose and it stops catching small systematic errors such as a slightly wrong coefficient on a term. - **Choose coordinates deliberately.** With more parameters than you can afford to probe, sample a couple of dozen at random and add fixed picks: at least one bias, one parameter in each distinct branch of the code, one in each tensor the derivation treats differently. Systematic errors are shared across coordinates, so a modest sample finds them with high probability. Keep the sample deterministic in the committed test; randomise it only in a nightly run where a new failure can be investigated rather than blocking a merge. - **Report per-coordinate, not pass or fail.** When it does go red, the distribution of errors across coordinates is what tells you whether you are looking at an algebra bug, a non-smooth operation, or a step-size problem. ## Stating the limit out loud A finite-difference check verifies internal consistency: that the backward code computes the derivative of the forward code that exists. It cannot see whether the forward code computes the objective you meant. A loss that squares the wrong residual, applies a factor in the wrong place, or reduces over the wrong quantity will be differentiated perfectly and the check will pass with room to spare. That is not a weakness of the technique so much as a boundary on what it buys you, and a lead should say it explicitly when the team treats a green gradient check as proof the layer is correct. The complement is a forward-value test: compute the layer's output for a small input by hand or from an independent implementation and assert on the number. The pair -- forward values against an independent reference, gradients against finite differences -- covers both halves. Either alone leaves a hole, and the gradient check is the half people remember to write.
- How many coordinates should the check probe on a model too large to sweep?The better move is to shrink the test model until a full sweep is cheap. Where that is impossible, a couple of dozen randomly chosen coordinates plus deliberate picks -- a bias, one parameter per distinct code branch, one per tensor the derivation treats differently -- is enough, because a systematic algebra error corrupts many coordinates at once and any sample is very likely to hit one. Keep the sample fixed in the committed test so failures are reproducible.
- A custom geometric loss passes its gradient check in double precision but fails in single. What does that tell you?That the derivation and the backward code are almost certainly fine and the check's own arithmetic is the problem. Subtracting two nearly identical losses in single precision leaves very few significant digits, so the achievable relative-error floor sits well above a tight threshold. Run gradient checks in double precision as a matter of policy; the failure is evidence about the measurement, not about the gradient.
- What class of bug does a passing gradient check completely miss?Any error in the forward pass itself. The check confirms only that the backward code differentiates whatever forward code is present, so a loss that implements the wrong objective is differentiated faithfully and passes with a relative error of 1e-11. Catching that needs a separate test asserting forward outputs against values computed independently -- by hand on a small input, or from a reference implementation.
saying these in an interview costs you the question
- Gradient-check the full model before every training run
- Checks are pointless because automatic differentiation is always right
- A green check proves the loss implements the intended objective
- Loosen the tolerance whenever the test goes red
- Let the test draw fresh random inputs on every run