Before you publish a robustness number from a wrapped model, how do you verify that the gradient surface your wrapper exposes really is the gradient of the model's loss with respect to its input?
answer
- central finite differences
- difference through predict only
- zeros / NaN / shape / sign
- sample coordinates, not all
- test reruns on wrapper change
basics
~20 sDo not assume it — test it. Compare the returned gradient against a central finite-difference estimate on a handful of coordinates and inputs, using only the prediction surface for the loss. Also check it is not all zeros or NaN, and that a small step along it moves the loss the way the sign says.
solid answer
~50 sThe check that actually proves the surface is a finite-difference comparison, because it uses only the prediction path. Pick a few inputs and a few random coordinates; perturb each coordinate up and down by a small step; recompute the loss through the prediction surface; the central difference should match the corresponding gradient entry within tolerance. If it matches, the two surfaces describe the same model. Around that, run cheap sanity checks: the gradient is not all zeros, contains no NaN or inf, has the input's shape and layout (channel order matters), and a single small step along the negative gradient measurably lowers the loss. Check on both a correctly classified and a misclassified example, since some bugs only show on one branch. When finite differences disagree, the usual causes are a detached preprocessing step, a mismatched loss, an input that is not being tracked, or a wrapper whose prediction path and gradient path run different code.
code
python · 9 linesg = wrapper_loss_gradient(x, y)
step = 1e-3
for i in sample_coordinates(x.shape, n=12):
xp, xm = x.copy(), x.copy()
xp[i] += step
xm[i] -= step
fd = (loss(wrapper_predict(xp), y) - loss(wrapper_predict(xm), y)) / (2 * step)
rel = abs(fd - g[i]) / max(1.0, abs(fd))
assert rel < 1e-2, (i, fd, g[i])go deeper
Should at least check the gradient is non-zero, finite, and the right shape before trusting a run.
Describes the finite-difference comparison and knows it must be computed through the prediction surface to be independent.
Chooses the step size deliberately, tests on both correct and incorrect predictions, and can name the usual causes of a mismatch: detached preprocessing, wrong loss, untracked input.
Makes the check a standing harness test tied to wrapper, checkpoint and preprocessing changes, because a degraded gradient always fails in the direction nobody questions.
### Why this check exists at all The gradient surface of an estimator wrapper — `loss_gradient(x, y)` in Adversarial Robustness Toolbox, the framework-native backward pass behind a Foolbox model — is the one component that can be completely wrong while looking completely fine. Shapes match, dtypes match, no exception is raised, and the attack runs to completion. Every way it fails pushes the result in the same direction: the attack gets weaker, so the model looks more robust, so nobody investigates. That asymmetry is the reason the check has to be explicit and automated rather than a thing someone did once in a notebook. ### The core test: central finite differences through `predict` only Pick an input, pick a coordinate, nudge it up by a step and down by the same step, evaluate the loss both times **using only the prediction surface**, and compare the central difference against that coordinate of the returned gradient: ``` fd_i = ( L(predict(x + h*e_i)) - L(predict(x - h*e_i)) ) / (2h) compare fd_i against loss_gradient(x, y)[i] ``` The whole value of the test is that the left side never touches the gradient code path. Agreement is therefore real evidence that the two surfaces describe the same function. Practical settings: sample a dozen coordinates on three or four inputs rather than sweeping the tensor; compare on relative error rather than absolute; and choose `h` deliberately — too small and float32 cancellation dominates the numerator, too large and curvature does. A mismatch that appears only at extreme step sizes is usually a bad `h`, not a bad surface. ### Cheap sanity checks that catch different bugs - **All zeros** — a detached or `no_grad` preprocessing step, an untracked input leaf, or a saturated softmax on a fully confident model. - **NaN or inf** — a log of zero in the loss, or a division by a zero scale in normalisation. - **Wrong shape or transposed layout** — wrapper and model disagree about channel ordering, so every perturbation lands in the wrong place while the array still looks plausible. - **Sign** — one small step along the negative gradient should lower the loss on most inputs. If it consistently raises it, the exposed surface's sign convention is inverted and your attack is a determined defence. - **Batch reduction** — this one is the classic. If the wrapper's loss is constructed with mean reduction over the batch, the per-sample gradient it returns is scaled by 1/N. Difference a single sample and you will see a discrepancy of exactly the batch size. Read as a "failure" it sends people hunting for a graph bug that is not there; read correctly it is a one-line reduction fix. Run all of these on both a correctly classified and a misclassified input, since some bugs only appear on one branch. ### What it costs Almost nothing, which is why there is no excuse. Twelve coordinates on four inputs is 96 extra forward passes plus one backward — seconds. The run it protects is a full attack sweep: iterative attacks at tens of iterations over hundreds or thousands of examples, plus the engineer-days of triage and write-up that follow. The real cost is the one you avoid: re-doing an engagement after the client's own team reproduces your evaluation and gets a different answer. ### Where this check misleads A passing finite-difference test proves **internal consistency between two surfaces of the same object**. It does not prove that object is the deployed system. All of the following pass the check with flying colours and still produce a worthless number: a wrapper built on last month's checkpoint while the service serves this month's; a full-precision local copy while the endpoint runs a quantised export; a wrapper whose preprocessing chain omits a resize the service performs; and a loss that is internally consistent but is not the objective the attack's threat model assumes. Fidelity to production is a separate check with a separate method — compare clean predictions between the wrapper and the live service on the same inputs — and finite differences will never surface it. The second misread is treating a localised disagreement as proof of a bug. ReLU kinks, clipping boundaries and max-pool ties are genuinely non-smooth points; the two-sided difference straddles a discontinuity and disagrees legitimately. Retest those coordinates at a different step, or on a different input, before concluding anything. ### Where it belongs In the harness, as a test that runs whenever the wrapper, the checkpoint, or the preprocessing changes — not as a one-off. Wrappers rot: somebody adds a resize, somebody swaps a normalisation constant, somebody upgrades the framework and an in-place operation starts breaking the tape. The failure is always silent and always flattering, so the only defence is a check that fails loudly on its own schedule.
- Finite differences match on most coordinates but diverge on a few. What is going on?Usually the step size against local curvature or float precision, or a genuinely non-smooth point such as a ReLU kink or a clipping boundary. Retest those coordinates at a different step before concluding the surface is wrong.
- The gradient is exactly zero everywhere on a batch. Name two causes.A detached or no-grad preprocessing step so the input is not on the tape, or a saturated loss where the model is fully confident and the numeric gradient underflows.
saying these in an interview costs you the question
- Validates the gradient by checking that the attack succeeds — circular, since a broken gradient makes it fail.
- Uses the model's own autograd to check its own autograd instead of differencing predictions.
- Only checks the shape and stops there.
- Treats a suspiciously robust result as good news rather than as a harness bug to rule out.