A model wrapper for an adversarial robustness toolkit can expose predictions only, or predictions plus the gradient of the loss with respect to the input. What must you implement for each surface, and what happens if you return a stub from the gradient one?
answer
- two surfaces: predict, loss gradient
- gradient must flow through preprocessing
- stub returns zeros, no error
- zeros look like robustness
- nondeterminism reads as robustness
basics
~20 sA predict-only wrapper maps a batch of inputs to model outputs. A gradient-exposing wrapper must additionally return the derivative of the loss with respect to the input, at the same input the model sees. Stub that and gradient attacks still run, but they follow a fake signal and report robustness that only measures your stub.
solid answer
~50 sThe prediction surface is the easy one: take a batch in the declared input domain, run the serving path, return scores in a fixed shape. It must be batched, deterministic, and free of per-request state. The gradient surface is a real derivative: for a given input and label, the gradient of the loss the attack is optimising with respect to the input tensor. Two details decide whether it is honest. It must be the loss the attack assumes — a wrapper that silently differentiates a different objective still returns well-shaped arrays. And it must be differentiated through everything the prediction surface does, including preprocessing, or the two surfaces describe different models. A stub is the dangerous case. Returning zeros or noise does not throw; the attack takes useless steps and ends with a low success rate. That looks exactly like a robust model, so nobody investigates it.
go deeper
Should know predictions in/out versus a gradient of loss with respect to the input, and that gradient attacks need the second surface.
Explains that the gradient must be taken through the same path the prediction surface runs, and that a stubbed gradient silently inflates robustness instead of erroring.
Adds the nondeterminism traps (dropout, sampling, caching, partial batches) and insists on falling back to a prediction-only evaluation rather than faking a gradient.
Treats the choice as a reporting decision: the surfaces the wrapper exposes define what the engagement's number is a claim about, and that has to be stated up front.
### Two surfaces, and what each one is From the library's point of view the wrapper **is** the model. Everything it can learn about the system under test, it learns by calling the methods you implemented. In Adversarial Robustness Toolbox those are `predict(x)` and `loss_gradient(x, y)` (plus `class_gradient` for attacks that want a per-class derivative); Foolbox reaches the same two capabilities through its framework-specific model objects. Which of the two you implement decides which attack families are even constructible against your wrapper, and — more importantly — decides what the resulting number is a claim about. **The prediction surface.** A batch of inputs in the declared input domain goes in, an array of scores comes out. The requirements people skip: it must be batched (attacks call it with thousands of rows), it must have a stable class ordering and a fixed output width, and it must be **deterministic**. Dropout left in training mode, a sampling temperature, a per-request cache, batch-norm still updating running statistics, or a rate limiter that silently returns a short batch all convert a deterministic optimisation into a noisy one. The attack cannot tell noise from curvature; it reads the noise as a rugged loss landscape, its steps stop being consistent, success rate falls, and the model looks robust for a reason that has nothing to do with the model. **The gradient surface.** The derivative of a scalar loss with respect to the *input tensor*, at the same point the prediction surface evaluates. Getting this honest means four things at once: the autograd tape is intact from the estimator's input, through your preprocessing, to the loss; the input is a tracked leaf and not a detached copy; no step in the middle sits inside a `no_grad` block or an in-place operation that breaks the graph; and the loss you differentiate is the loss the attack believes it is optimising. ART takes that loss as an explicit constructor argument (`PyTorchClassifier(loss=nn.CrossEntropyLoss())`) precisely so it is stated rather than inherited. ### What it costs A prediction-only wrapper against a model you can already call is usually an afternoon. A faithful gradient wrapper is days to weeks whenever the serving path is not a single Python process — preprocessing implemented in a Go service, a compiled inference runtime, a quantised export, or a remote endpoint you can only POST to. That gap is exactly why stubs get written: someone needs the gradient attacks to *run* by Friday, and a `return np.zeros_like(x)` makes them run. ### Where the number misleads This is the part that matters. Almost no toolkit can distinguish a wrong gradient from a genuinely flat loss region, because both are well-shaped arrays of the right dtype. | stub | what the attack does | what the report says | |---|---|---| | zeros | steps of length zero, exits at the iteration cap | robust accuracy near the clean accuracy | | Gaussian noise | random walk inside the epsilon-ball | a weak random search wearing PGD's name | | gradient of a different loss | consistent but misdirected steps | plausibly-low success rate, no error anywhere | | sign inverted | walks downhill on the loss | the "attack" makes the model more confident | Every row fails in the same direction: **toward a flattering number**. That is what makes it dangerous. A result that says the model is weak gets scrutinised by the model owner; a result that says it is strong gets pasted into a slide. And note the confound: a defence that deliberately flattens or obscures gradients produces output indistinguishable from a broken wrapper. Sorting those two apart is the defences topic's problem, but you cannot even begin until you have ruled out your own wrapper. ### What you would check Before any gradient attack result leaves the harness: assert the returned array is finite, non-zero on a batch you know the model gets right, and has the input's exact shape and channel layout. Take one small step along the negative gradient and confirm the loss actually decreases. Then difference the loss numerically through the prediction surface on a sample of coordinates and compare — that is the only check independent of the code path under suspicion. If you genuinely cannot differentiate the deployed path, **do not fake the surface**. Expose predictions only and drive decision- or score-based attacks that are honest about that access. A weaker attack you can defend is worth more in a deliverable than a strong attack fed a lie, because the second produces a robustness claim that collapses the first time a client's own team reproduces it.
- The prediction surface is nondeterministic because dropout was left active. What does the attack report?Inflated robustness. The optimiser reads the sampling noise as a rugged loss surface and its steps stop being consistent, so success rate falls for a reason unrelated to the model's actual robustness.
- Your wrapper reconstructs the model locally from a checkpoint while the service runs a quantised export. Is the gradient surface still valid?Only as an approximation. The gradients are the full-precision model's; you must say so, and confirm on the prediction side that the two copies agree on the examples you report.
- When is predict-only the right choice even though you have the weights?When the serving path contains a step you cannot differentiate faithfully, so a gradient surface would describe a model that is not the one being served.
A stubbed gradient is a compass with the needle glued in place: the search party still walks its full route, comes back empty, and reports that there was nothing out there to find.
saying these in an interview costs you the question
- Believes an incorrect gradient will make the attack error out.
- Reports high robustness from a gradient attack without ever validating the gradient surface.
- Leaves the model in training mode or with sampling enabled behind the prediction surface.
- Fills in an approximate gradient without saying so in the deliverable.