In an adversarial robustness library such as the Adversarial Robustness Toolbox or Foolbox, every attack runs against a wrapper object you supply for the model under test. If that wrapper can only send an input to a remote inference API and return the class probabilities it gets back, which attack families can you still run, and why do the rest fail?
answer
- wrapper interface gates attack set
- gradient / score / label access
- forward-only means query attacks
- silent gradient-estimation fallback
- surrogate plus transfer as the cheap route
basics
~20 sAttacks call methods on your wrapper. Gradient attacks need a gradient method, which a forward-only HTTP wrapper cannot implement, so they error or are unavailable. You are left with query attacks: score-based ones that read the returned probabilities, decision-based ones needing only the top label, and transfer from a local surrogate.
solid answer
~50 sAn attack in these libraries is written against the wrapper's interface, not against your model, and each family calls a different subset of it. Iterative gradient attacks ask the wrapper for a gradient of the loss with respect to the input. Score-based query attacks need only a probability vector per input. Decision-based attacks need only the predicted label. A wrapper backed by a remote endpoint can implement forward prediction and nothing else, so gradient-consuming attacks either refuse at construction or silently fall back to numerically estimating the gradient, which costs many calls per step. What is left is the query family, plus the transfer route: train a local surrogate on the target's answers and attack that with gradients. The tradeoff is that query attacks buy the same information by asking, so the cost moves off GPU time and onto per-example call count against a metered endpoint.
go deeper
Should know that gradient attacks need something more than a prediction API, and that black-box attacks work by repeated querying.
Should map the three access classes (gradient, score, label) onto attack families and explain why a remote wrapper lands in the query families.
Should catch the gradient-estimation fallback, question whether the returned score vector is faithful, and compare the query route against a surrogate-plus-transfer route on cost.
Should frame the wrapper's capability as the decision that fixes the assessment's cost model, and set a policy for what access the engagement asks for up front.
### The wrapper is a contract about what the target will tell you In the Adversarial Robustness Toolbox (ART) the object an attack is constructed against is an *estimator*: `ART's PyTorchClassifier` or `ART's TensorFlowV2Classifier` wraps a model you hold in process, while `ART's BlackBoxClassifier` wraps nothing but a Python callable that takes a batch of inputs and returns a batch of predictions. Foolbox plays the same role with `foolbox.models.Model`, which you call like a function on a batch. The crucial fact is that no attack in these libraries is written against *your* model. Each attack is written against the wrapper's interface, and each attack family calls a different subset of that interface's methods. That subset, not the model, decides what you can run. Three access classes, three method sets: | access class | what the wrapper implements | attacks it unlocks | cost per example | |---|---|---|---| | gradient | `loss_gradient()` / `class_gradient()` on top of `predict()` | FGSM, PGD, Carlini-Wagner, AutoAttack | tens of forward+backward passes on your own hardware; no per-call bill | | score | `predict()` returning a full probability or logit vector | `ART's SquareAttack`, `ART's ZooAttack`, NES-style search | hundreds to tens of thousands of calls | | label | `predict()` from which only the argmax is used | `ART's HopSkipJump`, boundary-walking attacks | thousands to tens of thousands of calls | A wrapper backed by an HTTPS inference endpoint can implement `predict()` and nothing else. There is no parameter tensor on your side of the wire and no backward pass to run, so `loss_gradient()` cannot be written at all — `ART's BlackBoxClassifier` simply does not define it, and handing that object to a gradient attack fails at construction rather than degrading quietly. What is left is the query family (score-based if the endpoint returns a vector, decision-based if it returns only a label), plus the transfer route: spend calls labelling a dataset, train a local surrogate on those labels, take free gradients on the surrogate, and spend metered calls only to check which candidates carry over. ### What it costs Gradient access is cheap because one backward pass hands the attack an entire search direction; a 40-step PGD is 40 forward and 40 backward passes, seconds of GPU time, zero API spend. Query attacks buy the same direction by asking, and the price shows up on an invoice. Order of magnitude: a decision-based run at a five-figure per-example evaluation budget over a few hundred examples is millions of calls. At a tenth of a cent per call that is thousands of dollars; at a few calls per second under a rate limit it is weeks of calendar, which usually binds before the money does. The surrogate route converts that recurring meter into a one-off labelling cost — the labels are paid for once and then amortise across every attack you try. ### Where the number misleads The dangerous case is not the attack that refuses to construct; it is the one that runs. Both libraries let you bolt a numerical gradient estimator onto a forward-only wrapper — Foolbox ships gradient-estimator wrappers for exactly this, and zeroth-order attacks like `ART's ZooAttack` are built on the idea. The estimator approximates each derivative by finite differences, which costs on the order of two calls per coordinate estimated, per step. A run that reads in your notes as "40-step PGD" is then a five-figure call count per example, and the reported attack-success rate is no longer a PGD result at all: it is the result of a coarse, budget-truncated approximation to PGD, and it will read *lower* than the true white-box number while looking like a white-box number. The second misreading is about the response body. An endpoint that returns top-k, rounded, or temperature-flattened probabilities is closer to label access than score access however rich the JSON looks. Score-based searches read the low-order digits of those numbers as their signal; quantise them and the search degrades toward a random walk, and the resulting "the model is robust" verdict is a statement about the serialisation format, not the model. ### What I would check Log the wrapper's actual call count and compare it against the attack's own counter — a gap means an estimator or a retry path you did not price. Print the returned vector for one input and confirm it is full-precision and full-length, not top-k. Confirm the attack class you think you ran is the class that ran, by name, in the run log. Read the attack's per-example evaluation cap out of the constructed object rather than from the docs. And price a surrogate-plus-transfer pilot on ten examples before committing the budget to direct queries.
- The library lets you wrap a forward-only endpoint and still construct a gradient attack. What has it actually done?Attached a numerical gradient estimator: every step now spends a batch of queries approximating the derivative, so the per-example call count jumps by orders of magnitude compared with a real backward pass.
- Your endpoint returns only the top class and its confidence, not the full vector. Which query family does that put you in?Effectively label access with one scalar. Score-based attacks that need the runner-up class's score are crippled; decision-based boundary-walking attacks still work, at a higher call count.
- Why might a locally trained surrogate be the cheapest option here?Because gradients on the surrogate are free once it exists, so the only metered spend is the calls used to label its training data, and one dataset amortises across many attacks.
saying these in an interview costs you the question
- Saying you can just run a gradient attack against any endpoint without noticing the gradient has to come from somewhere
- Not knowing that score-based and decision-based attacks need different response content from the target
- Accepting a gradient-estimation fallback without asking what it costs in calls
- Assuming a returned probability vector is faithful when the endpoint truncates or rounds it