skip to content

In an adversarial-robustness library such as the Adversarial Robustness Toolbox or Foolbox, attack classes differ in what they demand from the model object you hand them. What access levels can an attack class require, and how do you establish which level a given class needs before you run it?

level: middleimportance: must knowfreq 65%

answer

  1. gradients / scores / decisions
  2. constructor arguments leak the assumption
  3. requirement check refuses, does not degrade
  4. query-bound vs compute-bound
  5. finding must name the access level

basics

~20 s

Three levels. White-box classes need loss gradients through the model. Score-based black-box classes need the full confidence vector per query. Decision-based classes need only the predicted label. You establish which from the class's documented requirement and the arguments its constructor insists on; these libraries generally reject a model that cannot supply what the class needs.

solid answer

~50 s

Attack classes are grouped by the model surface they consume. **Gradient (white-box)** classes differentiate a loss through the wrapped model, so they need a real backward pass. **Score-based black-box** classes only call forward, but need the full output vector to estimate a search direction — cheap per step in compute, expensive in queries. **Decision-based** classes work from the top-1 label alone and are the most query-hungry of the three. The practical check before running anything is the class's declared requirement plus what its constructor insists on: a class that wants a loss needs gradients; a class that wants a query budget, a starting point or an initial adversarial example is black-box. These libraries typically validate the requirement when the model is attached and fail rather than silently degrading — but never rely on that. The real failure mode is not an exception; it is running a white-box class against a locally downloaded copy and reporting the number as though it described the deployed endpoint.

go deeper

for a junior

Should distinguish white-box from black-box and know that a gradient attack needs a local differentiable model, not a hosted endpoint.

for a middle

Names all three surfaces, explains that black-box splits into score-based and decision-based, and knows the constructor arguments reveal the assumption.

for a senior

Connects access level to cost shape (queries vs compute) and to remediation — rate limits help against one level and not another — and insists the finding state the level.

for a principal

Sets the standard that every reported number carries its access level, the model artefact used, and its query/compute cost, so results across engagements are comparable.

### Access level is the first assumption in the menu Every attack class in a library such as the Adversarial Robustness Toolbox (ART), Foolbox or Torchattacks consumes one specific **surface** of the model object you hand it. The wrapper you build — ART's estimator, Foolbox's model wrapper — is the adapter between your target and that surface. Which surface the class needs is not a preference; it determines whether the class can run against your engagement's target at all, and it determines what the resulting number means. **Gradients (white-box).** The class differentiates a loss with respect to the input and steps the input against that gradient. This requires a local, differentiable copy of the model *and* of its preprocessing chain, because the gradient has to flow back through resizing, normalisation and tokenisation to reach the raw input. A hosted HTTP endpoint cannot provide this at any price. Cost shape: **compute**. A handful of forward/backward passes per example for a single-step method, tens to hundreds for an iterative one; seconds to minutes of GPU time for a few hundred examples. **Scores (score-based black-box).** The class calls forward only, but needs the full confidence vector per query so it can estimate a descent direction numerically. Cost shape: **queries**, typically hundreds to a few thousand per example depending on the class, the input dimensionality and the perturbation budget. **Decisions (decision-based black-box).** The class sees only the top-1 label — one bit of useful feedback per query, "did it flip or not" — and walks along the decision boundary from a starting point that is already misclassified. Cost shape: **queries**, and far more of them: commonly thousands to tens of thousands per example for a tight perturbation. On a metered endpoint at, say, a tenth of a cent per call, one example can be a few dollars and a day of wall-clock against a rate cap. ### How to establish the requirement before you run 1. **Read what the class declares.** These libraries document, per attack, the model capability it assumes; ART additionally declares which estimator interfaces an attack requires, and the framework refuses an estimator that does not implement them. 2. **Read the constructor.** Required arguments leak the assumption. A class that wants a loss object or a gradient-bearing model is white-box. A class that wants a query budget, a starting adversarial example, or a sampling/step-count parameter is a black-box search. A class that wants a `mask` or a feature-constraint argument is telling you it expects a structured input domain. 3. **Attach the model and let the check speak.** When the wrapped model does not expose the surface, the run stops — refused at attach time, or erroring the moment it tries to differentiate. It does **not** silently degrade into a weaker method. That failure is the safe outcome. ### Where the number misleads The failure mode is almost never the exception. It is the workaround: a gradient class refuses the client's label-only endpoint, so someone downloads a checkpoint from a public hub, points the class at that, and the run succeeds. The percentage that lands in the report is about a different artefact under an access level nobody granted. The served model may be fine-tuned, quantised, wrapped in different preprocessing, or fronted by a filter — none of which the toolkit knows or reports. The second misreading is **cross-level comparison**. "Attack A: 96%; attack B: 31%" is meaningless unless both ran at the same access level, the same perturbation budget and against the same artefact, because the number is the joint product of all three. And "black-box" does not mean "cheap" — a decision-based run is the most expensive of the three per example, which is why an attacker with weights and an attacker with an API face genuinely different barriers. The third is **remediation mismatch**. Rate limiting and query-pattern anomaly detection are real mitigations against a query-bound decision-based attacker and no mitigation whatsoever against someone running a gradient class on leaked weights. A finding that does not name its access level cannot be routed to the right control. ### What you check on a real run Which surface the class required; whether the model attached was the served artefact or a local stand-in, and how you established that; the perturbation budget and norm the number was measured at; and the run's actual cost — queries spent per success, or GPU-minutes — because at black-box levels that cost *is* the attacker's barrier, and it belongs in the finding beside the success rate.

  • You have a local copy of the model but the client only ever exposes a label-returning API. Which access level should the headline number in your report come from?
    The one the client's attacker actually has — the label-only surface. The white-box run belongs in the report too, but labelled as what it is: a different, more permissive attacker position, useful as a bound and for diagnosis.
  • Why do score-based classes usually need far fewer queries than decision-based ones?
    A confidence vector gives a graded signal to estimate a search direction from; a bare label gives one bit of feedback per query, so the search has to spend queries discovering the boundary's shape.

saying these in an interview costs you the question

  • Picking an attack class by name recognition and never asking what it needs from the model.
  • Assuming a gradient class will fall back to a black-box search when gradients are unavailable.
  • Reporting a success rate without saying which access level produced it.
  • Believing 'black-box' means cheap — decision-based classes are the most query-expensive of the three.

context