skip to content

In a decision-based (label-only) attack run from an adversarial robustness toolkit against a remote endpoint, 60% of examples hit the per-example query cap without producing an adversarial example. What do you check in the run before you treat that as a property of the model, and what do you change for the next run?

level: seniorimportance: should knowfreq 42%

answer

  1. saturation is censored data
  2. queries-to-success tail near the cap
  3. perturbation still shrinking?
  4. count errors, throttles, duplicates
  5. subsample at ten times the cap

basics

~20 s

Treat it as censored data, not a model property. Check the queries-to-success distribution on the examples that worked, whether the perturbation was still shrinking when the cap hit, and whether errors, throttling or duplicate queries burned the allowance. Then re-run a stratified subsample at a much larger cap.

solid answer

~60 s

A cap-saturated run produces censored measurements: for 60% of examples you know only that success needed more than the cap, not that it is unreachable. The first thing to look at is the distribution of queries-to-success among the 40% that finished. If it has a long right tail that is still climbing near the cap, the cap is the binding constraint and the number is about your budget. If successes clustered far below the cap and the rest went nowhere, the failures may be genuinely harder. Second, check the per-example trajectory: was the perturbation norm still decreasing when the attack stopped? A still-improving search that was cut off is a budget verdict, not a robustness verdict. Third, audit where the queries went. Error responses, rate-limit rejections and repeated identical probes are often counted against the cap by the attack loop even though they returned no information. The next run: stratified subsample of the saturated examples at ten times the cap, with caching and a distinct counter for informative versus wasted calls.

go deeper

for a junior

Should at least say the result depends on the query budget and that hitting the cap is not the same as failing.

for a middle

Should ask for the queries-to-success distribution and re-run a subsample at a higher cap.

for a senior

Should treat it as censored data, audit where the queries actually went, inspect the perturbation trajectory, and design a stratified follow-up plus a transfer baseline.

for a principal

Should decide what claim the organisation can defensibly publish from a budget-bound run and set the reporting rule that the cap always travels with the number.

### Saturation is a property of the run, not of the model When a decision-based attack — `ART's HopSkipJump` or a boundary-style search in Foolbox — stops because it exhausted its per-example evaluation budget, the only thing you have learned about that example is that success needs *more than the cap you chose*. That is a censored observation, in the statistical sense: the true queries-to-success value exists but lies somewhere to the right of your cutoff. Reporting "60% of examples resisted the attack" collapses a censored measurement into a model property, and the number will move the moment somebody re-runs with a larger cap. It is a measurement of your budget wearing the model's name. ### Diagnose the censoring Take the 40% that finished and plot queries-to-success as a distribution or a survival curve. Two shapes, two verdicts. If the curve is still descending steeply as it approaches the cap, the population of successes is truncated and the observed 40% is a function of the cutoff — raise the cap and the share rises. If successes cluster far below the cap and then the curve goes flat well before it, the saturated examples are plausibly harder in kind, not merely in degree, and more budget alone may not close them. Either way, the cap travels with the number in every artefact you publish; a success share without its evaluation budget cannot be compared with anything, including your own next run. ### Diagnose the search Decision-based attacks walk a point that is already misclassified back toward the original, shrinking the perturbation norm as they go. Log that norm per query for a handful of saturated examples and read the trajectory: - **still decreasing at the cap** — budget-bound; the attack was working when you switched it off. - **plateaued high** — the search stalled; more budget buys nothing, and the fix is a different attack family or a better starting point (a nearer misclassified seed). - **oscillating** — the target is returning inconsistent labels for the same input, so the search is chasing noise; stabilise the decision function before spending anything more. Those three shapes lead to three completely different follow-up runs, and you cannot tell them apart from the success share alone. ### Diagnose the spend Count calls by outcome, not in aggregate: informative answer, HTTP error, rate-limit rejection, duplicate of a query already issued. Many attack loops decrement their evaluation budget on any call, so a run that burned a third of its allowance on 429s and 5xx never had the budget its log claims. If that share is large the result should be discarded and re-run, not reported with a caveat. ### Rule out the wrapper before blaming the model A batching or ordering bug that misaligns responses to inputs looks exactly like an unbreakable model: every probe comes back "wrong" from the search's point of view, so nothing ever converges. Cheap check — send a handful of known inputs through the wrapper singly and in a batch and confirm identical labels in identical positions. Also confirm any preprocessing in the wrapper (resize, normalise, clip) matches what the served pipeline does, since a mismatch quietly destroys small perturbations before they arrive. ### What it costs to answer the question properly Re-running the whole set at ten times the cap multiplies the bill by ten, which is usually unaffordable and is also the wrong experiment. Stratify the saturated examples — by class and by how far the perturbation had come — take a subsample you can afford at ten times the cap, and run a cheap transfer baseline from a local surrogate alongside it. Both outcomes pay: if the subsample breaks at the larger budget you now know the shape of the cost curve and can quote robustness as a function of attacker budget; if it does not, you have a far stronger statement about a smaller, named set of examples, which is a more defensible deliverable than a soft claim about all of them. ### What I would report Cap, share saturated, queries-to-success distribution for the finished examples, calls by outcome, and the trajectory verdict. Anyone downstream who sees a bare percentage will read it as robustness, and the correction is much harder to make later than the qualification is to write now.

  • What single number, published alongside the result, makes a cap-saturated run interpretable to a later reader?
    The per-example query cap itself, together with the share of examples that hit it — without those the success share cannot be compared with anything.
  • The perturbation trajectory oscillates instead of shrinking. What does that suggest?
    The target's label answers are not stable for the same input, so the boundary search is chasing noise; stabilise the response before spending more budget.

A cap-saturated run is a stopwatch you stopped at ten minutes: the runners who finished have times, and everyone else you can only record as 'more than ten minutes'. Averaging that as though the unfinished ones took exactly ten minutes tells you about your stopwatch, not about the runners.

saying these in an interview costs you the question

  • Reporting the saturated examples as a model property with no mention of the cap
  • Never plotting queries-to-success or the perturbation trajectory
  • Not distinguishing informative answers from errors and throttled responses inside the query count
  • Raising the cap for the whole set instead of a stratified subsample, and blowing the budget

context