skip to content

You are about to launch a query-based (black-box) attack from an adversarial robustness toolkit against a hosted model endpoint that bills per call. How do you estimate what the run will cost in calls before you start it?

level: juniorimportance: must knowfreq 55%

answer

  1. examples x cap x restarts, plus clean pass
  2. failures spend the full cap
  3. pilot ten examples, count real calls
  4. retries and batch padding are billed
  5. hard ceiling that raises, not warns

basics

~20 s

Multiply the number of evaluation examples by the attack's per-example query cap, times restarts, and add one clean pass to find which examples are classified correctly. Attacks spend the whole cap on examples they fail. Set the cap explicitly, pilot on ten examples, measure the real call count, then scale.

solid answer

~50 s

The worst case is close to the real case, because a query attack that fails burns its entire per-example allowance. So the arithmetic is: examples selected for attack, times the per-example query cap, times restarts or repeats, plus a clean evaluation pass over the whole set to pick the correctly classified examples worth attacking. Two corrections matter. Successful examples finish early and cost less than the cap, so the true number sits between the cap-times-examples ceiling and whatever the pilot measured. And retries against a flaky or rate-limited endpoint are billed calls too. The reliable move is to run ten examples with counting instrumented in the wrapper, read the observed calls-per-example distribution, then extrapolate with headroom rather than trusting a documented default. Never launch a full set against a metered target without an explicit cap set; the default is chosen for a local model where a call is free.

code

python · 9 lines
python
n_examples = 500
clean_accuracy = 0.9
per_example_cap = 5000
restarts = 3

clean_pass = n_examples
attacked = int(n_examples * clean_accuracy)
worst_case = attacked * per_example_cap * restarts
print(clean_pass + worst_case)

go deeper

for a junior

Should produce the multiplication — examples times per-example cap — and know to pilot on a handful first rather than launching the full set.

for a middle

Should add restarts, the clean pass, early-stop savings on successes, and instrument a counter in the wrapper.

for a senior

Should account for retries, batch padding and vote repeats, and enforce a hard spend ceiling that aborts the run.

for a principal

Should treat the call budget as a contract term agreed before the engagement, with an agreed behaviour when it is exhausted mid-run.

### The arithmetic A query-only run against a metered endpoint has four multipliers and one fixed term, and you can write it down before you start: ``` clean_pass = n_examples # one call each, to find correctly classified inputs attacked = n_examples * clean_accuracy # only correctly classified inputs are worth attacking worst_case = attacked * per_example_cap * restarts ceiling = clean_pass + worst_case ``` `per_example_cap` is the attack object's own evaluation budget — `max_eval` on `ART's HopSkipJump`, `nb_parallel` and `max_iter` together on `ART's ZooAttack`, `max_iter` and `nb_restarts` on `ART's SquareAttack`. Read it off the constructed object, not off a blog post: these defaults were chosen for a model running locally where a forward pass is free, and `ART's HopSkipJump` alone defaults into the five figures of evaluations per example. ### Why the ceiling is close to the truth Newcomers assume the ceiling is wild pessimism and quote an average instead. It is not, and the reason matters. An example the attack *solves* stops early — often far under the cap — and costs little. An example the attack *cannot* solve consumes every query it is permitted before the loop gives up. So the spend concentrates on the failures, and on a genuinely robust target the failures are the majority. Reasoning with a mean call count taken from a soft target is the single most reliable way to be wrong by an order of magnitude on a hard one. ### What inflates the estimate beyond the arithmetic - **Transport retries and rate-limit backoff.** The HTTP client resends below the wrapper, so the attack's internal counter never sees them, but they reached the endpoint and they are billed. - **Vote repeats.** If the endpoint is non-deterministic and you take a majority over three probes to stabilise a near-boundary label, you have multiplied every term by three. - **Batch padding.** A wrapper that always posts fixed-size batches pays for the slots it did not fill. - **A gradient-estimation shim.** Attaching one to a forward-only wrapper turns each attack step into a batch of calls, multiplying the per-example count by one to three orders of magnitude. ### What shrinks it A stratified subsample instead of the full evaluation set. Stopping at first success rather than continuing to minimise the perturbation. A query cache keyed on the exact input, which pays off more than people expect because boundary searches revisit points. And a cheap screening attack first, reserving the expensive family for the examples that survive it. ### Where the number misleads The estimate you can trust is a measured one. Run ten examples with a counter instrumented in the wrapper, read the observed calls-per-example distribution, then extrapolate with headroom. But extrapolating the *pilot mean* is its own trap: a ten-example pilot drawn from easy inputs will mostly show early successes and will under-predict a full run by a large factor, because the full run will contain the failures that spend the whole cap. Extrapolate the ceiling and treat the pilot as a downward correction with a stated confidence, not as the forecast. Equally, a bill that comes in far *under* the ceiling is not automatically good news — it usually means the attack stopped early for a reason you have not diagnosed, such as errors being counted as evaluations, and the resulting success rate is then a measurement of your error path. ### What I would check That a call counter exists in the wrapper and its number reconciles against the provider's usage dashboard for the same window. That a hard spend ceiling is installed which raises an exception and aborts, rather than logging a warning nobody reads. That the cap on the attack object was set explicitly by you, and appears in the run log next to the result. And that the clean pass and any retries are in the estimate — a plan that omits them is short by a predictable, avoidable amount.

  • Why is the worst-case estimate usually close to the actual bill on a hard target?
    Because examples the attack fails on spend their entire per-example allowance, and on a robust target those are the majority.
  • Name two sources of billed calls that the attack's own query counter will not show you.
    Transport retries and rate-limit backoff issued by the HTTP client, and padding in fixed-size batches sent by the wrapper.

saying these in an interview costs you the question

  • Quoting only an average cost per example and ignoring that failed examples consume the whole cap
  • Leaving the library's default query budget in place against a metered endpoint
  • No call counter in the run, so the spend is only known from the invoice
  • Forgetting the clean evaluation pass and any retries in the estimate

context