A query-based attack loop from an adversarial robustness toolkit assumes each call to the target answers quickly and consistently for the same input. You are pointing it at a rate-limited, non-deterministic hosted endpoint that bills per call. Of retries, throttling and answer variability, which spend your call budget, which spend only wall-clock, and what do you do about each?
answer
- retries and votes cost money
- throttling costs calendar
- variability costs validity
- counter at the transport layer
- canary replay detects target drift
basics
~20 sRetries and repeated probes spend billed calls. Throttling and latency spend wall-clock, which limits how many examples you finish per day. Variability is worst: a boundary answer that flips makes the search oscillate. Handle it by voting over repeated probes, which multiplies the budget, or requesting a deterministic setting if the endpoint offers one.
solid answer
~50 sSeparate the three, because the mitigations differ. **Retries** are real billed calls that the attack's own counter usually does not see, since the HTTP client issues them below the wrapper. They inflate the invoice without advancing the search. Instrument the counter at the transport layer, not in the attack. **Throttling and latency** cost wall-clock, not money. They set the throughput ceiling for the engagement: cap times examples divided by achievable calls per second is your calendar, and that number, not the cost, is usually what forces a smaller evaluation set. **Answer variability** is the one that corrupts the result. Query attacks assume a stable decision function; if the same input returns different labels near the boundary, a boundary-walking search wanders and a score-based search reads noise as signal. The fixes are to pin any determinism control the endpoint exposes, or to vote over repeated probes — which multiplies the per-example budget by the vote size, and must be in the estimate before the run starts.
go deeper
Should know that rate limits slow a black-box run down and that retries still cost money.
Should separate billed calls from wall-clock and note that unstable answers confuse a query-based search.
Should instrument counting at the transport layer, quantify the throughput ceiling, and price a voting scheme against the shrunken per-example budget.
Should negotiate rate and determinism terms with the target's owner before the engagement, and require canary checks for drift on any long run.
### Three failure modes that feel alike and are not A query-based attack loop is written on the assumption that a call to the target is fast, free of side effects, and returns the same answer for the same input. A hosted, rate-limited, non-deterministic, per-call-billed endpoint violates all three. The discipline is to sort what you observe into what costs **money**, what costs **calendar**, and what costs **validity**, because each has a different fix and only the third can silently poison the result. ### Costs money Every request that arrives at the endpoint is billed: attack probes, transport-level retries after a transient failure, resends after a 429, and each repeat in a voting scheme. The trap is a counting mismatch. The attack's internal counter records the probes *it decided to issue*; the invoice records the requests that *arrived*. A retrying HTTP client sits below the wrapper, so a run can be 30% over budget with a counter that looks perfect. Put the counter in the transport path — the requests adapter, the session hook, the proxy — so the two numbers reconcile, and install a hard ceiling there that raises and aborts. A warning in a log is not a budget control. ### Costs calendar Per-minute quotas, concurrency limits and per-call latency do not add requests to the bill; they cap throughput, and throughput decides whether an affordable plan is a finishable one. The arithmetic is blunt: examples times per-example cap divided by achievable calls per second is your wall-clock. Five hundred examples at a 5,000-evaluation cap is 2.5 million calls; at ten calls per second that is about three and a half days of continuous running, and at one call per second it is a month. This constraint usually binds before money does, and it is the real argument for a stratified subsample rather than the full evaluation set. ### Costs validity Two things sit here. **Answer variability**: a query attack assumes a stable decision function, and near the boundary — which is exactly where the search spends its time — a non-deterministic target may return different labels for the same input. A boundary walk then wanders, a score-based search reads sampling noise as gradient signal, and the failure looks like robustness. **Target drift**: a served model that is updated, A/B routed, or fronted by a cache mid-run is not one decision function, so the run's early and late phases are not comparable. Detect it cheaply by replaying a small fixed canary set at intervals through the run and checking the answers are stable; a shift there invalidates cross-phase comparisons and any before/after claim built on them. ### The voting arithmetic, which is the part people get wrong If you take a majority over three probes to pin down a near-boundary label, every query the attack thinks it is making is three billed calls. Under a fixed spend that means the *effective* per-example evaluation budget is a third of the nominal one, or the example count is a third, or the bill is triple. It has to come out of one of those three, and the decision belongs in the plan before the run, not in the post-mortem. Usually the trade is worth it: an unstable decision function does not make a large cheap run noisy, it makes it meaningless, and a third of the examples measured properly beats all of them measured against a coin flip. ### Where the number misleads A run against a throttled endpoint often reports a suspiciously low attack-success rate. Before believing it, ask what fraction of the evaluation budget went to non-answers. If the attack loop decrements its budget on 429s and timeouts — many do — then the caps were never spent on information, and the "robustness" you measured is your own error rate. Conversely, a run against a non-deterministic target can report a suspiciously *high* success rate: an input counted as adversarial because one sampled answer flipped may not flip again, and re-verifying candidates with repeated probes routinely deletes a meaningful share of them. ### What I would check and report Calls issued versus calls billed, reconciled against the provider dashboard. Achieved calls per second and the resulting projected finish date. Vote size and the effective cap after dividing by it. Canary stability across the run. And a re-verification pass over every candidate the run called a success, at the vote size, before any of them appear in a report.
- You add a three-probe majority vote to stabilise labels. What must change in the run plan?The effective per-example query cap drops to a third of nominal, so either the cap, the example count or the budget has to be raised to keep the same reach.
- How do you detect that the target changed underneath a multi-day run?Replay a small fixed canary set at intervals and compare answers; a shift means the run's phases are not comparable and the earlier results need re-baselining.
Voting over three probes to trust a wobbly answer is like insisting on three quotes before believing a price: the number you end up with is better, but you can only shop for a third as many items on the same budget. The mistake is planning the shopping list before deciding you need the three quotes.
saying these in an interview costs you the question
- Counting queries inside the attack only, so retries never appear until the invoice does
- Treating a non-deterministic target as a stable decision function and reporting the resulting numbers
- Adding a vote over repeats without shrinking the per-example budget to match
- Assuming the target is unchanged across a run that lasts days