skip to content

Rules of engagement give you a metered HTTPS endpoint that returns one predicted label per query, with a per-day request cap, and no model artefact. A teammate reports a very high success rate produced by a gradient-based attack class in an adversarial-robustness toolkit, run against a copy of the model downloaded from a public hub. How do you handle that number, and which attack class do you actually run against the granted surface?

level: seniorimportance: should knowfreq 48%

answer

  1. public checkpoint is not the served model
  2. label only means decision-based only
  3. pilot for queries-to-success first
  4. rate cap as the binding constraint
  5. two runs, two attacker positions, two labels

basics

~20 s

Do not report it against the deployed endpoint: it measured a different artefact under access nobody granted. Relabel it as a lab run on a public copy. For the granted surface, pick a decision-based class needing only the top-1 label, pilot a few examples to measure queries per success, and compare that to the daily cap.

solid answer

~50 s

Two separate problems. **The number**: it describes a public checkpoint under full white-box access, not the served model behind the endpoint. The served model may be fine-tuned, quantised, wrapped in different preprocessing, or fronted by a filter. Keep the run, but relabel the finding to name the artefact and the access level it assumed, so the client is not told their endpoint was tested when it was not. **The rerun**: the granted surface returns a bare label, so the only classes that can run are decision-based ones. Their cost is queries, and it is large per example, so treat the request cap as the binding constraint: pilot on a handful of inputs, measure queries-to-success and the perturbation size reached, then compute how many examples the cap allows. If the honest answer is 'three examples per day', that itself is a reportable result — the endpoint's rate limit is a real control against this attacker position, and saying so is more useful than a borrowed white-box percentage.

go deeper

for a junior

Should at least notice the tested model is not the deployed one, and that a label-only endpoint cannot supply gradients.

for a middle

Names decision-based classes as the only option on that surface and knows their cost is measured in queries.

for a senior

Pilots for queries-to-success against the daily cap, treats a truncated search correctly, and reports the two runs as two distinct attacker positions with different remediations.

for a principal

Makes the artefact-vs-endpoint distinction a reporting rule for the team, and negotiates query budget and alerting expectations into the rules of engagement before the work starts.

### What actually happened A teammate wanted a gradient-based class from an adversarial-robustness toolkit — ART, Foolbox, Torchattacks — and the class needs a differentiable model. The granted endpoint cannot supply one, so the missing precondition was filled by downloading a checkpoint from a public hub. The class then ran perfectly: it had weights, it had a loss, it produced a high success rate. Every step worked; the object under test was swapped. **Why the number is unactionable.** A finding is actionable when the client can name the attacker it describes and change something in response. This one names neither. The tested artefact is a public checkpoint; the deployed one is a served system that may be fine-tuned on client data, quantised to int8, wrapped in a different preprocessing chain (resize, normalisation, tokeniser version), batched, cached, and fronted by an input filter or a confidence threshold that never appears in the checkpoint at all. Any one of those changes the loss surface the gradient search exploited. No part of the toolkit detects the divergence — a wrapped model is a callable, and the class calls it. The access level is the second swap. The reported percentage assumes an attacker who holds the weights. The engagement's attacker holds one HTTPS route that returns a string. ### Selecting for the granted surface Work backwards from the output you are actually given: - one label per query -> **decision-based** boundary search only; - add full confidence scores -> score-based classes open up, and per-example cost falls sharply; - add weights -> gradient classes, and cost moves from queries to GPU time. The grant is the first case, so class selection is settled before anyone's preference about which attack is "strongest" enters the conversation. ### Budgeting the run, in numbers Decision-based search is query-bound, and the daily cap is therefore the binding constraint — not compute, not storage, not model size. The procedure is arithmetic, and it must happen before the main run: 1. **Pilot** on a handful of inputs. Record queries-to-success per example and the perturbation size reached. 2. **Divide.** Cap / queries-per-example = examples per day. If a pilot shows ~8,000 queries for a usable perturbation and the cap is 5,000 requests/day, the honest sentence is *one example every day and a half*, and the sample you can afford in a two-week engagement is single digits. 3. **Cost it in money and hours too.** At a metered price per call, 8,000 queries is a real line item; and an engineer babysitting a rate-limited loop for days is the larger cost. 4. **Decide the cut-off rule up front.** When the budget runs out mid-search, the example reports *the perturbation size reached at cut-off, marked truncated* — not "attack failed". Conflating a truncated search with a failure understates the risk and is the easiest way to hand a client a falsely reassuring number. A sample of three examples cannot support a percentage. Report the per-example results and the budget that bounded them; "3 of 3 examples flipped at a perturbation of X, at ~8,000 queries each, against a 5,000/day cap" is a stronger sentence than any success rate you could compute from it. ### Where the numbers mislead - **The borrowed white-box percentage** reads as the endpoint's exposure. It is a bound on a *different, more permissive* attacker. - **A low decision-based success rate** reads as robustness when it is often just budget exhaustion. The denominator is examples attempted; the missing variable is queries allowed per example. - **A rate cap looks like a nuisance** and is in fact a finding: if the granted surface makes the attack cost days per example, that is a real, reportable control — and it evaporates the moment the client raises the limit or an attacker parallelises across accounts, which is worth stating. ### What you check Whether identical queries return identical labels — a stochastic sampler, an A/B split or a probabilistic filter breaks a boundary search's feedback signal, and you then need repeated queries per probe, multiplying the cost by that repetition factor. Whether the pilot's traffic tripped any client-side alerting or throttling (if it did, say so; if it did not, that is also a finding). And the report should carry both runs, separately labelled: (a) endpoint, decision-based, N examples, queries per example, perturbation reached, bounded by the cap; (b) public checkpoint, gradient class, lab conditions, offered as the weight-leak scenario. Two attacker positions, two remediations — query-rate controls and pattern detection for the first, artefact control and provenance for the second.

  • The client offers to hand over the served weights for the duration of the test. Does that change the deliverable?
    It adds a run, it does not replace one. The white-box run on the real artefact is now meaningful as a stronger-attacker bound and is far cheaper to run, but the endpoint-level number is still the one that describes today's exposed surface.
  • Identical queries to the endpoint sometimes return different labels. What does that do to a decision-based class?
    It breaks its feedback signal — the search reads noise as a boundary crossing. You need repeated queries per probe to stabilise the answer, which multiplies the query cost, and that multiplier belongs in the report.

Testing the public checkpoint and reporting it as the endpoint's result is like picking the lock on a display model in the showroom and writing it up as a break-in at the client's building. The technique is genuine; the door was not theirs.

saying these in an interview costs you the question

  • Reporting the white-box lab number as the endpoint's result because it is more impressive.
  • Assuming the public checkpoint and the served model behave identically.
  • Starting a decision-based run at scale without a pilot to measure queries-to-success.
  • Reporting a budget-truncated search as a failed attack rather than as the perturbation reached at cut-off.

context