Why is a white-box attack result treated as a ceiling and a query-only result as a floor?
answer
- The classes are nested by containment
- One dominates the other by construction
- A floor moves up with budget
- The gap is a bill, not a wall
- Only comparable at the same radius
basics
~20 sA white-box adversary holds everything a weaker one could obtain, so its measured success is the most any adversary achieves — a ceiling. A query-only run measures one particular under-informed adversary, so its success can only be raised by more budget or better technique — a floor.
solid answer
~50 sAccess classes are ordered by containment. Anything an adversary with only predictions could learn, an adversary holding weights and gradients already has, so the white-box success rate dominates every weaker adversary's on the same model, perturbation family and radius. That makes it an upper bound on attack success and a lower bound on surviving accuracy. A query-only figure is the opposite kind of object: it reports what one adversary managed with one budget and one method, and a bigger query budget or a better search only moves it up. Reporting only the query-only number understates risk; reporting only the white-box number can overstate the adversary you actually face. The gap between them is not a security boundary — it is the price of obtaining or approximating the gradient, and price is something an adversary can pay.
go deeper
Be ready to state that an adversary holding weights can do everything a query-only adversary can, plus more. That containment is the reason one number is a ceiling and the other is not.
An interviewer expects you to explain why the query-only figure moves upward with budget and technique while the granted-access figure bounds weaker adversaries, and to insist both figures carry their access class, norm and radius.
Show that you read a claim by asking which of the two it is before looking at the number, and that you treat the gap between them as a cost control with an expiry date rather than a boundary you can plan around.
Own the reporting standard: which figure goes to a decision-maker, what caveat travels with it, and how you stop a query-only result being quoted as assurance in a deck that nobody re-reads.
## The classes are nested, and that is the whole argument Access classes in adversarial ML form a containment order. An adversary assumed to hold weights, architecture and input gradients can do anything an adversary holding only returned predictions can do — it can send queries too — plus more. There is no capability the weaker one has and the stronger one lacks. From that single fact the ceiling-and-floor reading follows directly, with no experiment required: - the **white-box attack success rate** is an upper bound on the success rate of any weaker adversary against the same model, under the same perturbation family and radius; - the **accuracy surviving** that white-box attack is correspondingly a lower bound on the accuracy surviving weaker adversaries; - a **query-only attack success rate** is a lower bound on what a query-only adversary can eventually do, because it reflects one budget and one method. ## Why the two numbers are different kinds of claim They are not two estimates of one quantity. The white-box figure answers 'how bad can this get, at most, within the assumed perturbation family'. The query-only figure answers 'how bad did it get for this attacker, with this budget, using this technique'. The first is stable in the sense that a *weaker* adversary cannot exceed it. The second is unstable upward by construction: give the same attacker more queries, a better search, or a stand-in model that agrees with the target near the boundary being attacked, and the number rises. Nothing about a query-only result caps anything. This is why quoting a query-only number alone understates risk, and why an evaluation that reports only 'our model resisted an external attacker' is a weak claim. It is also why quoting only a white-box number can mislead in the other direction: it may describe an adversary far stronger than the one your deployment actually faces, and treating it as the expected case can push spending toward the wrong control. The honest report carries both, each with its access class named: | Reported figure | What it bounds | | --- | --- | | Accuracy under the granted-access attack | The floor for every weaker adversary | | Accuracy under a query-only attack | Nothing; it moves down with budget and technique | | A figure with no access class stated | Nothing at all | ## What the gap between them is worth The distance between the ceiling and the floor is the value of not handing out weights, and it is real. It shows up as the adversary's bill: queries that are metered or paid for, wall-clock spent probing, or the effort of fitting a stand-in that need only agree with the target near the boundary under attack. That last point is why the gap is narrower than people expect — the substitute does not need the target's accuracy or its architecture, only local agreement where the attack is working. What the gap is *not* is a boundary. It is a cost control. Costs are paid by adversaries who care, and they fall over time as techniques improve, so a defence whose entire value lives in that gap is a defence with an expiry date and no floor underneath it. Reducing what a prediction reveals shifts an adversary between families and raises the bill; it does not remove the possibility of attack, and the ceiling is unchanged by it. ## The direction discipline Getting the direction of these claims right is most of the skill. A white-box run that leaves high accuracy means *the attacks that were run failed*; it does not certify anything, and a stronger search later can only lower it. A query-only run that leaves high accuracy means even less. And the ceiling property holds only when the two runs share everything else — same model version, same perturbation family, same radius, same task. Two rows compared across different radii or different norms are not comparable at all, and the containment argument does not apply between them. ## Using it in practice When you read a robustness claim, the first question is not 'what is the number' but 'what access did the attacker have'. Then: is this the ceiling or the floor? If it is the floor, the model may be far worse than stated. If it is the ceiling, the number is the one you can act on — it is the worst case you have evidence for, and everything you actually face sits under it.
- If two runs used different perturbation radii, does the ceiling argument still hold between them?No. The containment argument compares adversaries at the same model, same perturbation family and same radius. Change the radius or the norm and you have changed the threat model, so neither row bounds the other. A larger radius admits more inputs and will generally show more successful attacks regardless of access, which is why a robustness figure without its norm and radius is not comparable to any other.
- A team argues their query-only result shows the model is fine. What is wrong with that?A query-only result bounds nothing. It records what one attacker achieved with one budget and one technique, and more budget or a better search moves it up. It is also partly a measurement of the attacker's information rather than of the model. It is useful as a realism datapoint beside a granted-access result, never as the primary evidence.
- Why is the gap between the two numbers smaller than people expect?Because an adversary approximating the gradient does not need to replicate your model. A stand-in only has to agree with the target near the boundary being attacked, not across the whole distribution, so a much less accurate stand-in still transfers usefully. The remaining gap is mostly a bill in queries and time, and bills get paid.
saying these in an interview costs you the question
- Treats the two numbers as estimates of the same quantity
- Reports a query-only result as evidence the model is robust
- Calls the ceiling a guarantee rather than a measured worst case
- Compares robustness figures across different norms or radii
- Believes withholding weights removes an attack rather than pricing it