skip to content

Two robustness claims for the same image classifier use different units — can you rank them?

level: middleimportance: should knowfreq 45%

answer

  1. two honest numbers, no ordering
  2. money, geometry and write access
  3. there is no exchange rate
  4. the stricter grant makes the bolder claim
  5. fix one adversary, then sort

basics

~20 s

No. A query count, a perturbation radius and a poisoned fraction restrict different resources and do not convert, so two claims can both be true and still not rank. Ranking is only meaningful inside one fixed threat model.

solid answer

~50 s

You cannot, and the reason is that the budgets are stated in units that do not convert. One deck says the model survived an adversary who could send images and read back only the top-1 label, over five thousand paid queries. The other says it survived an adversary handed the weights but confined to a small per-coordinate perturbation radius. One of those prices *search effort*; the other bounds *how far the input may move*. There is no exchange rate between them, so both numbers can be simultaneously correct and neither is evidence about the other. What you can say is which adversary each number describes, and that the weaker-looking figure was obtained under the stronger access — which usually makes it the more informative of the two. If you need a ranking, you have to fix one adversary and have both re-tested under it.

code

text · 11 lines
text
Bid A   robust accuracy 71%
  adversary sees   : returned top-1 label only
  adversary reach  : 5,000 queries per input
  stage            : inference, deployed model

Bid B   robust accuracy 68%
  adversary sees   : weights, architecture, input gradients
  adversary reach  : L-infinity radius 8/255
  stage            : inference, deployed model

... neither deck reports any training-stage result

go deeper

for a junior

Know that two robustness percentages are only comparable when the same adversary and the same budget produced them, and that a lower number is not automatically a worse model.

for a middle

Be ready to explain why a paid query count and a perturbation radius restrict different resources, and why granting the adversary the weights makes a result stricter rather than weaker.

for a senior

Demonstrate the move a reviewer makes: state what each claim covers, refuse the manufactured ranking, and specify one adversary both parties must re-test under if a ranking is genuinely required.

for a principal

Own where the unit gets pinned — in the request for proposal and the acceptance criteria — and accept the schedule cost of a re-test rather than shipping a comparison that was never real.

## Two true numbers that do not rank Procurement wants a single column it can sort. Adversarial-ML results resist that, and not because the field is sloppy — because the restrictions that define an adversary are genuinely different kinds of quantity. - A **query count** prices search: how many times the adversary may interrogate the deployed model before the bill or the rate limiter stops them. It is money and time. - A **perturbation radius in a named norm** bounds geometry: how far the input is allowed to move, and in what shape. Every coordinate may move a little; or a few coordinates may move a lot; or total change is capped. Those are three different adversaries with three different real-world counterparts. - A **poisoned fraction or row count** bounds write access to the training data. - A **share of participants** bounds influence in a distributed training scheme. - A **privacy parameter and its accounting** bounds what a released model may reveal about one record. None of these converts into another. Asking whether five thousand queries is 'more than' a small per-coordinate radius is like asking whether an hour is more than a kilogram. ## Why the lower number can be the stronger result Access classes are not equally generous. Granting the adversary weights, architecture and input gradients is the strictest test, and it is granted deliberately so that the result does not silently measure your own obscurity instead of your model. A figure obtained under that grant is a stronger statement than the same figure obtained against an adversary who only ever saw returned labels. So a deck reporting 68% under weights-in-hand access has, on the access axis, made a bolder claim than a deck reporting 71% against a label-only adversary — provided the second axis, the reach, is also comparable, which is exactly what a radius and a query count fail to be. The honest reading of the pair is therefore two sentences, not one ordering: *this model withstood a well-informed adversary confined to small changes; that model withstood a poorly-informed adversary given a fixed number of tries; neither deck says anything about the other's adversary.* ## The temptation to manufacture a comparison Three bad moves show up in real reviews. **Averaging.** Combining the two figures into one score invents a measurement nobody made. The average is not a number about either adversary. **Normalising by eye.** 'A five-thousand-query black-box attacker is roughly as strong as a small-radius white-box one' is a guess dressed as a conversion, and it is usually wrong in a direction nobody checks: hiding scores raises an attacker's bill and removes the family of attacks that estimates gradients by differencing returned scores, but it does not remove the family that needs only the returned label and walks the boundary. Withholding information is a cost control, never a boundary. **Declaring the whole thing unmeasurable.** The opposite failure. Robustness is measurable; it is just measurable *relative to a stated adversary*, and a review that gives up stops asking the one question that would have made the claims usable. ## What actually produces a ranking Fix the adversary yourself and make both suppliers report under it: one access class, one reach expressed in one unit, one stage, one dataset. Then the numbers are on the same axis and sorting them means something. That costs schedule and it costs goodwill with a bidder who has to re-run work, which is why it belongs in the request for proposal rather than in the review meeting. Where a re-test is impossible, the deliverable is not a rank but a written statement of what each claim covers and what it leaves blank — and, importantly, whether either adversary is the one your deployment exposes at all. A model retrained on data your own operating environment collects faces a training-stage adversary that neither inference-time number touches. ## The related trap The same incommensurability appears within a single supplier's own history. A team that changes the reported unit between releases — a radius one quarter, a query budget the next — has produced a trend line that measures nothing. Pinning the unit is what makes a robustness number a metric rather than an anecdote.

  • Both decks are for models you would retrain on your own collected data — what does that change?
    It makes both numbers partly beside the point. Each covers an adversary who perturbs an input at inference time; retraining on data your environment collects exposes an adversary who can write into the training set instead, which neither deck measured. The stage axis is the one your deployment moves, and it is blank on both sides.
  • Is a label-only result automatically weaker evidence than a weights-in-hand one?
    On the access axis, yes — granting weights is the stricter test, and it is granted on purpose so the result does not measure obscurity. But access is only one axis. A weights-in-hand result at a radius so small that nothing meaningful fits inside it is a weak claim despite the generous access, so read the reach alongside it.
  • Procurement wants one sortable column anyway — what do you give them?
    One threat model, written down and identical for both bidders, with the number reported under it. If that cannot be bought before the deadline, give them a two-line statement per bid saying which adversary it covers and which axis it leaves blank, and let them sort on criteria that are actually comparable.

Two fuel-economy figures, one measured on a motorway and one in city traffic. Both honest, neither ranks the cars, and averaging them invents a third number nobody measured.

saying these in an interview costs you the question

  • Sorts robustness figures taken under different adversaries
  • Averages two numbers from different threat models
  • Invents a conversion between query count and radius
  • Says robustness simply cannot be measured objectively
  • Assumes hiding confidence scores blocks black-box attacks

context