A supplier says their produce-recognition classifier is 'adversarially robust' — what must the claim name?
answer
- robustness is never unconditional
- a number describes an experiment somebody ran
- what could they see, touch, and when
- labels only, or weights and gradients
- a radius, a query count, a poisoned fraction
basics
~20 sIt must name the adversary it was tested against: what they could see (a returned label, a score, or the weights), how much they could touch (queries, a perturbation radius, poisoned rows), and at which stage they acted.
solid answer
~50 s'Adversarially robust' on its own is not a readable claim, because robustness is always robustness *against a stated adversary*. Three things fix it. What the adversary was allowed to **see**: only the top-1 label the endpoint returns, a confidence score, or the weights and input gradients themselves. How much they were allowed to **touch**: a number of paid queries, a perturbation radius in a named norm, or a fraction of rows written into the training set. And **when** they acted: at inference against the deployed model, or earlier, into the data it was trained on. A shared classification of ML attacks — NIST's adversarial-ML taxonomy is the usual reference — exists to force those three into the open. Until the deck states them, the number tells you which experiment somebody ran, not what the model resists.
go deeper
Be ready to say out loud that robustness is always robustness against somebody, and to list the three things a claim must name: what the adversary could see, how much they could change or write, and at which stage.
An interviewer expects you to explain why each slot changes the number — a label-only adversary and a weights-in-hand adversary are doing different work, and a radius and a query count restrict different resources.
Show that you would refuse to read the figure before the assumptions arrive, and that you would ask whether the assumptions match the adversary your own deployment actually exposes.
Own the consequence: unstated assumptions make results unrepeatable across quarters and vendors, so require the three slots in the acceptance criteria rather than arguing about numbers afterwards.
## Robustness is a property of an experiment, not of a model When a deck says a classifier is 'adversarially robust', it is reporting that some adversary, working under some restriction, failed to break it some fraction of the time. Change any part of that sentence and the number changes, often by tens of points. So the number without the sentence is not a weaker claim — it is not a claim at all. The reader cannot even tell whether it is impressive. This is why a shared vocabulary for describing ML attacks exists. Its job is not to enumerate attacks; its job is to give three slots that every result has to fill in before it can be set beside another result. ## The three things a claim has to name **1. What the adversary could see.** The realistic settings are a small ladder, and they are assumptions granted by the evaluator, not events that happened: | The adversary holds | What that opens | | --- | --- | | the returned top-1 label only | search that walks the decision boundary from a point already classified the way they want | | a confidence score or full probability vector | a gradient direction bought by probing and differencing returned scores | | weights, architecture and input gradients | the direction read straight off the loss with respect to the input | | write access to the training corpus or labelling queue | influence over what the model becomes, not just what it answers | A claim that says 'we tested black-box' has still not answered this: the label-only and score-returning settings are both black-box and they cost an attacker completely different amounts. **2. How much they could touch.** This is the budget, and it comes in units that suit the setting: a count of paid queries; a radius in a named norm (every coordinate may move a little, or a few coordinates may move a lot, or total change is capped — these are different adversaries); a count or fraction of rows they got into the training data; a share of participants in a distributed training scheme. A stated attack with no budget is the name of a paper, not a threat model. **3. When they acted.** An adversary who perturbs an input at inference time and an adversary who writes into the training corpus are not stronger and weaker versions of each other; they need different access, they are caught by different controls, and a result about one asserts nothing about the other. A deck reporting only inference-time numbers for a model that is retrained on data the public can influence has left the stage that matters to you completely unmeasured. ## Why naming the attack is not a substitute Candidates often answer this question by naming methods — a single-step gradient attack, an iterative one, a boundary-walking one. That describes the tool, not the assumptions. The same named method run with weights in hand and run against a label-only endpoint are different experiments with different results, and two methods run under identical assumptions are broadly comparable even if their names differ. The assumptions are the durable part; the method inventory turns over. ## What to do with a claim that lacks the three Ask for them, in writing, before you read the number. In practice the answer arrives in one of three shapes. Sometimes the assumptions exist and were simply omitted, and the claim becomes readable. Sometimes the evaluation granted the adversary very little — a handful of queries, a tiny radius, no training-stage access at all — and the number was never a strong statement. And sometimes nobody can say, which is itself the finding: an evaluation whose assumptions cannot be reconstructed cannot be repeated, so it cannot be checked and cannot be compared to next quarter's run of the same model. ## Readings that are wrong - 'Robust' is not a property the model carries around; it is always relative to a named adversary and budget. - A high number does not mean the model resists attacks; it means the attack that was run, under the assumptions that were granted, mostly failed. - A perturbation that flips a classifier is not noise. Random change of the same size essentially never flips a trained model; the attack works because it is a *direction*, and that is exactly why the size alone is meaningless without the direction-finding power you granted. - 'Imperceptible to a human' is not a budget. It is not measurable, it does not transfer between data types, and on discrete data — text, binaries, network records — there is no small change at all, only edits that preserve behaviour.
- The deck adds 'tested black-box' — is that enough to read the number?No. Black-box is one access class covering two very different adversaries: one who gets a confidence score back and can buy a gradient estimate by probing and differencing it, and one who gets only a label and must walk the boundary. It also says nothing about how many queries they were allowed, or which stage they acted at. Two of the three slots are still blank.
- Why isn't 'we ran the strongest known attack' a substitute for stating assumptions?Because 'strongest' is only defined inside an access and budget setting. The same method is a different experiment with weights in hand than against a label-only endpoint, and its strength depends on how many steps and how large a change it was allowed. Naming the method tells the reader which tool was used, not what the tool was permitted to do.
- A supplier says the model is robust because attacks were 'imperceptible' — what do you push back on?Imperceptibility is not a budget. It has no unit, no threshold anyone can re-measure, and it does not exist at all for data a person does not look at — a tabular claim record or a network flow. Ask for the norm and radius they actually enforced, or for the query and row counts if the setting has no radius.
'Waterproof' means nothing until the label says to what depth and for how long. Two jackets both stamped waterproof can be honest and still not comparable.
saying these in an interview costs you the question
- Treats 'robust' as a fixed property the model carries
- Thinks black-box means the attacker learned nothing useful
- Ranks two robustness numbers taken under different assumptions
- Names an attack method instead of the adversary's access and reach
- Calls an adversarial perturbation 'noise' of a certain size
- Assumes an inference-time result covers a training-stage adversary