In evaluating a model defense, what makes an attack adaptive rather than stock?
answer
- knowledge is granted, not stolen
- aim at what the pipeline decides
- the objective changes, not the step count
- more attacks is not adaptation
- adaptation is authorship, and it costs days
basics
~20 sAn adaptive attack is designed after reading the defense. The adversary is assumed to know the mechanism and its parameters, and aims the attack at the quantity the defended pipeline actually decides on, rather than at the model underneath it.
solid answer
~50 sThree things change, and only the first two are about the attack itself. **The assumption**: the adversary is granted the defense's design and parameters by fiat, with only genuine per-inference randomness left unknown. **The objective**: the attack targets what the deployed system decides on — the aggregated vote of an ensemble, the behaviour averaged over a randomized stage, the output after preprocessing — rather than the bare model the defense was bolted onto. **The effort**: several attack designs are tried, weak ones abandoned, and what was abandoned is recorded. Note what is *not* adaptivity: more optimization steps, more restarts, or more published attacks from the same shelf. Those keep an objective that ignores your defense. This is also why a non-differentiable or randomized component is a cost, not a boundary — the adaptive adversary attacks a stand-in for the awkward stage, or attacks its average behaviour.
go deeper
Know the one-line contrast: a stock attack was written for some other model, an adaptive attack was written after reading yours. Recognise that the second is what an evaluation needs.
Be able to explain the mechanics: which quantity the deployed system decides on once a defense is added, and why an attack that keeps aiming at the undefended output tells you nothing. Say why randomness and non-differentiability are costs rather than boundaries.
Demonstrate that you would reject an internal result whose attack ignores the mechanism, and that you would ask for the formulations tried and abandoned before believing a non-break.
Own the consequence: adaptivity is authorship, authorship is analyst-days, and someone must fund them. Decide what standard of adaptation your claims must meet and how the wording changes when nobody funded it.
### The word does real work *Adaptive* here means adapted **to this defense**, not adapted during the run. It describes how the attack was authored: somebody read the description of the defended system and formulated an attack against that system. The contrast is a stock attack — an off-the-shelf search written for a plain classifier, which treats whatever you added as scenery. ### 1. The knowledge assumption The adaptive adversary is *granted* the defense's design, architecture and parameter settings by assumption. This is an evaluation convention, not a prediction about leaks: it is made so the resulting number cannot quietly include credit for the attacker not knowing what you built. The one thing not granted is genuine randomness drawn per inference — the adversary knows the distribution and the mechanism, not the individual draw. Candidates often get this backwards and argue that a realistic attacker would not know the design. Perhaps not today. But a claim resting on that ignorance decays with every document you publish, every binary you ship, and every hour somebody spends probing the deployed behaviour, and it is not a property of the model at all. ### 2. The objective is re-pointed This is the technically substantive half. A stock attack optimises against some quantity the undefended model exposes. A defense typically changes *which quantity a decision comes from*: | What the defense adds | What the deployed system now decides on | | --- | --- | | A preprocessing or transformation stage | the output after the transform, not before it | | A randomized component | behaviour in expectation over that randomness | | An ensemble with a vote rule | the aggregated verdict, not any single member | | A detector in front of the classifier | the joint outcome of accept-and-classify | An adaptive attack aims at the right-hand column; a stock attack keeps aiming at the left. That is the entire difference, and it explains the failure mode: the stock attack does not fail because the inputs are not there, it fails because it is optimising something the deployed system no longer decides on. ### 3. The awkward-stage argument Two design choices come up constantly as claimed defenses: make a stage non-differentiable, or make it random. **Non-differentiable.** Removing a usable derivative removes one *technique* for searching, not the inputs being searched for. The adaptive response is to work against a differentiable stand-in for the stage, or to use a search that needs only the returned decisions. So the defense buys the attacker a harder search, and it is the search — not the model — that the stock number was measuring. **Randomized.** Randomness makes each observation noisy, so more observations are needed and results vary between runs. It does not create a boundary, because the adversary can target what the system does on average over the randomness. The honest description of both is *this raises the attacker's cost by roughly this much*, which is a claim you can defend, rather than *this stops the attack*, which is a claim the first adaptive attempt refutes. ### 4. Effort, and its record The third component is unglamorous and is where most weak evaluations actually fail. An adaptive evaluation involves trying several formulations, discovering that some do not work, and continuing. A single defense-aware attempt that returned nothing is barely stronger evidence than a stock run. What raises confidence is a record: which formulations were tried, which were abandoned and why, and how much effort was spent before stopping. Effort spent is the unit of a negative result here, in a way it is not for a positive one. ### What adaptivity is not - **Not more steps.** Ten times the optimization effort on an objective that ignores your defense is ten times as much of the wrong search. - **Not more attacks.** A shelf of published attacks shares the assumption that broke all of them. - **Not an access class.** Adaptivity and access are independent axes. An attacker handed weights who runs an unaware attack is doing an unaware evaluation with excellent access; an attacker with only returned decisions who formulated the attack against your ensemble rule is being adaptive on a thin vantage. - **Not automatic.** No suite makes an evaluation adaptive on your behalf, because the adaptation is authorship. ### Why interviewers ask it Because it separates people who can recite the names of attacks from people who understand what an evaluation measures. The follow-up is usually the honest consequence: since adaptivity is authorship and authorship costs analyst-days, the strength of every non-break result you will ever read is capped by how many days somebody was willing to spend, and by what they could see while spending them.
- Does an adaptive attacker also get the random values a defense draws at inference time?No. They are granted the mechanism, the parameters and the distribution, not the individual draw. The standard consequence is that the attack targets the system's average behaviour over that randomness, so randomization raises the number of observations needed and the variance of the result. It is a cost multiplier on the attacker, not a boundary they cannot cross.
- A stage in our pipeline is non-differentiable. Doesn't that stop a gradient-based attacker outright?It stops one search technique, not the existence of inputs the system reads wrongly. An adaptive adversary either works against a differentiable stand-in for the awkward stage or switches to a search that needs only returned decisions. If your reported number came from a gradient-following attack failing on that stage, the number is describing the technique, not the model.
- How many adaptive attacks are enough for a defense paper or an internal claim?There is no fixed count, because the strength of a non-break scales with effort and vantage rather than with a number of runs. What a reader needs is which formulations were designed against this defense, which were abandoned and why, what access was granted, and how much effort was spent before stopping. That makes the bound auditable instead of implied.
saying these in an interview costs you the question
- Thinks adaptive means rerunning a suite with more steps
- Treats randomness or non-differentiability as a boundary
- Assumes the attacker must steal the design rather than be granted it
- Counts attacks instead of designing one
- Attacks the bare model rather than the deployed pipeline