skip to content

A 55% robust-accuracy figure came from a ten-step, one-restart attack - what do you ask for?

level: seniorimportance: should knowfreq 38%

answer

  1. ten steps is where they stopped paying
  2. ask for the number as a curve
  3. steps and restarts buy different things
  4. steps compare only within one access class
  5. still falling means it never converged

basics

~20 s

Ask for the figure as a curve over optimisation steps and restarts, run until it stops falling, plus the attack access class and the stated edit budget. Ten steps is a stopping point, not a result.

solid answer

~40 s

Ten steps and one restart describe how hard the adversary tried, and that is the parameter that hides the most. Ask for the **saturation curve**: robust accuracy against optimisation steps, and against restarts, until it flattens - a figure taken while the curve is still descending has not converged and its distance from the truth is unknown. Ask what the attack could see, because ten steps under full access and ten steps against a returned decision are different adversaries and the step counts are not comparable across access classes. Ask how the edit budget was defined, since on source diffs there is no radius and "semantics-preserving" names a category rather than a limit. And treat an attack that does *worse* with more access as a sign the evaluation broke rather than the model held.

go deeper

for a junior

Know that an attack is a search that gets stopped somewhere, and that how long it ran is part of what the reported percentage is measuring.

for a middle

Explain why iterating beats a single step at the same budget, and why steps and restarts are separate parameters that a report must give separately.

for a senior

Demonstrate the reviewer's move: ask for the figure as a curve over steps and restarts until it flattens, qualify it by access class, and call an unconverged evaluation a finding.

for a principal

Set the bar your organisation accepts - a converged curve, a stated access class and edit budget - and decide who funds re-running the evaluation when a supplier will not.

## Why the step count is the headline, not a footnote An attack against a trained model is an optimisation: it searches the permitted modifications of an input for one the model reads wrongly. Every reported robust-accuracy figure is therefore the score of a search that was stopped somewhere. "Ten steps, one restart" tells you exactly where it was stopped, and that single fact often accounts for more of the reported number than anything about the defence. The reason is structural. Iterating beats a single step at the same budget because each small step re-reads the model and re-aims before moving again, staying inside the permitted set the whole way. Extending that iteration keeps finding inputs the shorter run missed - until it stops, at which point the number has **converged** and further search buys the adversary nothing. Before that point, the figure is a ceiling of unknown height. ## What to ask for, in order **1. The saturation curve.** Robust accuracy reported at several step counts - and separately at several restart counts - so a reader can see whether it has flattened. A defensible write-up shows the curve and states that it plateaued; a write-up quoting one figure at ten steps has reported the point where they stopped paying, not the point where the adversary stops gaining. **2. Restarts, separately from steps.** They buy different things. Steps deepen one search from one starting point; restarts sample fresh starting points, so a search that stalls in a poor region gets another chance. One restart is a meaningfully weaker adversary than several, and the two parameters are not interchangeable - a report that gives only their product has given neither. **3. The access class.** Access is an assumption, not an event: weights and gradients available, a score returned, or only the decision returned. A step under full access moves along a direction read from the model itself; a step under decision-only access is one probe of a boundary walk. Ten of one is nothing like ten of the other, so **step counts only compare within an access class**, and a report that omits the class has made its step count uninterpretable. Where the access is metered, the honest cost unit is queries as well as steps. **4. The budget definition.** For a source diff there is no small epsilon to quote. The budget must be stated as which edit operations the adversary could draw from and how many per input. Without it, the percentage counts failures over a space nobody has described. **5. Clean accuracy and the test set**, so the figure is not being lifted by a detector that simply flags more. ## Reading the answer you get back A few shapes in the returned numbers are worth more than the numbers themselves: - **Still falling at the largest step count reported.** The evaluation has not converged. The correct interpretation is that the true figure is somewhere below what was printed, by an amount nobody has bounded. - **A big gap between one restart and several.** The search was getting stuck; the single-restart number was measuring the search's bad luck. - **The less-informed attack scoring lower than the better-informed one.** Something is wrong with the evaluation rather than right with the model - an attack given more information should never do worse. Diagnosing that failure mode is its own topic, but spotting it is part of reading the table. ## Why this is a senior question Juniors are asked what robust accuracy means. Seniors are asked what to do when the write-up is thin and the model is not in your hands. The move is not to argue about whether 55% is a good number; it is to name the parameter that would change it and to ask for the figure as a function of that parameter. If the authors cannot produce the curve, you have learned that the evaluation was never pushed to convergence, which is a finding you can write down. If they can, you have learned the real ceiling. Either way you have converted an unqualified percentage into something with a shape. And the cost side matters in the report too: an attack's step count is the adversary's expense, so a converged curve tells you not just where the model fails but roughly what it costs an adversary to make it fail - which is the quantity a defender actually has to reason about.

  • What does it mean if robust accuracy is still falling at the largest step count reported?
    The evaluation has not converged, so the printed figure is a ceiling whose distance from the truth nobody has bounded. The honest report is the curve plus a statement that it had not flattened; the honest reading is that the model's real robustness at that budget is somewhere lower by an unstated amount.
  • What do restarts buy an adversary that more steps do not?
    Steps deepen a single search from one starting point, so a run that stalls in a poor region stays stuck. Restarts sample fresh starting points and give the search another chance at inputs the first run failed on. A one-restart figure can sit several points above the multi-restart figure on identical weights.
  • How does the access class change what 'ten steps' means?
    Under full access a step follows a direction read off the model, so ten steps is already a serious search. Under decision-only access each step is one probe of a boundary walk and ten of them barely start. Step counts are therefore only comparable within an access class, and where access is metered the report should also state queries.

saying these in an interview costs you the question

  • Treats step count as an implementation detail
  • Accepts a figure whose curve has not flattened
  • Compares step counts across different access classes
  • Assumes one restart is as good as several
  • Argues about whether 55% is good instead of what moved it

context