skip to content

A code-review gate reports 60% robust accuracy against edited diffs - what does that claim?

level: juniorimportance: must knowfreq 66%

answer

  1. it describes the contest, not the model
  2. a ceiling, never a floor
  3. 60% of what, under whose budget
  4. attack, access, steps, edit budget
  5. flag every diff and score 100%

basics

~20 s

Robust accuracy of 60% means one specific attack, run at one stated effort inside one stated edit budget, failed on 60% of the tested inputs. It bounds that attack, not the model's resistance to a better-resourced adversary.

solid answer

~50 s

Robust accuracy is the share of a test set the model still gets right when an adversary is allowed to modify each input inside a stated budget first. So 60% is a fact about the attack that was run: which edits it was allowed, how hard it searched, and what it could see of the model. Because an attack is a search, a better search can only find more failures - an empirical robust-accuracy figure is an **upper bound** on true robustness, never a floor and never a guarantee. That means the number quoted alone tells you the ceiling and hides the gap. A usable report names the attack and its access class, its optimisation steps and restarts, the edit budget an input was held to, and clean accuracy beside it - otherwise a detector that flags every diff scores 100% robust and is worthless.

go deeper

for a junior

Be ready to say in one sentence that robust accuracy scores the attack that was run, inside a budget somebody chose, and that it is an upper bound rather than a guarantee.

for a middle

An interviewer expects you to list what must be reported beside the figure: the attack and its access class, optimisation steps and restarts, the permitted edits and their count, and clean accuracy.

for a senior

Show that you would refuse to act on a bare percentage - name the columns you would demand, and say what you would do when the evaluation cannot be re-run against a model you do not hold.

for a principal

Own the reporting standard: decide what your organisation requires before any robustness figure is allowed to gate a release, and who pays for the evaluation that produces it.

## What the number is **Robust accuracy** is the fraction of a test set that a model still classifies correctly *after an adversary is allowed to modify each input, inside a stated budget, before the model sees it*. For a merge-gating code model - a learned detector deciding whether a source diff is flagged for human review - the adversary is whoever submits the diff, and the modification is a set of edits that leave the diff's behaviour as they intend while changing how the detector reads it. "60% robust accuracy" therefore means: on 60% of the tested diffs, **the attack that was actually run** failed to move the gate's decision. That phrasing is the whole lesson. Robust accuracy is not a property the model carries around. It is the score of a contest, and a score is meaningless until you know who was on the other side. ## Why it is a statement about the attack An attack is a **search** over the modifications the adversary is allowed to make, looking for one the model reads the wrong way. Searches differ enormously in quality: how much of the model they can see, how many optimisation steps they take, how many times they restart from a fresh starting point, how well the search is suited to the input type. A poor search fails to find modifications that exist. This gives the single most important structural fact about the number: **a strictly stronger search finds at least as many failures as a weaker one, so the reported figure can only fall.** Empirical robust accuracy is an *upper bound* on the quantity people think it is ("the share of inputs no adversary inside this budget can flip"). The distance between the reported ceiling and the truth is set by how good the attack was - and that distance is exactly what a bare percentage hides. History here is unkind: a long series of published defences reported high figures and lost most of them to later, better-optimised attacks against the same models at the same budgets. ## The four things that must sit beside it 1. **Which attack, and what it could see.** Access is an assumption, not an event: weights and gradients in hand, a returned score, or only the returned decision. An attack given less access is a weaker search and therefore reports a *higher* number, so a figure without its access class may be measuring how little the evaluator was allowed to know rather than how the model behaves. 2. **How hard it searched** - optimisation steps and restarts. Step count alone routinely moves such a figure by tens of points, which makes it a headline parameter, not an implementation detail. 3. **The budget the adversary was held to.** For continuous inputs that is a norm and a radius. For a source diff there is no small epsilon - you cannot move a token by 0.01 - so the budget has to be stated as *which* edit operations were permitted and *how many* per input. "Semantics-preserving edits" names a category, not a budget; two evaluations using that phrase can be describing adversaries an order of magnitude apart. 4. **What it was measured on**, and clean accuracy beside it. Without clean accuracy the metric is trivially gameable: a detector that flags every diff is never fooled, scores a perfect robust accuracy, and is useless. The pair is the minimum honest report. ## What the number never claims - It is **not a guarantee**. A guarantee - a certified radius, say - is a different kind of object with its own stated confidence and cost; an empirical figure is the outcome of one experiment. - It is **not about every adversary**, only the one that ran. - It says nothing about a *different* budget than the one it was measured under; carrying robustness between perturbation families is a separate question with its own answer. - A high figure proves the attack that was run failed. It does not prove the model held. ## How to use this in an interview When a number is quoted at you, do not argue about whether 60% is good. Ask what it is 60% *of*: which attack, at what access, at how many steps and restarts, inside which stated edit budget, against what clean accuracy. If those columns are absent, the correct answer is that the figure describes no adversary at all, and the right next move is to obtain them or to re-run one attack of your own choosing under a budget you state. Interviewers ask this because reading a robustness claim correctly is most of the job, and because the wrong reading - treating the percentage as a model property - is the default one.

  • Can a stronger attack ever push a reported robust accuracy up?
    No. Each attack finds some subset of inputs it can flip inside the budget, and a strictly stronger search finds at least that subset, so the figure can only fall. A number that rises later means a different budget, a different test set, or a misconfigured attack - not a more robust model.
  • What replaces a radius when the inputs are source diffs rather than images?
    An enumerated set of permitted edit operations plus a limit on how many an adversary may apply per input. Discrete inputs have no small epsilon, so a phrase like 'semantics-preserving edits' is a category, not a budget; without the operation set and the count, the percentage is measured over an unnamed space and cannot be reproduced or compared.
  • Why insist on clean accuracy in the same table?
    Because robust accuracy alone is gameable by a useless model: a detector that flags every diff is never fooled and scores perfectly, while blocking every merge. The pair tells you what the defence costs on ordinary traffic as well as what it withstood, and a defence is only interesting where both numbers are acceptable.

A lock advertised as resisting 60% of attempts tells you nothing until you know who was trying, with which tools, and for how long.

saying these in an interview costs you the question

  • Treats robust accuracy as a fixed property of the model
  • Reads an empirical figure as a guarantee
  • Quotes the number with no attack, access class or budget
  • Thinks random edits and adversarial edits cost the same
  • Ignores clean accuracy, so a flag-everything detector looks perfect

context