skip to content

Two merge-gate detectors report 55% and 48% robust accuracy against edited diffs - is the first better?

level: middleimportance: should knowfreq 54%

answer

  1. two numbers, two different adversaries
  2. each one is a ceiling, not a measurement
  3. the weaker evaluation flatters its model
  4. steps and restarts move it by tens of points
  5. same attack, same access, same budget, same set

basics

~20 s

Not from those numbers. Robust accuracy ranks two models only when both faced the same edit budget, the same access class and at least as strong an attack. Optimisation steps alone can move such a figure by tens of points.

solid answer

~40 s

Two robust-accuracy figures are comparable only when four things match: the attack and what it could see of the model, how hard it searched in steps and restarts, the budget an adversary was held to, and the test set. Each figure is an upper bound produced by its own attack, so **the weaker evaluation flatters its model**. If the 48% model was attacked with far more optimisation steps under full access while the 55% model faced a decision-only attack at ten steps, the 48% model is plausibly the more robust of the two. The honest move is to re-run one attack you specify, at one budget you state, against both - and if you hold neither checkpoint nor endpoint, to record the table as unranked rather than pick the larger number.

code

text · 7 lines
text
Table 3 - merge-gating detectors, "robust to semantics-preserving edits"

detector  attack access   steps  restarts  edits/diff   clean acc  robust acc
A         decision-only      10         1  not stated       96.1%       55.0%
B         full access       200         5  not stated       94.8%       48.0%
...
(no row states which edit operations were permitted, or how many per diff)

go deeper

for a junior

Know that a robust-accuracy percentage belongs to one attack under one budget, so two such percentages from two different write-ups are not a ranking.

for a middle

Be able to name the columns that must match - attack and access class, steps and restarts, the permitted edits and their count, the test set - and explain why a weaker evaluation reports the higher figure.

for a senior

Show how you would settle it in practice: one attack you specify against both models, steps pushed until the figure stops falling, and an explicit 'unranked' verdict when you cannot re-run the evaluation yourself.

for a principal

Decide what your organisation will accept as a comparison before a number is allowed to choose a vendor, and who funds the independent evaluation that makes two claims comparable at all.

## The mistake this question exists to catch The reflex answer is arithmetic: 55 is bigger than 48, so model A is better. That reflex is wrong for a reason that has nothing to do with statistics and everything to do with what the number is. **Robust accuracy is the score of a contest between a model and one specific attack under one specific budget.** Two numbers produced by two different contests do not rank the contestants any more than two runners' times rank them when one ran uphill. ## Why the weaker evaluation wins the table An attack searches the space of permitted modifications for one the model reads wrongly. A better search finds at least as many as a worse one, so each reported figure is an **upper bound** on the true robustness of that model at that budget. Upper bounds do not compare: the model whose evaluator tried less hard gets the flattering ceiling. In practice the difference is not marginal. Two parameters do most of the moving: - **Optimisation steps.** A search stopped early has explored a fraction of what it would have. Extending steps until the curve flattens routinely takes a reported figure down by tens of points on the same weights. - **Access class.** An attack that can see weights and gradients searches far more efficiently than one working from the returned decision alone. Handing the attacker less access does not make the model stronger; it makes the measurement weaker, and it silently converts a robustness claim into an obscurity claim. **Restarts** matter for the same reason as steps: a single starting point can stall in a poor region of the search space, and sampling several starting points finds failures one run misses. ## What must match before a comparison means anything | Column | Why a mismatch breaks the ranking | | --- | --- | | Attack and access class | A decision-only attack and a full-access attack are different adversaries; their numbers are on different scales | | Optimisation steps and restarts | Sets how much of the truth the number hides; the largest single lever on the figure | | The permitted edits and their count | The budget *is* the adversary; without it, neither figure describes anyone | | Test set and clean accuracy | A different set, or a detector that simply flags more, changes both numbers for reasons unrelated to robustness | Notice that on a source diff there is no radius to quote. The budget has to be stated as which edit operations an adversary could draw from and how many per input. A table where both rows say only "semantics-preserving edits" has stated no budget twice, and the two rows may be describing adversaries that differ by an order of magnitude in what they were allowed to change. ## What a defensible comparison looks like One attack, chosen by whoever is doing the comparing, run against both models at one stated budget and one access class, at a step count pushed until the figure stops falling, with restarts stated, reported beside clean accuracy on the same test set. Anything short of that is a claim about two evaluations rather than about two models. If you cannot do that - because you are reviewing a write-up and hold neither checkpoint nor endpoint - the correct output is not a ranking. It is a list of the columns that must be reported before the table supports one. As a reviewer that is a stronger position than it sounds: the authors chose which attack to run and how long to run it, and "we cannot tell which model is better from this table" is a finding, not an admission. ## The trap on the other side Do not over-correct into treating clean accuracy as the tiebreaker either. Defences that buy robustness generally give up some clean accuracy, so the model with the lower clean figure may be the one that actually did something. The two axes are read together, and neither one alone orders a pair of systems.

  • If you could demand only one extra column, which would it be?
    The budget: which edit operations were permitted and how many per input. Without it neither figure describes any adversary, so nothing else in the row can be interpreted. Steps and restarts come next, because they set how far each reported ceiling sits above the truth.
  • How would you actually rank the two models?
    Run one attack you choose against both, at one access class and one stated edit budget, increasing optimisation steps until each figure stops falling, and report restarts and clean accuracy on the same test set. If you hold neither checkpoint nor endpoint, record the pair as unranked and say which columns would settle it.
  • Does higher clean accuracy break the tie?
    No. Clean accuracy is a separate axis, and a defence that genuinely buys robustness usually costs some of it, so the lower clean figure can belong to the stronger defence. Read the two columns together and rank only when both were measured under one shared evaluation.

saying these in an interview costs you the question

  • Ranks two robust-accuracy figures with no shared budget
  • Assumes the higher number faced the harder attack
  • Calls steps and restarts an implementation detail
  • Treats 'semantics-preserving edits' as a stated budget
  • Uses clean accuracy as the tiebreaker between the two

context