A review deck shows a hardened ranker at 61% accuracy under attack and 1.8 clean points lower — what do you ask for?
answer
- two numbers, two different populations
- one is conditional, one is not
- a number about the attack, not the model
- ask what the eval attack was
- and what fraction of traffic qualifies
basics
~20 sAsk for what makes the two numbers comparable: the edit set and its size, the strength of the attack run and whether it was the training attack, per-slice clean scores, and the share of live impressions manipulated inside that edit set.
solid answer
~40 sNeither number means much alone, and together they still do not compare. `61% under attack` says the specific attack that was run, at whatever strength, failed that often against this checkpoint — it is not a property of the model. So ask which perturbation set was used, how strong the search was, and whether the evaluation attack is the one the hardening trained against, because a defence graded by its own training attack reports an upper bound on itself. The `1.8 clean points` is charged on all traffic, so ask for it per slice: aggregate parity routinely hides a category carrying the whole loss. Then ask the only figure that puts both into one unit — what fraction of live impressions is manipulated *and* inside that edit set.
code
text · 10 linescandidate clean NDCG@5 acc. under attack attack family eval set
-----------------------------------------------------------------------------
baseline 0.412 0.19 seller-edit 10k held-out
hardened 0.394 0.61 seller-edit 10k held-out
...
not reported: edit set size (fields, per-field range, cost to seller)
not reported: attack strength (steps, restarts, adaptive to defence?)
not reported: is the eval attack family the one used in hardening?
not reported: clean NDCG@5 per category / per seller segment
not reported: share of live impressions manipulated inside that edit setgo deeper
Know that a figure quoted under attack describes the attack that was run, not the model, and that a clean-accuracy loss applies to all traffic rather than only to attacked traffic.
Be able to name the missing columns: the edit set and its size, the attack strength, whether the evaluation attack is the training attack, and the per-slice clean numbers.
Show the reviewer's move — refuse to compare the two figures until both are expressed per impression, and identify the manipulated-traffic share as the missing multiplier that makes the comparison possible.
Set the reporting standard rather than fixing one deck: define what a robustness claim must carry before it reaches a go/no-go, so this argument does not get relitigated every quarter.
## What each number actually asserts The review deck has two figures and they are doing very different jobs. **`61% accuracy under attack`** is a statement about an *attack*, not about a model. It says: the attack that was run, at the strength it was run, against this checkpoint, on this evaluation set, failed 61% of the time. Change any of those four things and the number changes. Read the other way — the direction that matters — a high figure proves the attack that was run failed, never that the model is robust. **`1.8 clean points lower`** is a statement about all served traffic. It is unconditional: every impression the hardened checkpoint serves is served slightly worse. One is conditional on an adversary arriving with a specific capability; the other is charged unconditionally. They cannot be subtracted. ## The columns that have to be on the table For the attacked number: - **The perturbation set and its size.** On a listing surface this is not an epsilon; it is a set of permitted edits — how many fields, how far each may move, what it costs the seller. A robustness figure quoted without its edit set is quoting nothing, and a tiny set produces both a flattering number and a nearly worthless defence. - **The strength of the attack.** Optimisation steps, restarts, and whether the search was allowed to adapt to this specific defence. A weak search produces a high number for free. - **Whether the evaluation attack is the training attack.** If the hardening was trained against exactly the moves it is graded on, the figure is close to a self-report: an upper bound on the defence, not an estimate of it. A held-out family of edits is what makes the number informative. - **Whether an unhardened baseline was run under the identical attack.** Without it there is no way to attribute the 61% to the hardening rather than to the attack being weak. For the clean number: - **Per-slice results.** The loss from hardening concentrates on marginal, rare and thin-data cases. Aggregate parity is compatible with one category or one seller segment absorbing several times the average. A per-slice table is the single most informative artefact in this review. - **The same evaluation set, unshifted.** Two checkpoints scored on different traffic windows are not comparable, and on a marketplace the window moves fast. ## The conversion nobody brings Even with every column filled, the comparison is still not decidable, because the deck is missing the multiplier: **what share of live impressions is manipulated *and* inside the stated edit set.** That fraction is what turns a conditional benefit into a per-impression expected value, so it can be set beside an unconditional per-impression cost. Teams almost never measure it, and the substitute — 'we know sellers game the ranking' — is not the same statement, because manipulation outside the trained-for edit set collects almost none of the benefit. A second conversion is needed too: the two harms are not the same *kind*. One undeserved top slot may cost the marketplace far more than one slightly mis-ordered ordinary page, or far less. Straight accuracy-point arithmetic silently assumes they are equal. ## Other traps in a deck like this - **Aggregate accuracy holding flat is not evidence nothing happened** — it is equally consistent with a loss that landed entirely on one slice. - **A defence whose attacked number is high while a cheaper, weaker attack scores better against it** is showing a search failure rather than robustness; that pattern is a reason to rerun the evaluation, not to ship. - **Robustness bought at one edit set does not transfer to another,** and none of it transfers to manipulation that happens off the ranking surface entirely. ## What you say in the room 'Right now I can conclude that one specific attack, at one strength, on one evaluation set, succeeded less often against the hardened checkpoint, and that the hardened checkpoint serves all traffic 1.8 points worse. I cannot yet conclude that shipping it is net positive, because nothing here tells me how much of our traffic the defended condition actually covers, or which slice is paying the 1.8. Give me the edit set, the attack strength, a held-out attack family, the per-slice clean table and the manipulated-traffic share, and this becomes a decision.'
- Why does it matter whether the evaluation attack is the same family the hardening was trained on?Because a defence graded by its own training attack is close to self-reporting: the model has seen exactly those moves and the figure becomes an upper bound on the defence rather than an estimate of it. A held-out family of edits, and a stronger search than the training one, is what turns the number into evidence. Without that, a high figure is consistent with a defence that only covers what it memorised.
- The deck shows aggregate clean accuracy within 0.2 points. Is the trade-off gone?Not established. Aggregate parity is fully consistent with the loss landing on one slice — a thin category, a rare query class, a small-seller segment — while the bulk of traffic is unaffected. Ask for the per-slice table before accepting parity, because that slice is where the bill actually went and it may be a segment the business cares about disproportionately.
- What single missing figure most changes the decision?The share of live impressions that are manipulated and inside the stated edit set. It is the multiplier that converts a conditional benefit into a per-impression expected value so it can be set against an unconditional per-impression cost. Everything else refines the estimate; this one decides whether there is anything to weigh at all.
saying these in an interview costs you the question
- Reads 61% as a property of the model
- Subtracts robust and clean points as one unit
- Accepts a robustness figure with no edit set
- Ignores whether the eval attack was the training attack
- Takes aggregate clean parity as proof of no loss