A red-team report says every out-of-spec unit was accepted - what must its cost line state to be usable?
answer
- success rate alone is not a finding
- say what access was granted
- steps and restarts per unit
- same allowance on every row
- the random baseline should be near zero
basics
~10 sSuccess alone is not a finding. State the access granted, the optimisation steps and restarts per unit, the time that implies, the fixed allowance, and a same-size random-change baseline that should be near zero.
solid answer
~40 sA 20-of-20 line says nothing about who can reproduce it or what it costs them. The cost line needs four things. **Access**: the attacker was handed the weights and a machine that could differentiate the model, which bounds this to whoever can reach a copy of the artefact. **Effort**: steps and restarts per unit and the GPU-minutes they imply, including where added steps stopped buying success. **The allowance**, held fixed across every row so the rows compare. **The baseline**: an equally large random change on the same units, which should accept close to none - without it nobody can tell whether aiming mattered. The report should also say what it does not establish: it is silent about an adversary who sees only the pass or hold decision.
code
text · 11 linesengagement: in-line defect classifier, weights read on the supplier workstation
permitted per-unit change: identical on every row below
method steps restarts units accepted / 20 GPU-min / unit
random change - - 0 / 20 0.00
single step 1 1 6 / 20 0.01
iterated 40 1 17 / 20 0.35
iterated 400 1 18 / 20 3.50
iterated 40 5 20 / 20 1.70
...
not simulated: an adversary who sees only the pass / hold decisiongo deeper
Know that a reported attack success rate is meaningless without the access the attacker was given and the effort it took per unit.
Be able to name the columns: access, steps, restarts, time per unit, the fixed allowance, and a same-size random-change baseline.
Show the judgment: read the saturation point and the restart count, refuse rows at different allowances, and state plainly what the engagement did not simulate.
Own the conversion from a technical result to a decision - the prerequisite is a differentiable copy, so the lever this finding exposes is artefact distribution, not model retraining.
## Why success rate alone is not a finding A red-teamer finishing a factory engagement can honestly write that a chosen out-of-spec unit was accepted by the in-line inspection model on every attempt. The person who has to act on that sentence - a plant owner deciding whether to change anything, and what - cannot do anything with it, because the sentence omits everything that determines whether the attack is a realistic threat or a laboratory result. ## The four things the cost line must carry **1. The access that was granted.** The attack requires differentiating the model end to end: the weights, the architecture, and a machine that will report derivatives. That is the strongest assumption available in this family, and it defines the population who can reproduce the result - roughly, anyone who can obtain a copy of the model artefact. Reported without it, a 20-of-20 line reads as though a passer-by could do this. Reported with it, the owner's next question becomes the right one: who has a copy, and how many of them are there. **2. Effort per unit.** The natural unit here is optimisation **steps** and random **restarts**, because each step costs about one backward pass. Report both, plus the GPU-minutes and the attacker-hours they translate to. Two shapes are worth calling out explicitly because they change the interpretation: where the success curve **saturated** in steps (beyond which more steps bought nothing), and how many **restarts** the stubborn units needed. "Every unit, but three of them needed five restarts" is a materially different finding from "every unit on the first trajectory". **3. The allowance, held fixed.** Every row of a comparison must be at the same permitted change, or the rows are not comparable and the table means nothing. The allowance itself is a claim about the adversary and belongs stated in the threat model, not silently varied per row. **4. The baseline that isolates aiming.** An equally large random change, on the same units, reported in its own row. It should accept essentially none. If it accepts a meaningful fraction, the finding has changed character entirely - the model is fragile to ordinary variation, which is a quality problem the QA team owns, not an adversary problem. ## What the report must decline to claim - It does not establish anything about an adversary who cannot differentiate the model. That adversary - one who sees only the pass or hold decision - was not simulated, and their cost is a different and much larger number. - It does not establish that the change would survive being physically realised. A change written into a digital image is not the same object as an alteration to a unit in front of a fixed-mount camera, where resampling, lighting, focus and compression are all applied afterwards. - It does not establish that the model is poorly trained. Ordinary accuracy is measured on inputs nobody chose adversarially, and it can be excellent alongside this result. ## Triage when the result is unstable A finding that reproduces on two attempts in five is not automatically noise. With this attack the usual cause is the starting point: the search is not convex, so different randomly displaced starts inside the allowance reach different outcomes. Before writing "intermittent", check whether restarts fix it. If five restarts turn 2-of-5 into 5-of-5, the honest report is "reproducible at a cost of five restarts", which is a stronger and more actionable claim than "sometimes". ## What a reader should be able to compute from your table A good cost line lets someone else convert your result into their own decision without rerunning anything: the money and the hours per accepted unit, the profile of who could sit in that chair, and the sensitivity of the whole thing to the one control the plant actually holds - which is not the model's weights but who gets a copy of them. That is the difference between a finding that changes something and a screenshot of a flipped label.
- The attack succeeds on two attempts in five - do you file it as intermittent?Not before testing restarts. The search is non-convex, so the starting point inside the allowance often decides the outcome. If five restarts make it reliable, report it as reproducible at a cost of five restarts per unit. Calling it intermittent when it is merely restart-hungry understates the finding and invites the wrong response.
- The owner asks what to fix. What does this result actually point at?Primarily at who holds a copy of the model, because the attack's prerequisite is a differentiable copy. Hardening the function itself is a separate and expensive programme with its own accuracy cost. The cheapest lever the result exposes is the distribution of the artefact and the physical control of the workstations that hold it.
- How would the cost line change if the attacker only saw the pass or hold decision?Steps and restarts stop being the unit of cost and queries become it, because the direction has to be inferred from returned decisions rather than computed. The per-unit price rises by orders of magnitude and becomes rate-limitable. That is a different engagement and needs its own rows; it cannot be extrapolated from a white-box result.
saying these in an interview costs you the question
- Reports success rate with no access assumption stated
- Omits the random-change baseline entirely
- Varies the permitted change between compared rows
- Files a restart-hungry result as intermittent noise
- Extrapolates a white-box result to a decision-only adversary