An evasion run reports 66% success against a malware classifier — what is it not telling you?
answer
- a benign verdict is only half a result
- count the candidates that broke
- starts, or actually does its job?
- the edit budget belongs in the report
basics
~20 sIt is not telling you how many of those files still worked. A benign verdict counts as evasion only if the artefact still runs, so the report also needs the functional pass rate, the test bar and the edit budget.
solid answer
~50 sA benign verdict on an edited file is half a result; the other half is whether the file still does what it was built to do. So the first missing column is the functional pass rate — of the candidates that scored benign, how many survived being exercised — and the second is what the functional test actually was, because 'the process started' and 'it completed its job in a realistic environment' are very different bars. After that: which edit families were used and under what budget, including the ceiling on file-size and start-up-latency growth beyond which the artefact stops being usable; how many queries were spent and what the channel returned; and whether successes were rebuilt, re-extracted and re-scored rather than counted in the feature space. Read the direction too: a high rate shows that these edits beat this model under this budget, not that the model class is weak.
code
text · 10 linesrun: static file classifier, verdict-only channel
move family candidates scored benign still ran net
append to ignored region 500 412 ? ?
section padding 500 377 ? ?
no-op insertion 500 198 ? ?
...
headline reported: "66% evasion" <- 987/1500, third column only
not reported: functional test used, size / start-up growth cap,
queries spent per success, re-score after rebuildgo deeper
Know that an evasion figure on a file classifier is incomplete on its own: the edited file has to still work, and somebody has to have checked that.
Explain which columns make the number interpretable — functional pass rate, the test's bar, the edit families and their growth ceiling, the queries spent — and why each one changes the reading.
Show the reviewer's instinct: ask what fraction still ran, what counted as running, and what the edit budget was, and state the finding narrowly enough that nobody can quote it as a claim about the model class.
Own the reporting standard your team publishes under, and the trade it implies: stricter evidence per finding means fewer findings per quarter, and you have to be willing to defend that to people who count findings.
## Why one number is never the result here On a continuous input, an attack success rate is close to self-explanatory: within the stated radius, this fraction of inputs was misclassified. Two columns — the norm and the radius — make it comparable to somebody else's number. On an artefact that has to keep working, the same figure is ambiguous in a way that matters, because the adversary can fail in two completely different places. The model can refuse to be fooled, or the file can stop working. A run that counts only the first has measured something that is not the attack. ## The columns a credible report carries **The functional pass rate.** Of the candidates that scored benign, what fraction still performed their function? This is the column whose absence turns the headline into a number about byte sequences rather than about artefacts. It is also the column that tells a reader how hard the constraint really bit — a family of edits with a 90% break rate is a very different finding from one with a 5% break rate, even at identical evasion. **What the functional test was.** 'It loaded without error' is a floor. 'It performed its task end to end in an environment resembling the one it must run in' is the bar that makes the claim mean something. State the bar; the number is uninterpretable without it. **The move families and their budget.** Which edits were permitted, how many were applied, and what ceiling was placed on growth in file size and start-up latency. This is the analogue of a radius and it is what makes two runs comparable at all. A run allowed unbounded appended content is not comparable to one held to a few percent growth. **The access and the query count.** What the channel returned — a label, a score — and how many queries were spent per successful sample. A result that needed thousands of interactive queries and one that needed six describe very different adversaries even at the same success rate. **Whether the success was realised end to end.** Rebuilt as a file, exercised, re-extracted, re-scored. A count taken before the rebuild describes the feature space, and over-states the adversary. ## Reading the direction of the claim This is where candidates most often overreach. A 66% figure establishes that *these edit families*, under *this budget*, against *this model*, produced benign verdicts at that rate on *this sample set*. It does not establish that static file classifiers are weak; it does not establish that the artefacts would survive deployment; and it does not transfer to a retrained version of the same model without being re-run. The claim is narrow, and stating it narrowly is the mark of someone who has actually run these campaigns. The same discipline applies to the negative result. If the run scores 4%, that bounds the edits that were tried, not the model. The attacker who tries a different family, or the same family with a larger growth ceiling, is not covered by the finding at all. ## Reproducibility, which is its own column These campaigns produce flaky findings, and the flakiness has two sources with opposite meanings. Either the model's verdict is unstable across near-identical candidates, or the artefact only sometimes survives the edit. Those demand different follow-ups, and a report that just says 'reproduces intermittently' has not separated them. Give the run count, the success count and, where the variance sits — in the verdict or in the functional test. ## What this looks like as a habit When you are the reviewer, ask three questions in order and you will have covered most of it. *Of the ones that scored benign, how many still worked?* *What was the bar for 'still worked'?* *What was the edit budget?* Anything that cannot answer all three is a preliminary observation, not a finding — and preliminary observations should be labelled as such rather than shipped as a headline percentage that somebody else will quote out of context.
- What functional test would you accept for this kind of run?One that exercises the artefact's actual job in an environment resembling where it has to run, recorded pass or fail per candidate — not merely that a process started. Put the bar in the report, because the evasion number is only interpretable against it, and a weaker bar silently inflates every figure downstream.
- The attack reproduces once in five tries. How do you triage that?Report it as a rate with the run count rather than as a success, then separate the two possible causes: the model's verdict is unstable across near-identical candidates, or the artefact only sometimes survives the edit. Those have different consequences and different fixes, and a finding that has not separated them is not yet triaged.
- What does a 66% number license you to say about the model class?Only that these edit families, under this budget, produced benign verdicts at that rate against this model on this sample set. It does not generalise to file classifiers as a category, does not survive a retrained model without being re-run, and says nothing about edit families that were never tried.
saying these in an interview costs you the question
- Reports evasion without a functional pass rate
- Counts a broken artefact as a successful evasion
- Accepts that the process started as the functional test
- Generalises one run to the whole model class
- Omits the edit budget and the query count