skip to content

A report claims 99% evasion success against your malware classifier. What do you require before funding a response?

level: principalimportance: should knowfreq 38%

answer

  1. a percentage with no goal attached
  2. which direction did the flips go
  3. attempts per file, and how many
  4. only one direction is a bypass
  5. restate it or shelve it

basics

~10 s

Require the goal before the number: any wrong verdict or one chosen verdict, and in which direction. Fund against the malicious-read-as-benign rate on realistic files at a stated attempt budget.

solid answer

~50 s

A percentage with no goal attached is not yet a claim, so the first requirement is the goal: was success any wrong verdict, or one chosen verdict, and in which direction did the flips go? On a benign-or-malicious scanner an untargeted headline can be substantially benign files pushed to malicious, which costs us in false positives but gives an adversary no bypass. I also want the denominator and the attempt budget per file, the access the adversary was assumed to have, and where the samples came from. Then the funding decision follows the restated number: the malicious-read-as-benign rate on files resembling live traffic, at a budget an offline adversary would really spend, which is large. If the report cannot be restated, I treat it as unscoped rather than as either proof or noise, and say publicly what would change my mind.

code

text · 12 lines
text
Evasion evaluation (as printed in the report)
---------------------------------------------
model             : static binary classifier, benign/malicious
samples           : 500 files
attempts / sample : up to 50 local re-scores
access            : returned verdict only
success           : 99.2%
---------------------------------------------
not stated: the goal (any wrong verdict, or one chosen
verdict); the direction of the flips; whether 99.2% is
per attempt or per file; how the 500 files were selected
...

go deeper

for a junior

Know that an attack success rate is not interpretable until someone says what counted as success, and that asking which goal it measured is a reasonable first question.

for a middle

Be able to list the columns a claim needs before it means anything: goal and direction, denominator, attempt budget, access assumed, and where the samples came from.

for a senior

Show that you can restate someone else's headline into a scoped sentence about an adversary and a budget, and say what that restated result does and does not establish about the deployed system.

for a principal

Own the funding call and the standing rule behind it: which restated number enters the risk register, what you tell a customer who read the same report, and how you say no to an unscoped claim without dismissing it.

## The decision, not the number Someone hands the team a result: 99% evasion success against the classifier the product ships. Engineering wants to rewrite the model, sales wants a rebuttal, and a customer has already read it. The judgment being asked for is what, if anything, gets funded, and on what evidence. That decision cannot be made from a percentage alone, and the reason is the goal. ## What the number is missing Four columns turn a percentage into a claim. **The goal, and its direction.** Untargeted success counts any verdict other than the correct one. Targeted success counts one named verdict. On a two-class scanner these coincide per file but diverge in the aggregate, because an untargeted rate measured over a mixed sample set includes benign files pushed to malicious. That direction is a real defect, it costs analyst hours and customer trust, and it buys an attacker nothing. Only malicious read as benign is a bypass. A single headline that mixes the two is not the number a security decision should hang on. **The denominator and the budget.** Successes out of attempts on one file is a cost statement. Successes out of files tried is a coverage statement. Neither means anything without the attempts allowed per file, and against a classifier the adversary runs locally that budget is effectively unbounded, so a low per-attempt rate does not protect anything. **The access assumed.** Verdict only, scores returned, or parameters in hand are three different adversaries. Withholding scores raises the price of a chosen-verdict search; it does not remove an attack that works from labels alone. **The sample population.** A curated set of known-malicious files is not what the deployed system sees. A result on files that do not resemble live traffic bounds a laboratory, not a product. ## Restating it, and then deciding The restated form is a sentence, not a percentage: *against an adversary who sees only a verdict and re-scans locally, this fraction of files resembling live traffic reached the benign verdict within this many attempts.* Once the claim is in that shape, three organisational outcomes are available and each is defensible. - **The number restates high in the direction that matters.** Fund it. That is a bypass with a demonstrated cost, and the roadmap moves. - **The number restates low, or is dominated by the harmless direction.** Do not fund a model rewrite, and say why in writing. Then treat the false-positive half seriously on its own merits, because that half is a product-quality problem even though it is not a security one. - **The report cannot be restated because the author will not supply the columns.** Treat it as unscoped. Not disproved, not accepted. State exactly which columns would change the decision, and put the request in writing, so the position is a standard rather than a dismissal. The trap on both sides is symmetric. Reading "99%" as proof the model is broken funds work against a number that may correspond to no adversary at all. Reading "it is only untargeted" as an all-clear is equally wrong: for some deployments any error is a loss, and an untargeted result is then exactly on point. Which applies is a fact about what your product promises, and that is a call the person who owns the product has to make rather than delegate to the researcher. ## What you tell people The customer wants to know whether they are exposed. The honest answer names the adversary and the budget rather than quoting a rival percentage: what the demonstrated capability is, in which direction, at what cost, on which population, and what is not measured. A counter-number is a worse answer than a scoped sentence, because it invites the same mistake in the other direction. The executive question, "are we broken", has one honest shape too: the model behaves as designed on ordinary inputs, and against an adversary who modifies inputs deliberately it can be pushed to a chosen verdict at a stated cost, which is a property of every deployed classifier and not a defect unique to this one. What varies between products is the cost, and that is the thing worth improving and worth measuring over releases. ## The standing rule Set it once so the team stops relitigating it: no evasion number enters the risk register without its goal, its direction, its denominator, its attempt budget, its access assumption and its sample population. Numbers arriving without those get restated or shelved. That rule costs one meeting to establish and saves every subsequent quarter from funding against whichever percentage was largest.

  • The author restates it as 41% targeted bypass. Does the lower number lower the priority?
    No, it usually raises it. The 99% was a mixture that included flips in the direction nobody pays for, while 41% of malicious files reaching a benign verdict is the harm the product exists to prevent, demonstrated at a stated cost. Severity follows what the number maps to, not its size, and a smaller well-scoped number outranks a larger unscoped one.
  • What do you tell an executive who asks whether the model is broken?
    That against ordinary inputs it behaves as designed, and against an adversary who modifies inputs deliberately it can be pushed to a chosen verdict at a measurable cost, which is true of every deployed classifier. What differs between products is that cost. Then give the scoped sentence: which direction, which population, how many attempts, and what has not been measured.
  • The false-positive direction turns out to dominate the 99%. Is there nothing to do?
    There is, but it is not a security spend. Benign files pushed to malicious cost analyst hours and customer trust, so it belongs on the quality roadmap with its own owner and its own measurement. Keeping it separate from the bypass work is what stops one headline percentage from funding the wrong team.

saying these in an interview costs you the question

  • Accepts a headline percentage without asking which goal it measured
  • Treats an untargeted result as proof the deployed model is bypassed
  • Dismisses a finding as untargeted when every error is a loss anyway
  • Ranks findings by percentage rather than by the harm each maps to
  • Assumes an unstated attempt budget is small because the rate looks low

context