skip to content

Targeted or Merely Wrong

Any misclassification is close to free and forcing one chosen class is expensive - at inference. Interviewers ask because headline rates quote the cheap number and the ordering is not universal.

on this pageshow

explore

questions

4

Against a malware classifier, what separates an untargeted evasion goal from a targeted one, and which costs more attempts?

level: juniorimportance: must knowfreq 70%

answer

  1. define winning before measuring anything
  2. any wrong answer, or one named one
  3. one acceptable outcome instead of many
  4. the direction of the flip is the payoff

basics

~20 s

Untargeted means any wrong verdict is success; targeted means one named verdict on one chosen file. Targeted is harder: the attacker must arrive somewhere specific rather than merely leave the right answer, so it costs more attempts.

solid answer

~50 s

The goal is what defines success, and until it is named no success rate means anything. An untargeted adversary wins on any verdict other than the correct one; a targeted adversary wins only when one chosen file is read as one named class. Targeted is the harder job: untargeted search can stop at the first outcome that is wrong in any way, while targeted search has to keep going past every wrong-but-useless outcome until it reaches the one class it wants, so at the same attempt budget it succeeds less often. On a two-class benign-or-malicious scanner the two goals coincide for a single file, but the aggregate still mixes both directions of flip, and only one direction, malicious read as benign, is a bypass. That is why an untargeted headline number rarely maps to a harm.

go deeper

for a junior

Be ready to state the two goals in one line each and say which is harder. Untargeted accepts any wrong verdict; targeted accepts one named verdict on one chosen input, and that narrower target costs more attempts.

for a middle

Expect to explain why the cost differs rather than assert it: untargeted search ends at the first wrong outcome, targeted search must pass through wrong outcomes that earn it nothing. Note that the targeted cost varies by which class is named.

for a senior

Show that you read a quoted success rate by asking which goal and which direction it measured, then say what it does and does not establish about a deployed system before anyone acts on it.

for a principal

Own the framing that success is undefined until the goal is named, and be able to say which of the two numbers belongs in a risk register for your product and why the other one is a curiosity.

## The goal is the definition of success An adversary who does not own a malware classifier but can feed it files must decide, before anything is measurable, what counts as winning. Two answers are possible, and the whole taxonomy of this area rests on the difference. **Untargeted (sometimes called an indiscriminate or "any wrong answer" goal).** Success is any verdict other than the correct one. The adversary does not care which way the model is wrong. On a multi-class model, every one of the k-1 incorrect classes ends the job. **Targeted.** Success is one named outcome on one chosen input: *this* file must be read as *that* class. Nothing else counts, including all the other ways the model could be wrong. Everything else follows from that one line. Robust accuracy, attack success rate, cost per sample and the business meaning of a finding are all different numbers under the two definitions, and a report that omits which one it measured has not reported an attack at all. ## Why the targeted goal costs more at inference The untargeted stopping rule is a disjunction: the search ends the moment the verdict is anything but correct. The targeted stopping rule is a single condition: it ends only when the verdict is the one named class. Intermediate wrong answers, which would have ended an untargeted run, earn a targeted run nothing and it keeps working. So the targeted adversary spends more of whatever budget they have. Against a locally installed scanner that returns a verdict for free, the budget is denominated in attempts per file, plus restarts after failures. Against a metered endpoint it is queries, and therefore money. In every setting, at the same budget, the untargeted success rate is the higher of the two, and the gap widens as the named class sits further from where the model currently puts the input. There is a second consequence: the targeted cost is not one number. Forcing a class that the model already considers a near neighbour is much cheaper than forcing a distant one, so a targeted rate quoted without saying which class it was measured on is close to meaningless. ## The two-class case, which is where malware sits A benign-or-malicious scanner has only two outcomes, so for one malicious file "any wrong verdict" and "the benign verdict" are the same event. It is tempting to conclude the distinction does not apply. It does, in a different place: in the aggregate. A success rate measured over a mixed set of files counts both directions of flip. A benign file pushed to *malicious* is an error the model made, and it inflates an untargeted number, but it buys an attacker no bypass at all; it is a nuisance the defender absorbs as a false positive. A malicious file pushed to *benign* is the bypass. So on two classes, the targeted question becomes "which direction, on which population of files", and that is the only version of the number that maps to the harm the product exists to prevent. ## Why this matters when you read a claim The failure mode is well worn: a result reports a very high evasion success rate, and the reader concludes the model is broken. Before that conclusion is available you need the goal. Untargeted success on a curated benchmark is close to free for the adversary and rarely corresponds to anything a business loses; the targeted rate, in the direction that pays, on files resembling what the system actually sees, is usually far lower and is the one worth putting in a risk register. The mirror error is just as bad. "It is only untargeted, so it does not matter" is wrong for systems where every error is a loss, for instance a detector whose whole value is not missing anything. Whether an untargeted number is meaningful is a property of the deployment, not of the attack. ## What to ask, every time - Which goal was measured, and if targeted, to which class and in which direction? - What is the denominator: attempts on one file, or the fraction of files that ever flipped? - What could the adversary see while working, a verdict only, a score, or the parameters? - Where did the files come from, and do they resemble live traffic? None of these is a detail. Each one changes the number by a factor that dwarfs the difference between models. ## The ordering is not universal One last thing worth carrying: this cost ordering belongs to inference-time attacks. An adversary who instead gets rows into the training corpus is in the opposite situation, where aiming at one specific behaviour is the cheaper goal and making the model broadly worse is the expensive one. The goal axis is the same axis; its price tag depends on when the adversary acts.

  • On a two-class benign-or-malicious scanner, do the two goals collapse into one attack?
    For a single file, yes: the only wrong verdict is the other one, so untargeted and targeted name the same event. The distinction moves to the aggregate. A success rate over a mixed file set counts benign files flipped to malicious, which is an error but not a bypass, alongside malicious files flipped to benign, which is. Only the second direction is a harm, so the number must be reported per direction.
  • Is an untargeted success rate ever the number that matters?
    Yes, when every error costs the operator the same. For a detector whose entire value is not missing events, any misclassification of a real event is a loss, so the untargeted figure maps directly to harm. It stops being meaningful when only one direction of error benefits an adversary, which is the usual case for a classifier that gates an action.
  • Why do published results usually quote the untargeted number?
    It is the cheapest to produce and the highest to print. It needs no choice of target class, it runs to completion faster because any wrong outcome ends the search, and it is comparable across datasets with different label sets. None of those are reasons to treat it as an estimate of what an adversary gains.

A pickpocket who will take any wallet can work a whole crowd. One who must take your wallet is doing a much harder job with the same skills.

saying these in an interview costs you the question

  • Says targeted just means attacking one input instead of many
  • Treats a high untargeted success rate as proof the model is broken
  • Assumes every misclassification is worth the same to an attacker
  • Claims both goals cost the same because both need one flip
  • Quotes a targeted rate without naming the class it aimed at

context

open as a page

With unlimited free re-scans of a local malware classifier, why does forcing one chosen verdict still cost more attempts?

level: middleimportance: should knowfreq 54%

basics

~10 s

The stopping rule is narrower. An any-wrong-verdict run ends at the first outcome that is not correct; a chosen-verdict run must pass those and keep going, burning more attempts and more restarts.

open as a page

A report claims 99% evasion success against your malware classifier. What do you require before funding a response?

level: principalimportance: should knowfreq 38%

basics

~10 s

Require the goal before the number: any wrong verdict or one chosen verdict, and in which direction. Fund against the malicious-read-as-benign rate on realistic files at a stated attempt budget.

open as a page

In a red-team report on a malware classifier, how do you scope a chosen-verdict flip that reproduces once in five attempts?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Ask what the five are. One success in five attempts on one file, against a scanner the attacker runs offline, is a capability costing five attempts; one file in five is partial coverage. Report it separately.

open as a page