skip to content

In a red-team report on a malware classifier, how do you scope a chosen-verdict flip that reproduces once in five attempts?

level: seniorimportance: nice to knowfreq 27%

answer

  1. ask what the five actually are
  2. attempts on one file, or five files
  3. retries are free on a local scanner
  4. two findings, not one weak one
  5. rank by harm, not by percentage

basics

~20 s

Ask what the five are. One success in five attempts on one file, against a scanner the attacker runs offline, is a capability costing five attempts; one file in five is partial coverage. Report it separately.

solid answer

~50 s

The first move is to fix the denominator, because "one in five" names two different findings. If it means one success in five attempts on the same file, and the scanner runs locally where retries are free, the adversary has the capability at roughly five times the attempt cost, which is close to no cost at all. If it means one file in five ever reaches the named verdict, that is a partial capability and the interesting question is what distinguishes the four that did not. Second, do not merge it with the any-wrong-verdict result. Those are two findings with different success criteria, different denominators and different harms, and averaging them produces a number that describes neither. Usually the chosen-verdict finding is the more important of the two, because it is the one that maps to a bypass, even when its rate is the lower number.

go deeper

for a junior

Know that a reported success rate needs a denominator: successes out of attempts on one input is a different claim from successes out of inputs tried.

for a middle

Be able to explain why a one-in-five per-attempt result against a locally held model is a cost, not unreliability, and why the failed attempts bound the effort spent rather than the model's strength.

for a senior

Show the triage: split the chosen-verdict and any-wrong-verdict results into separate findings with their own denominators and budgets, and rank them by the harm each maps to rather than by rate.

for a principal

Own how findings of this class are written up across the team, so severity tracks the harm rather than the headline percentage and mitigations premised on metering are not accepted against an offline adversary.

## Two claims are sitting inside one report A red-team run against a locally installed malware classifier has produced two results. Modified variants of malicious files stop being read correctly almost every time. And on one chosen file, the specific verdict the adversary wanted, benign, came back on one attempt out of five. The instinct is to write this up as "evasion works reliably; the targeted variant is flaky". That is the wrong shape, and triage is where the shape gets fixed. ## First: what is the denominator? "One in five" is ambiguous in a way that changes the severity by an order of magnitude. **One success per five attempts on one file.** The adversary can repeat. Against a classifier they hold on their own machine, a re-scan costs a fraction of a second and nobody is counting, so a one-in-five per-attempt rate is not unreliability, it is a price: five attempts. Written honestly, that finding reads "the chosen verdict is reachable on this file at roughly five attempts". Calling that flaky imports an assumption from metered systems that does not hold here. **One success per five files.** Now the claim is about coverage, not cost, and the useful work is characterising the other four. Do they share a property? Were they already far from the boundary in the model's ordering? Is the named verdict simply unreachable for some inputs inside the budget? A per-file rate with the attempt budget stated is a real capability statement; without the budget it is nothing. The two readings can even coexist, in which case the finding needs both numbers: the fraction of files reached, at N attempts each. ## Second: it is a different finding, not a weaker one The any-wrong-verdict result and the chosen-verdict result have different success criteria, so they are not two measurements of one thing. Merging them, or presenting one as a degraded version of the other, destroys the only distinction that matters to whoever reads the report. Their harms differ too. On a benign-or-malicious scanner, a broad any-wrong-verdict result over a mixed sample set may be substantially made up of benign files pushed to malicious. That is real, and it costs the operator in false positives and analyst time, but it is not a bypass. The chosen-verdict result, malicious read as benign on a file the adversary picked, is the bypass. A lower number attached to a worse harm outranks a higher number attached to a milder one, and a report that ranks by percentage will get that backwards. So: two findings, two severities, two denominators, each with its own budget stated. If a single summary line is required, it names both goals rather than averaging them. ## Third: what the numbers do and do not establish Be careful about direction of claim in the write-up. - A high any-wrong-verdict rate establishes that the model is fragile on the sample population that was used. It does not establish that an adversary gains anything, and it says nothing about the population the product actually sees. - A one-in-five chosen-verdict rate establishes that the capability exists on the files tested at the budget spent. It does not establish that four fifths of files are safe, because the budget, not the model, may be what stopped the other four. - A failure to reach the named verdict at all bounds the effort that was spent, not the model. Nothing about "we could not do it in fifty attempts" implies "it cannot be done in five hundred". ## Fourth: anticipate the fix that is not a fix Someone will propose capping attempts per sample. Against a scanner distributed to endpoints and run locally, there is nothing to cap: the adversary owns the copy, the loop and the clock. Any control premised on metering is void in this threat model, and saying so plainly in the report is part of scoping it, because otherwise the finding is closed against a mitigation that was never in force. ## The habit to carry The triage question on this class of finding is always the same three words: *which goal, which denominator, which budget*. Answer those and a result that looked like one flaky observation resolves into two clean findings, correctly ordered by the harm each maps to rather than by the size of its percentage.

  • Someone proposes capping attempts per sample as the mitigation. What do you say in the report?
    That the control is premised on metering the adversary, and this adversary holds the classifier on their own machine. There is no request to rate limit, no counter to enforce, and no log the defender sees. The finding should say so explicitly, otherwise it gets closed against a mitigation that was never in force in this threat model.
  • Should the two results share one severity rating?
    No. Severity follows the harm each maps to, not the size of the percentage. On a benign-or-malicious scanner, the broad any-wrong-verdict result is dominated by errors the operator absorbs as false positives, while the chosen-verdict result is the bypass. The lower number frequently deserves the higher severity, and one shared rating hides that.
  • What extra evidence would you ask for before signing off the chosen-verdict finding?
    The attempt budget spent per file, the number of files attempted rather than only those that succeeded, the verdict the flips went to and from, and what the adversary was assumed to see while working. Without those four the result cannot be reproduced or compared to any later run.

saying these in an interview costs you the question

  • Calls a one-in-five result flaky without asking the denominator
  • Merges the chosen-verdict and any-wrong-verdict results into one rate
  • Ranks the two findings by percentage rather than by harm
  • Reads four failures as evidence the model resisted rather than the budget ran out
  • Proposes rate limiting an attacker who holds the classifier locally

context