skip to content

An extraction run reports 97% agreement with a paid malware-verdict API - what do you ask before believing it?

level: seniorimportance: nice to knowfreq 30%

answer

  1. ninety-seven percent of what
  2. compare it against always answering benign
  3. which pool, and chosen when
  4. the class skew hides the interesting side

basics

~10 s

Ask which pool agreement was measured on and the target's base rate there. On a pool where the service answers benign 97% of the time, a copy that always says benign also scores 97%.

solid answer

~50 s

First: which pool was agreement measured on, and what is the target's class base rate on that pool? If the evaluation inputs were drawn the same way as the queries - random byte-strings, say - then a copy that always answers benign already scores in the high nineties, and 97% is the base rate rather than a result. Second: was the evaluation pool fixed in advance, drawn from the traffic the copy is meant to stand in for, and identical across the allocations being compared? Otherwise the three rows are not comparable at all. Third: per-class agreement, because a single aggregate on a heavily skewed two-class problem hides everything on the malicious side, which is the side anyone cares about. Only after those does the ranking of allocations mean anything - and it frequently inverts once they are answered.

code

text · 10 lines
text
extraction run - agreement after 50,000 submissions ($0.02 each)

allocation                      submissions   agreement
random byte-strings                 50,000       97.1%
public in-domain executables        50,000       88.4%
boundary-seeking selection          50,000       91.2%

... evaluation pool not stated (same as query pool?)
... target's benign rate on that pool not stated
... per-class agreement not broken out

go deeper

for a junior

Know that an agreement percentage means nothing without knowing which inputs it was measured on and what a trivial copy that always gives one answer would have scored.

for a middle

Explain why the evaluation pool must be fixed in advance and shared across allocations, rather than being whatever inputs each allocation happened to submit.

for a senior

Interrogate a reported extraction result the way you would any evaluation - base rate, pool provenance, per-class breakdown, spend - before you accept the ranking of allocations it implies.

for a principal

Own what you will and will not put your name to: an agreement figure without its pool and base rate is not a finding, and the fact that the cheapest allocation won is itself the report.

## Why a bare agreement figure is not a finding Agreement - the fraction of inputs on which a fitted copy returns the same verdict as the target it was extracted from - is the right metric for extraction. It is also almost useless on its own, for exactly the reasons any evaluation number is useless without its columns. Three questions turn it into a claim. ## 1. Measured on what, and against what base rate? Agreement is a property of a *pool*. Say what the pool was and how it was drawn. Then say what the target itself answers on that pool, because that sets the floor: if the service returns benign for 97% of those inputs, a copy that always answers benign scores 97% agreement and has learned nothing. The failure mode is specific and common. An allocation that spends its whole budget on uniformly random or synthetic junk gets nearly-constant verdicts back; the copy learns the constant; and if the evaluation pool is drawn the same way, it reports the highest agreement of any allocation on the sheet. The tell is that the allocation with the *least* information per paid label produced the *best* number. That pattern is nearly always a base-rate artefact, not a discovery. ## 2. Was the evaluation pool fixed in advance and shared? Each allocation must be scored on the same pool, chosen before the campaign, drawn from the traffic the copy is intended to replace. If each row was scored on whatever it happened to submit, the rows measure different functions on different distributions and cannot be ordered against each other. This is where extraction write-ups fall apart most often, because the query pool is sitting right there and reusing it is free. ## 3. Aggregate or per class? A verdict service over real traffic is heavily skewed. A single aggregate agreement number is dominated by the majority class, and the entire question of whether the copy reproduces the target's *malicious* decisions can hide inside three percentage points. Break agreement out per class, or the number tells you about the easy side only. ## The cost columns, once the number survives Only after the above is it worth asking what the result cost: submissions, unit price, total spend, and how many submissions went to locating good inputs rather than to labelling them. That last one matters because with a top-1 verdict and no score, an attacker inferring which files sit near a boundary spends budget on the inference itself. A campaign that reports 50,000 submissions may have spent a meaningful slice of them on search, and comparing it to an allocation that spent all 50,000 on labelling a free in-domain corpus is comparing different things. ## What the sheet should look like A defensible reporting line names, for each allocation: the submission count, the unit price, the evaluation pool and how it was drawn, the target's base rate on that pool, aggregate agreement and per-class agreement. Anything less and the ranking of allocations is not established, whatever the headline says. ## Why this question is the one a red-teamer gets asked The chair here is the person writing the engagement report, or reading a teammate's. The deliverable is a claim about what a fixed budget bought - and the client's follow-up question is always some version of "so how good is the copy, really?". A number that collapses under "measured on what?" is worse than no number, because it will be quoted back. The honest report usually says something less dramatic and far more useful: on a fixed in-domain evaluation pool, with the target's base rate stated, this allocation reached this agreement overall and this much on the class that matters, for this spend. ## What a well-founded agreement number still does not establish Even a clean figure is a statement about the evaluation pool and nothing beyond it. It does not say the copy behaves like the target on inputs unlike anything in that pool; it does not say anything about the target's parameters; and it does not say the copy would remain in agreement after the service is retrained. Agreement is a snapshot of behavioural overlap on one distribution, priced in submissions - which is exactly what an extraction engagement is entitled to claim, and no more.

  • In that table, why is the random-input row the suspicious one rather than the impressive one?
    Because it is the allocation with the least information per paid label yet the best number. That combination almost always means the evaluation pool was drawn the same way as the queries, where the target's verdict is nearly constant and a copy that always answers one way already scores in the high nineties. The row is reporting a base rate wearing the costume of a result.
  • What would make those three rows comparable?
    One evaluation pool, fixed before the campaign, drawn from the traffic the copy is meant to stand in for and used identically for all three; the target's own class base rate on that pool printed beside each figure; per-class agreement broken out; and the submission count split between labelling and the search for which inputs to label.
  • Once the number is clean, what does it still not establish?
    That the copy behaves like the target anywhere outside that pool, that anything was learned about its parameters, or that agreement survives the service being retrained. Agreement is behavioural overlap on one distribution at one point in time, priced in submissions. That is a legitimate thing for an extraction engagement to claim, and it is the whole of what it can claim.

saying these in an interview costs you the question

  • Reads a high agreement number as a good copy
  • Compares agreement across differently drawn evaluation pools
  • Ignores the base rate a constant answer would score
  • Reports one aggregate number on a heavily skewed problem
  • Never asks what the campaign actually cost per point of agreement

context