skip to content

Why does spending a malware-API query budget on random byte-strings buy far less agreement than real executables?

level: middleimportance: should knowfreq 48%

answer

  1. the classifier never saw inputs like these
  2. the reply is predictable before you pay
  3. a predictable label carries no information
  4. high agreement on a pool nobody cares about

basics

~20 s

Random byte-strings sit far off the distribution the classifier was fitted on, so nearly all of them come back with the same verdict. A label you could have predicted before paying for it teaches the copy nothing.

solid answer

~50 s

Because the service was fitted on real executables, and random byte-strings sit nowhere near that distribution - it returns the same verdict, almost always benign, for nearly all of them. A label the attacker could have predicted before paying for it carries no information, so those submissions constrain the copy nowhere that matters. Worse, the copy dutifully learns the constant rule and then scores very high agreement when tested on that same random pool, which reads as success and is not. A free public corpus of real in-domain executables spans the regions the target actually separates, so each paid verdict is one the copy could not already guess. Synthetic inputs are not automatically wasted: ones steered toward where the partly-fitted copy is unsettled behave quite differently from uniform junk, but they still need many more submissions than a free in-domain corpus for the same agreement.

go deeper

for a junior

Recall that a classifier's answers on inputs unlike anything it was built for are close to meaningless, and that paying for such answers burns the budget without moving the copy.

for a middle

Explain information per label: a verdict the copy could already predict changes nothing in it, which is exactly why off-distribution submissions are near-worthless.

for a senior

Spot the failure signature - very high agreement measured on the same odd pool the queries came from - and be able to say what it actually proves and what it does not.

for a principal

Be able to say where the cheap allocation stops: with no free in-domain corpus, steered synthetic inputs become the fallback at a much higher submission count, which changes what a campaign is worth quoting.

## What the paid verdict is for An attacker copying a hosted verdict service - submit an executable, get back malicious or benign, pay per submission - is buying labels for files they already hold. The copy is then fitted to reproduce those verdicts. So the value of any one submission is exactly the amount it changes the copy, and that is set by how surprising the returned verdict is. ## Why random inputs are close to worthless The target was fitted on real executables: things with headers, sections, imports, code that runs. A uniformly random byte-string has none of that structure and sits far outside the region the classifier ever had to separate. Two consequences follow, and both cut the same way. **The answer is predictable in advance.** Off-distribution inputs collapse onto whichever side of the boundary that whole region falls on - in practice, nearly always the same verdict for nearly all of them. After a few hundred submissions the attacker can guess the rest for free. Paying for the remaining 49,000 buys confirmations, not information. **The constraint lands where nobody cares.** A fitted copy is only pinned down near the inputs it was trained on. Verdicts on junk pin the copy down in a region of file space that no real submission ever visits, and leave the region that matters - the neighbourhood where real clean software and real malicious software are separated - completely unconstrained. ## The trap this creates in the write-up There is a failure signature that turns up constantly in this kind of work. The copy is fitted on the random pool, evaluated on inputs drawn the same way, and reports very high agreement. That number is real and means nothing: on a pool where the target answers benign 97% of the time, a copy that always answers benign scores 97%. The measurement has to be made on a pool drawn from the traffic the copy is meant to stand in for, with the target's base rate on that pool reported beside it, or the whole comparison is a base-rate artefact. ## Why a free in-domain corpus is the strong baseline Public executables - project releases, packaged installers, binaries shipped with an operating system, and whatever already flows through an attacker's own gateway - cost nothing and are exactly the kind of thing the target was built to classify. They spread across the regions where the target's verdict genuinely varies, so the labels bought on them are labels the copy could not have guessed. This baseline is why "they would need a corpus like ours" is not the defence it sounds like: the inputs are free, and the target itself supplies the labelling. ## Where synthetic inputs do earn their place It would be wrong to conclude that anything not scraped from the real world is useless. Inputs generated to sit where the partly-fitted copy is unsettled - where two of the attacker's own fitted copies, or successive versions of one, return different verdicts - behave nothing like uniform noise. They concentrate exactly where the answer is unpredictable, which is where a label is worth the most. What they cost is extra submissions to locate: with a top-1 verdict and no score, the attacker cannot read proximity to a boundary directly and has to infer it from disagreement, which itself burns budget. So the ordering by submissions-per-unit-of-agreement typically runs: free in-domain corpus first, steered synthetic inputs when no such corpus exists, uniform random essentially never. ## The information framing, in one line A submission is worth what its answer could not have been guessed for. That single sentence explains the ordering above, explains why more queries do not automatically mean a better copy, and explains why a defender who reasons about *volume* alone - counting submissions - misses that two campaigns of identical volume can differ by an order of magnitude in what they extracted. ## Where the reasoning stops The cheap allocation depends on a free in-domain corpus existing and covering the regions that matter. In domains where public in-domain data is thin, the attacker falls back to steered generation and the bill rises steeply. And a copy fitted on any allocation still only agrees where its inputs came from - a corpus with a blind spot produces a copy with the same blind spot, whatever the aggregate agreement figure says.

  • Does this mean synthetic inputs are always a waste of budget?
    No. Uniform junk is near-worthless, but inputs steered toward where the attacker's partly-fitted copies disagree with each other concentrate near a boundary, which is where a label is worth most. They cost extra submissions to locate, because a top-1 verdict gives no direct read on proximity to a boundary. So they are the fallback when no free in-domain corpus exists, not the first choice.
  • A copy scores 97% agreement on the pool it was queried on. Why is that not evidence of a good copy?
    Because agreement is only meaningful next to the target's base rate on that pool. If the service answers benign for 97% of those inputs, a copy that always answers benign already scores 97% without having learned anything. Agreement has to be measured on a pool fixed in advance and drawn from the traffic the copy is meant to replace, with per-class figures broken out.
  • Would a rate limit change this analysis?
    It changes the calendar and the price, not the value of a label. A limit stretches a fixed submission count over more time or more accounts; it does not make a well-chosen input worth less or a junk input worth more. The allocation question - which files receive the submissions - is unchanged, which is why volume alone is a poor summary of what a campaign extracted.

saying these in an interview costs you the question

  • Assumes more queries always mean a better copy
  • Treats every input as carrying equal information
  • Reads high agreement on the queried pool as success
  • Thinks any input the model accepts is a useful probe

context