A perturbed feature vector fools a malware classifier — why may no such file exist?
answer
- the model never sees the file itself
- the extraction step runs one way only
- most vectors have no file behind them
- coordinates move in bundles, not alone
basics
~20 sFeatures are computed from a file by a fixed extraction step, and most perturbed vectors have no preimage that is also a working program. A feature-space success is therefore an upper bound on evasion, not a demonstrated one.
solid answer
~50 sA static classifier never sees the file; it sees a vector derived from it — sizes, counts, ratios, presence flags, summary statistics. The adversary can move freely in that vector space, but to actually evade they need a file whose extraction yields that vector and which still runs. Two things break the round trip. The extraction map is many-to-one and its image is a thin subset of the space, so most perturbed vectors correspond to no constructible file at all. And the coordinates are coupled: appending content moves size, section counts and summary statistics together, so a vector that shifts one coordinate in isolation is usually unreachable. The direction of the error matters when you read a result: a feature-space evasion rate over-states what the adversary can realise, and the honest number exists only after the artefact is rebuilt, exercised, re-extracted and re-scored.
go deeper
Know that a static file classifier scores a vector derived from the file, not the file itself, and that an adversary has to produce a real file rather than a vector.
Explain why the extraction map cannot be run backwards: it is many-to-one, its image is a thin subset of the space, and the coordinates move in correlated bundles rather than one at a time.
Demonstrate that you read the direction of the claim correctly — a feature-space number bounds the adversary from above — and that you would demand the rebuild-exercise-rescore chain before calling it an attack.
Be ready to decide what your evaluation programme is allowed to claim: whether feature-space probing is an acceptable cheap screen, and what evidence a finding needs before it drives a product change.
## Two spaces, and only one of them is real A classifier that scores files before they run almost never consumes raw bytes end to end. Something extracts a vector first: how large the file is, how many sections it declares, what fraction of it looks compressed, which capabilities it announces, counts of this and ratios of that. The model is a function on that vector. The adversary, however, must ship a *file*. That gives you two different spaces and a one-way bridge between them. Extraction maps files to vectors. Nothing maps vectors back to files. ## Why the bridge does not run backwards **The map is many-to-one.** Extraction is lossy by design — it summarises. Many different files produce the same vector, so a vector does not identify a file. **The map is nowhere near onto.** The set of vectors that any real file can produce is a thin, irregularly shaped subset of the space the model is defined over. An optimiser wandering that space will happily land on points describing a file with a negative-length region, mutually contradictory counts, or a compression ratio no real content achieves. Those points are not files. They are arithmetic. **The coordinates are coupled.** Almost nothing an adversary can do moves exactly one feature. Append content and you have changed total size, the count of regions, the byte-frequency statistics and probably an entropy estimate, all at once. So even a perturbation that lands inside the reachable set is usually reachable only as part of a bundle of correlated moves — the single-coordinate step the optimiser found is not on the menu. **And the artefact must still work.** Even among reachable vectors, the adversary can only use ones produced by a file that still does its job. That prunes the set again, and it prunes it by a criterion the feature space cannot see. ## The direction of the error This is what an interviewer is really probing. A result obtained purely in feature space **over-states** the adversary, because it grants them freedom they do not have. So a feature-space evasion rate is an upper bound on realizable evasion, and a defender reading one should not treat it as a demonstrated compromise. Equally, a red-teamer reporting one has not shown an attack; they have shown that the model is fragile in a space nobody can actually inhabit. There is a symmetric point worth knowing. Extraction throws information away, so behaviour the features never look at can change without moving the vector at all. A file that already scores benign can gain function the model has no coordinate for. The same gap that inflates paper attacks also conceals real ones, and both facts follow from the map being lossy. ## Which coordinates are genuinely cheap Not all features are equally hard to move. Ones that count content the author freely controls — bytes in regions nothing executes, padding, additional resources — move monotonically and predictably, and the adversary can push them as far as their size and start-up-latency ceiling allows. Ones derived from the code that must actually run — structure of control flow, capabilities the program genuinely needs — are expensive, because moving them risks the behaviour the adversary is trying to preserve. A useful mental model of this attacker is: they hold a few nearly free coordinates and a great many pinned ones, which is not what an unconstrained optimiser assumes. ## What an honest claim requires If you are handed a feature-space result and asked what it is worth, the answer is a pipeline, not a number: | stage | what it establishes | | --- | --- | | perturbed vector scores benign | the model is fragile somewhere in vector space | | a file producing that vector was constructed | the point is reachable at all | | the file was exercised and still worked | the artefact survived the edit | | the file was re-extracted and re-scored | the verdict holds on the real thing | Report the fraction lost at every stage. The number that survives all four is the claim; anything earlier describes the vector space rather than the threat. ## Why this matters beyond files Any model that consumes a derived representation of an artefact inherits this gap — flow summaries computed from packets, statistics computed from a document. Wherever the model's input is a summary and the adversary's lever is the underlying object, the two spaces come apart, and an attack demonstrated in the summary has not yet been demonstrated at all.
- Which feature coordinates are actually cheap for this adversary to move?The ones that only count content the author controls and can add without consequence: bytes in regions nothing executes, padding, extra resources. Those move monotonically and predictably up to the file-size and start-up-latency ceiling. Anything derived from code that must run — control-flow structure, capabilities the program genuinely needs — is expensive, because moving it risks the behaviour the adversary is trying to keep.
- How would you turn a feature-space result into a claim you would defend?Add a realizability stage. Construct the artefact, exercise it against a stated functional test, re-extract the features, re-score it, and count only candidates that pass all of those. Report the fraction lost at each stage. Without that chain, the number characterises the vector space rather than the adversary, and it over-states them.
- Does the gap between the two spaces ever run the other way?Yes. Extraction discards information, so behaviour the features never look at can be changed while the vector barely moves — a file that already scores benign can gain function the model has no coordinate for. The lossy map inflates paper attacks and hides real ones for exactly the same reason.
saying these in an interview costs you the question
- Reports a feature-space evasion rate as a working attack
- Assumes every vector has some file behind it
- Moves one feature without the coordinates coupled to it
- Thinks the classifier consumes the raw file directly
- Treats the extraction step as reversible