A static malware classifier scores a file before it runs — why can't an attacker add a small perturbation?
answer
- the constraint is not a distance
- the artefact must still work afterwards
- bytes do not move by fractions
- a finite catalogue of harmless edits
basics
~20 sBecause the file is discrete and must still work after the edit. Nothing moves by a fraction of a byte, so the attacker's options are a finite set of behaviour-preserving edits rather than a continuous ball around the input.
solid answer
~50 sOn an image the threat model is a norm and a radius: every coordinate may move a little, and every point in that ball is a legal input. An executable has neither property. Bytes are discrete — there is no 0.01 of a byte — and the artefact has to keep doing what its author built it to do after the edit, which is a functional predicate, not a distance. So the adversary's move set is a finite catalogue of behaviour-preserving edits: content appended where the loader never executes, section padding, reordering of independent operations, no-op insertion, equivalent instruction sequences. The attack becomes a combinatorial search over that catalogue with a functional check on every candidate, rather than a gradient step with a projection back into the ball. That is why an evasion figure quoted at an L-infinity radius carries nothing here, and why image-domain attack numbers do not transfer.
go deeper
Recall the contrast: an image threat model states a norm and a radius, while an artefact that must keep running gives the adversary a set of allowed edits instead. Be able to say why a tiny change is meaningless on bytes.
Explain what replaces the ball: a finite catalogue of behaviour-preserving edits, a combinatorial search over it, and a functional check on every candidate because feasibility here is a predicate rather than a distance.
Show that you would ask for the move set and the functional pass rate before believing any evasion number, and that you would refuse to carry an image-domain robustness figure across to a classifier that scores files.
Own the framing call: which threat model your product actually faces, and whether spending on norm-bounded robustness buys anything at all when your inputs are artefacts nobody can perturb by a fraction.
## What a threat model actually states An adversarial threat model has two halves: what the adversary can see, and what they are allowed to change. In the image setting the second half is written as a norm and a radius. An L-infinity ball says every coordinate may move by at most some amount; an L-2 ball bounds total energy; a sparse budget says a few coordinates may move a lot. Those are different adversaries, and a robustness number quoted without both the norm and the radius is quoting nothing. That framing carries a hidden property which makes the whole machinery work: **every point inside the ball is a legal input**. You can take a step, land anywhere in the neighbourhood, and still hold something the model will happily consume. That is why an iterative attack can take many small steps along a direction read off the loss with respect to the input and simply clip back into the allowed set after each one. ## Why that breaks on an artefact that has to keep working A file that a static classifier scores before it is allowed to run breaks the framing twice. First, it is **discrete**. A byte takes an integer value. There is no small move; the smallest change you can make is a whole byte, and a whole byte is not small in the sense the continuous framing means. Second, and far more important, **legality is not proximity**. A file one byte away from a working program is very often not a working program at all — a length field no longer matches, a checksum fails, an offset points somewhere wrong, an instruction stream stops making sense. The adversary's requirement is not 'stay close to the original'; it is 'still do the job'. That is a predicate over behaviour, and a predicate does not define a neighbourhood you can project onto. There is no clip step, because there is nothing continuous to clip into. Feasibility can only be decided by exercising the candidate and seeing whether it still works. ## What replaces the ball The adversary works from a finite catalogue of edits known to leave behaviour intact: content appended to regions that are never executed, padding inserted inside a section, reordering of operations that do not depend on each other, insertion of instructions with no effect, substitution of one instruction sequence for an equivalent one. Some of those moves carry a size parameter; none of them is a distance. So the adversary's budget is stated in different units. Which edit families are permitted at all. How many edits may be applied. How much growth in file size and start-up latency the artefact can absorb before it stops being usable in the wild. That budget is the analogue of a radius, and confusing the two is the mistake this question exists to catch. ## Why continuous results do not carry over Three reasons, and they are worth being able to say out loud. 1. **The search is combinatorial, not differentiable.** There is no useful gradient with respect to a byte you can only move in whole steps and only in ways that keep the program running. 2. **The feasible set is thin and irregular, not a ball.** Much of the continuous literature leans on the geometry of a convex neighbourhood; here the set of legal points is a sparse, awkward subset of the space, and results about balls say nothing about it. 3. **The cost profile inverts.** In the image setting the expensive resource is queries to the model. Here the expensive resource is verification: proving that each candidate still runs, and still does its job, before you even bother scoring it. A red-teamer's time on this problem goes mostly into functional testing. ## The same shape appears wherever the input is an artefact A network flow must still parse and still deliver its payload. A sentence must still mean what it meant. In each case a distance budget is replaced by a functional constraint over a discrete space, and the honest way to state the threat model is to name the edits available and the ceiling on how many you may apply — not to borrow a radius from a different modality. ## The contrast worth remembering In the continuous setting the useful fact is that an adversarial perturbation is a *direction*, not noise: random change of the same magnitude essentially never flips a trained classifier, while a structured one does. On a discrete artefact even that direction has nowhere to go, because the coordinates are not free to move — most of them are pinned by the requirement that the thing still executes. The attack is not a smaller version of the image attack. It is a different problem wearing the same word.
- Iterative attacks project back into the allowed set after each step. What replaces that projection here?Nothing continuous. Feasibility is a predicate you can only decide by exercising the artefact, so the loop becomes: propose an edit, verify the behaviour survived, then query the model. Infeasible candidates are discarded outright rather than projected onto a boundary, which is why the search is combinatorial and why its cost is dominated by functional verification rather than by model queries.
- Does the smaller move set make a discrete classifier harder to evade?Not reliably. The set is smaller and much harder to search automatically, but the adversary needs only one artefact that both runs and scores benign, and a large share of a static model's features read regions whose contents the file's author fully controls. Fewer moves, but several of them are cheap and effectively unbounded in size.
- Where does the norm-and-radius framing still apply outside images?Where the model genuinely consumes continuous values and nearby points are all legal inputs. Even then, validity constraints narrow the ball hard — integer-valued fields, allowed ranges, relationships that must hold between coordinates, and in some domains a real cost the adversary must pay per unit of change. The norm alone is still not the whole threat model.
Nudging a photograph is like turning every dial a fraction. Editing a working program is like editing a contract: only certain substitutions leave the meaning intact, and there is no such thing as a one-percent change.
saying these in an interview costs you the question
- Quotes an L-infinity radius for a file's bytes
- Treats a perturbed feature vector as if it were a file
- Assumes image evasion results transfer directly
- Calls the adversarial edit random noise
- Thinks any byte in the file can be changed freely