Why can one message within a fixed edit budget satisfy both an inbound-mail classifier and its detector?
answer
- not two independent gates
- same object, two conditions
- intersection of two error regions
- text is discrete, no small radius
- budget = edits that keep it useful
basics
~20 sThe attacker searches one allowed set of edited messages for a point on the wrong side of the classifier's boundary and the ordinary side of the detector's. Two conditions in one budget means more search, not a wall.
solid answer
~50 sBoth models read the same message, so the attacker is not facing two independent gates - they are looking for a single point that lands in both error regions at once, inside one budget. On text that budget is a count of behaviour-preserving edits that leave the message still useful for the sender's purpose, not a small numeric radius; email is discrete, so there is no such thing as nudging a character by a hundredth. The intersection is usually non-empty because both boundaries are fit from finite data over an enormous input space, and both have large regions they read wrongly. What the second condition costs is search: more candidate messages, more edits spent, sometimes a larger budget before anything works. The detector genuinely wins only when satisfying both requires more edits than leave the message useful to the attacker at all.
go deeper
Know that the detector and the classifier see the same message, so one crafted message can satisfy both. Avoid describing the pipeline as two independent checks whose chances of failing multiply.
Be able to state the search as finding one point in the intersection of two error regions inside one budget, and to say what the budget is for text: behaviour-preserving edits, not a numeric radius.
Demonstrate that you would measure effort-to-success with and without the detector under an attack that knows about both, and that you know the attacker's real limit is edits that keep the message useful.
Frame for the business what the second model bought: a multiplier on attacker effort and a false-alarm bill, not a reduction in the class of inputs the classifier can be fooled by.
### Two conditions, one object When a detector sits in front of an inbound-mail classifier, it is tempting to picture a pipeline of independent checks whose failure probabilities multiply. They do not multiply, because the two checks are not independent draws - they are two conditions evaluated on the *same message*, and the adversary controls that message. The adversary's situation is therefore: within an allowed set of messages (the original message plus at most so many behaviour-preserving edits), find one that the classifier delivers and the detector calls ordinary. Geometrically, they are looking for a point in the intersection of the classifier's error region and the detector's error region, restricted to the allowed set. Adding the detector shrinks the target; it does not make it empty by construction. ### Why the intersection is usually non-empty Both models are learned functions fit from finite samples over a space vastly larger than those samples. Neither draws its boundary where the true concept lies; each draws one that happens to separate its training data, and each has substantial regions where it is confidently wrong. A detector's false-negative region contains, by definition, crafted messages it reads as ordinary. Nothing in how it was trained makes that region disjoint from the classifier's false-delivery region - the two models see the same features of the same messages, and their errors are correlated rather than independent. This is also why detector accuracy does not compose into a security number. A detector that is right on 95 percent of a held-out mail sample says nothing about whether the worst point inside the attacker's allowed set is in its remaining 5 percent, because the attacker is not sampling from that distribution - they are searching for the exception. ### What the budget actually is here On continuous inputs a perturbation budget is stated as a norm and a radius: every coordinate may move a little, or total energy is bounded. Email text is discrete and that framing does not carry over - you cannot move a word by a hundredth. The realistic budget for a message is a count of edits that preserve what the message is for: substituted wordings, inserted or reordered filler, formatting changes, added unused content. The constraint that matters is not a number of pixels but whether the message still does the sender's job after the edits. That constraint is what makes the whole thing finite, and it is what the detector is competing against. ### What the second model costs the adversary The measurable effect is a multiplier on effort. Satisfying two conditions takes more candidates tried, more edits spent, and often more submissions against the deployed pipeline before one message gets through. That factor - effort to a delivered message with the detector, divided by effort without it - is the detector's value stated as something you can defend. It is a real quantity and it can be large. It is simply not the same quantity as "crafted messages do not reach the classifier". ### Where it actually stops working for the attacker The honest failure case for the adversary is budget exhaustion in the domain's own terms. If the edits needed to satisfy the detector as well as the classifier push the message past the point where it still achieves the sender's purpose - the link is gone, the instruction is unreadable, the payload no longer parses - then the attack failed for a reason that is about the domain, not about the detector's accuracy. That is the shape of a genuine win from a screen on discrete inputs, and it is worth naming precisely because it is checkable: you measure the edits required, not the catch rate. ### The evaluation consequence If the two conditions are evaluated jointly by the attacker, they must be evaluated jointly by whoever measures the defence. A number produced by an attack that only ever targeted the classifier tells you what happens to an adversary who ignored your detector, which is the adversary you were least worried about. The measurement that carries information is: under an attack that knows both models are there, at a stated edit budget, how much more effort does a delivered message cost, and what fraction of real mail does the detector quarantine at that operating point.
- Why can't you multiply the two models' error rates to get a pipeline number?Because multiplying assumes independent draws, and there is only one message. The adversary chooses it, and the two models read the same features of it, so their errors are correlated and adversarially selected rather than sampled. The pipeline's worst case is the intersection of two error regions inside the allowed set, which no product of average accuracies describes.
- How is the attacker's budget expressed for text, since there is no small perturbation radius?As a count of behaviour-preserving edits - substitutions, filler, formatting, reordering - bounded by the requirement that the message still serves the sender's purpose. Results carried over from continuous inputs do not transfer: there is no coordinate you can nudge by a hundredth, and the search space is discrete, which changes both what the attack can do and what a budget means.
- Does making the detector deliberately hard to differentiate help?Only against attacks that need a smooth signal from it. It raises cost for that family and leaves search that needs nothing but the final verdict untouched, and a defence that mainly makes the attacker's search awkward has a history of being reported as robustness it does not have. Treat it as a price increase and measure it as one.
saying these in an interview costs you the question
- Multiplies the two models' accuracies into a pipeline number
- Treats detector and classifier as independent gates
- Assumes their error regions cannot overlap
- Applies a small numeric perturbation radius to email text
- Cannot say what the attacker's budget is measured in