Why can a substitute trained on a target ranker's verdicts fool it despite far lower accuracy?
answer
- accuracy is the wrong yardstick
- agreement, not truth
- only near the attacked boundary
- trained on the target's own replies
- labels are metered, inputs are free
basics
~20 sThe attack needs agreement with the target near the boundary it crosses, not accuracy against ground truth. A stand-in well below the target's accuracy still points the right way locally, which is why its query budget stays small.
solid answer
~50 sAccuracy measures how often a model matches the *true* label across a whole distribution. An evasion attack needs something much narrower: that the stand-in and the target disagree about the same class in the same direction, in the small neighbourhood of the inputs actually being pushed across. The attacker fitted the stand-in on the target's own verdicts, so it is trained toward *agreement with the target*, not toward truth — and only in the region their query pool covered. A model that is wrong about most of the distribution can still be right about which way the boundary runs where the attack operates. That decoupling is what makes the exercise affordable: the labelled-query budget is a small fraction of what a faithful full-distribution copy would need, because most of that fidelity is not being used.
go deeper
Remember the headline: the attacker's model is trained on the target's answers, not on the truth, so being inaccurate about the world does not stop it from imitating the target where it matters.
Be ready to separate accuracy, fidelity across the distribution, and local agreement near the attacked boundary, and to say which of the three the attack actually consumes.
Show you know what to measure — label agreement in the attacked region and a replayed transfer rate — and that you would not size a defence against the cost of retraining your model from scratch.
Own the consequence for threat modelling: the attacker's entry price is set by local agreement, not by your training-data scale, so 'our dataset is proprietary and huge' is not a control you can put in a risk register.
## Three different things that all get called 'how good the substitute is' The confusion this question exists to clear up is that a substitute has three separate quality measures, and only the third one matters for evasion. 1. **Accuracy** — agreement with the *true* label, averaged over a distribution. This is what people quote by reflex, and it is the least relevant number here. 2. **Fidelity** — agreement with the *target's* label, averaged over a distribution. Closer, but still an average over regions nobody is attacking. 3. **Local agreement** — whether the stand-in and the target put the boundary in the same place, and orient it the same way, in the small neighbourhood of the inputs being crafted. This is the one the attack consumes. A crafted input is produced by following a direction until it crosses the stand-in's boundary. It fools the target when the target's boundary lies in roughly the same place, facing roughly the same way, *right there*. Everything the stand-in gets wrong about creatives that are nowhere near the attacked region costs the attacker nothing. ## Why the substitute is biased toward agreement in the first place The attacker did not train toward ground truth. They took inputs they already owned, submitted a fraction of them to the ranker, and used the returned verdicts as labels. If the ranker is wrong about a creative, the stand-in is trained to be wrong about it in exactly the same way. The training signal *is* the target's behaviour, so 'accuracy' measured against real eligibility rules is measuring the wrong thing twice over. This is also why the query budget is small. The attacker is not reconstructing a general-purpose eligibility model; they are learning a local map of one boundary. Unlabeled in-domain inputs cost them nothing — they own a large pool of creatives — so the only metered resource is labels, and the number of labels needed to place a boundary in one region is far below the number needed to learn the task. ## The wrong answer, stated plainly The common senior-level mistake is: *'a substitute has to be accurate, so this needs almost as much data as training the original — therefore it is impractical against a serious model.'* Both halves are wrong. The requirement is agreement where the attack operates, not accuracy; and a stand-in whose held-out accuracy sits well below the target's routinely produces directions that carry. Sizing the defence against the cost of retraining the target from scratch badly overestimates what the attacker has to spend. ## What actually predicts transfer - **Task similarity and distribution overlap.** The attacker's pool has to look like the traffic the target sees. This, not architecture, is the binding requirement. - **Coverage of the attacked region.** If the queries never landed near the kind of creative being pushed through, the stand-in never learned that stretch of boundary and transfer collapses there. - **Attack type.** Untargeted evasion — any wrong answer will do — transfers markedly better than targeted evasion, which demands a specific chosen outcome and therefore agreement about a specific piece of the boundary rather than merely about which side is wrong. - **Perturbation size.** Directions pushed further past the local boundary transfer more often, at the cost of being more visible and more likely to be caught by other controls. ## How you would measure it The honest instrument is not the stand-in's accuracy report. It is: take a sample of crafted inputs, replay them at the target, and record the fraction it misreads. Alongside that, label agreement between stand-in and target on a held-out sample of inputs drawn *from the attacked region* is a far better predictor than either model's accuracy. A defender reading an attacker's write-up should look for those two numbers and treat a quoted substitute accuracy as decoration. ## Where the decoupling stops helping Agreement near a boundary is cheap only while the boundary stays put. Retraining moves it. Randomised or input-dependent preprocessing on the target means the stand-in is approximating something that is not a fixed function. And a target that was itself hardened against evasion has a different boundary shape from the naively fitted stand-in, so local agreement degrades exactly where the attacker needs it most. None of these make the technique fail outright; they raise the query budget and lower the transfer rate, which is the currency this whole exercise is denominated in.
- What would you measure instead of the substitute's accuracy?Two things: label agreement between substitute and target on held-out inputs drawn from the region being attacked, and the transfer rate — the fraction of locally crafted inputs the target actually misreads when replayed. Both are measured against the target's behaviour rather than against ground truth, which is the only comparison the attack cares about.
- Why does a targeted attack transfer worse than an untargeted one?An untargeted attack succeeds if the target lands on any wrong class, so it only needs the two models to agree that the input has been pushed off its original side. A targeted attack demands one specific outcome, which requires agreement about a particular stretch of boundary between two particular classes — a much stronger condition, and one that fails more often.
- If a defender retrains the ranker weekly, how does that change the picture?It caps how long a fitted stand-in stays useful, so the attacker's fixed query cost has to be amortised inside one retrain interval. That is a genuine cost control: it does not stop the technique, it shortens the window over which the up-front spend pays back, and pushes the attacker toward fewer, higher-value inputs.
saying these in an interview costs you the question
- Says the substitute must be nearly as accurate as the target
- Estimates the query budget as a full training set
- Quotes substitute accuracy as evidence of transfer
- Confuses agreement with the target and agreement with ground truth
- Assumes transfer is uniform across the input distribution