A report shows an evasion attack at 96% white-box and 38% transferred to an unqueried model — what explains the gap?
answer
- the same attack against two models
- how far past the boundary it landed
- direction generalises, location does not
- any wrong class versus one chosen class
- ask which columns the number is missing
basics
~20 sWhat transfers between two models is the direction the attack pushed in, not the exact point it landed on. Crafted inputs sit barely past the source model's boundary, and the target's lies close but not identically.
solid answer
~50 sWhite-box, the attacker optimises against the exact model being scored and stops as soon as the input crosses its boundary — so the point sits barely past it. Transferring that point to a second model asks a different question: is it also past a boundary that lies nearby but not in the same place? The direction generalises because both models learned overlapping features; the location does not, so a large share of crafted inputs fall back on the correct side. The gap narrows with a larger input budget, and with attacks built against several source models at once rather than converged tightly on one. It widens sharply when the goal is a specific wrong class rather than any wrong answer. Before reading 38% as a robustness figure, ask which budget, which goal, and which source model produced it.
code
text · 9 linestransfer evaluation - 1,000 held-out documents
source model target access edit budget goal success
own fine-tune A itself white-box <=12 edits untargeted 96.2%
own fine-tune A employer model B transfer <=12 edits untargeted 38.4%
own fine-tune A employer model B transfer <=12 edits targeted 5.2%
own fine-tune A employer model B transfer <= 4 edits untargeted 11.3%
ensemble of 3 employer model B transfer <=12 edits untargeted 58.0%
...go deeper
Remember the shape of the result: an attack run against a model you hold succeeds far more often than the same attack handed to a model you have never touched.
Explain the geometry — the crafted point sits barely past the source model's boundary while the target's boundary lies nearby but not identically — and name what moves the rate: budget, goal, and how many source models were used.
Demonstrate that you read the columns before the number. Say which figures are not comparable, and note that a low transfer rate bounds one adversary's yield rather than establishing anything about your model.
Weigh what the transfer row costs the attacker against what the white-box row costs them, and decide which adversary your assurance argument is actually written for before signing anything that quotes either number.
### Reading the two numbers The two rows are answers to different questions. The white-box row asks whether the attacker can find a point that crosses the boundary of a model whose loss surface they can evaluate directly. The answer is almost always yes, because that is an optimisation with full feedback; a figure near 100% there is the expected result and says very little on its own. The transfer row asks something much harder: does a point found against **one** model also sit on the wrong side of **another** model's boundary, when the attacker never touched that second model at all? ### Why the second question is harder Two models trained for the same task over overlapping data agree about roughly where the boundary goes, because they learned overlapping features. That agreement is what makes transfer possible, and it is agreement about a **region and a direction**, not about a surface to the millimetre. Meanwhile the crafted point is typically as close to the source model's boundary as the attacker's search allowed — an attack that stops when the label flips lands just barely over. Barely over one boundary is frequently still on the correct side of a neighbouring one. That single geometric fact accounts for most of the drop. ### The terms that move the transfer number - **The input budget.** More allowance means the attacker can push further past the source boundary, buying margin that survives the mismatch. It also makes the change more conspicuous to anyone who inspects the input, so the budget is a real trade rather than a free dial. - **Overfitting to the source.** An attack driven to convergence against a single model exploits idiosyncrasies of that one model's surface. Building against several source models at once, or against varied versions of the input, yields directions that fewer models disagree with — which is why the ensemble row in a well-built report is higher. - **Targeted versus untargeted.** 'Any wrong answer' is a huge region and models agree about a lot of it. 'This specific wrong class' asks two models to agree about which particular alternative sits nearest, which they frequently do not, and targeted transfer rates run far below untargeted ones. - **Distance between the models.** Divergence in training data, in architecture family, and in how much of any shared base representation survived fine-tuning all reduce agreement. ### What the number does not establish A transfer rate is only meaningful with its columns. Reported alone, 38% does not say what the input budget was, whether the goal was targeted, how many source models were used, or whether the source and the target descend from the same public base — and each of those can move the figure by tens of points. Two rows from different reports with different budgets are not comparable at all. Be careful about the direction of the inference as well. A low transfer rate means **the attack that was actually run** mostly failed. It does not establish that the target is robust: a stronger source, a bigger budget, or several sources combined is the adversary's ordinary next step, so the figure is a floor on their capability and never a ceiling. ### Reading it as a defender The useful reading is comparative and longitudinal. Track transfer rate against a fixed, stated source and a fixed budget across releases, and treat a change in it as a signal about how your model's boundary has moved relative to publicly buildable models. Treat the absolute number as an estimate of one adversary's yield, not as a property of your model. And note what the transfer row costs the attacker in a way the white-box row does not capture at all: nothing. No queries, no rate limit, no probing pattern in your logs — 38% of attempts succeeding at zero interaction cost is a different operational picture from 96% of attempts succeeding only if the adversary already holds your weights. ### On discrete inputs Where inputs are text or binaries rather than continuous signals, the budget is a count of behaviour-preserving edits rather than a magnitude, and the same reasoning applies with one extra caveat: results measured on continuous inputs do not carry over numerically, so a transfer rate borrowed from a different modality is not evidence about yours.
- Why does building the attack against several source models raise the transferred rate?Because it stops the attack overfitting one model's surface. A direction that all of several models agree pushes toward error is far more likely to be a direction the unseen target also agrees with, whereas a direction driven to convergence against one model partly exploits that model's idiosyncrasies. The cost is that the attacker must own or obtain several suitable models.
- Why is targeted transfer so much weaker than untargeted?Untargeted only needs two models to agree that the input has left the correct class, which is a large region of agreement. Targeted needs them to agree on which specific alternative it landed in — a much finer coincidence between two boundaries that merely lie near each other. The rate difference is routinely an order of magnitude.
- Does a low transfer rate let you call the model robust?No. It says the attack that was run, from that source, at that budget, mostly failed. A stronger source model, a larger budget or an ensemble of sources is the adversary's normal next move, so the figure is a floor on their yield rather than a bound on it. Report it with its columns and treat it as one measurement, not a property.
saying these in an interview costs you the question
- Reads a low transfer rate as proof of robustness
- Quotes a transfer rate without budget or goal
- Thinks targeted and untargeted transfer at similar rates
- Blames the gap on the target being more accurate
- Compares transfer numbers from reports with different budgets