skip to content

A red-team demo only ever evaded the attacker's substitute — what does that establish about the production ranker?

level: seniorimportance: should knowfreq 42%

answer

  1. the attacker owns that model
  2. success there was guaranteed
  3. one number is missing
  4. replay a sample at the target
  5. low transfer is a cost, not an all-clear

basics

~20 s

It establishes that a model the red team built themselves is evadable, which was true by construction. Without crafted inputs replayed against the production ranker and a measured transfer rate, the finding states a hypothesis, not a demonstrated exposure.

solid answer

~50 s

A white-box attack on a model you own succeeds close to always — that number carries no information about the target. The missing column is the **transfer rate**: replay a stated sample of the crafted creatives at the production ranker, under its real rate limit and preprocessing, and report the fraction it accepts. Ask for the rest of the threat model too: the labelled-query cost paid to fit the stand-in, the perturbation family and radius used, the steps and restarts, and whether the replay went through the same front door a real advertiser would use. And be careful in both directions: a low transfer rate is not 'no finding', because the attacker can craft many candidates cheaply offline and only needs the ones that land; while a failure to transfer proves nothing about robustness, only that this stand-in did not agree with the target well enough this time.

code

text · 9 lines
text
finding R-104  creative eligibility ranker (production)
  access assumed        : verdict only, free tier, 40 submissions/hour
  substitute            : fitted on 4,000 paid verdicts, in-domain pool
  threat model          : L-infinity, radius stated
  attack steps/restarts : 40 / 1
  success vs substitute : 98.4%   <- white-box, on a model the red team owns
  success vs production : --      <- not measured
  reproduced            : 1 run
  ...

go deeper

for a junior

Remember that beating a model you built yourself is not evidence about someone else's model; the crafted inputs still have to be tried against the real one.

for a middle

Be able to name the missing measurement — transfer rate on a stated sample replayed through the real interface — and the threat-model details that must accompany any evasion number.

for a senior

An interviewer expects you to read the number in both directions: a low transfer rate is an exploitation cost, and a zero is a statement about one attack, not about the model's robustness.

for a principal

Own the reporting standard: define up front what evidence converts a surrogate-only demonstration into a rated finding, so severity is never argued from an experiment whose result was fixed by its own setup.

## Why the demo, as reported, is not yet a finding The red team fitted a stand-in on verdicts bought from the free tier, then attacked it with full knowledge of its weights. Near-total success against that model is the expected outcome of a white-box attack on an undefended model you built yourself. Reporting it as an evasion rate is reporting the experiment's setup, not its result. What the demo does establish is narrower and still worth writing down: that a stand-in *can* be fitted at that query cost from the access level the attacker had, and that the crafted inputs exist. Everything about the production system is still hypothesis. ## The column that is missing The bridge from stand-in to target is transferability, and it is measured, never assumed. The reviewer's job is to send the report back for one number and its denominator: of N crafted creatives replayed at the production ranker, how many were accepted? That replay has to happen through the real interface, because the production path may include preprocessing, re-encoding, size normalisation or a separate detector that the local stand-in never modelled — any of which can destroy a direction that worked offline. ## Reading the number once it arrives This is where reviewers get the direction of the claim wrong, in both directions. - **A low transfer rate is not an all-clear.** Crafted candidates are free once the stand-in exists; the attacker can generate thousands offline and submit the ones they like. If the payoff is one ineligible creative entering an auction, a rate in the low tens of percent is an operational attack, not a curiosity. The right reading of a low rate is a *cost*, expressed in submissions per success against the rate limit. - **A failure to transfer is not evidence of robustness.** It says this stand-in did not agree with the target well enough in this region, at this perturbation size, with this query pool. A better-covered pool or a larger radius may change it. The claim 'the model resisted' is only ever about the attack that was run. - **A high transfer rate is a finding about the boundary, not about a leak.** Nothing was stolen; two models trained on overlapping data simply agreed about direction. Do not let it be triaged as a credential or weight-disclosure issue. ## What else the reviewer should demand - **The threat model as a pair**: the perturbation family and its radius, stated together. A rate quoted without both describes nothing repeatable. - **The access level assumed**, and whether it matches what a real advertiser account has: verdict only, rate limit, per-account throttling, any review step in the path. - **The query cost actually paid** to fit the stand-in — this is the attacker's entry price and it is the number that decides how the risk is rated. - **The attack's strength**: steps and restarts. A weak attack that transfers is more alarming than a maximal one that transfers. - **Reproducibility**: a demo that reproduces once in five tries is reporting a distribution, and the report should give the spread rather than the best run. ## Closing the gap when you cannot touch production If policy forbids submitting crafted creatives to the live auction path, the gap is still closable: replay against a staging deployment carrying the same weights and the same preprocessing chain, and state clearly that the measurement is on staging. What is not acceptable is closing the gap rhetorically — presenting the stand-in's own success rate as the production number, or projecting a transfer rate from a different system. ## The judgment call Most of the time the replay is cheap: a few hundred submissions inside an existing rate limit, on an account the red team already has. The reviewer should treat 'we never ran it against the real thing' as an open item rather than a verdict either way, and should say what evidence would close it — a stated sample size, a transfer rate with its spread across repeats, and the query cost of the stand-in — before the finding is rated, accepted or dismissed.

  • The replay comes back at 12% transfer. Do you close the finding?
    No — you re-rate it. Candidates are free to generate once the stand-in exists, so 12% converts into roughly eight submissions per success, which the rate limit prices in time rather than in difficulty. The right output is an exploitation cost with an assumed payoff, not a pass. Closing it would treat a rate as a barrier.
  • The replay comes back at 0%. What may you honestly claim?
    Only that this stand-in, fitted from this query pool, at this perturbation radius and attack strength, produced nothing that carried through the production path. It is evidence about one attack, not about the model. State the pool, the radius, the steps and the sample size, and say explicitly that a differently covered pool or a larger radius is untested.
  • Policy forbids submitting crafted creatives to the live auction. How do you close the gap?
    Replay against a staging deployment carrying the same weights and the same preprocessing and re-encoding chain, and label the number as a staging measurement. Note any known difference between the paths, since preprocessing is exactly what destroys transferred directions. What is never acceptable is quoting the substitute's own success rate as though it were the production figure.

saying these in an interview costs you the question

  • Reports the substitute's own success rate as the target's
  • Treats a low transfer rate as no finding
  • Reads failure to transfer as proof of robustness
  • Quotes an evasion rate with no perturbation family or radius
  • Triages successful transfer as a weight or data leak

context