Why does an attacker who only sees a ranker's accept/reject verdict train their own model on those verdicts?
answer
- the reply carries no direction
- own a model you can differentiate
- verdicts become your training labels
- one fixed cost, many free attacks
- the stand-in is disposable
basics
~20 sA verdict-only endpoint returns no gradient. Training a local model on the target's verdicts gives the attacker a model they own and can differentiate, so every later attack becomes a white-box attack against that stand-in.
solid answer
~40 sAn evasion attack needs a direction to push an input in, and that direction is read off the gradient of a loss with respect to the **input**. A hosted eligibility ranker returns an accept/reject verdict and nothing else, so the attacker cannot compute one against it. Two options remain: walk the target's boundary with many queries per input, or spend a query budget once labelling their own collection of in-domain creatives with the target's verdicts and fitting a local model to them. That local model — a substitute, or surrogate, in the usual vocabulary — has weights the attacker owns, so directions can be optimised against it offline and for free, then replayed against the target. It is disposable: it exists only to be differentiated, and it is thrown away when the target changes.
go deeper
Be ready to say plainly that a verdict-only endpoint gives the attacker no direction to move in, and that fitting a local model on the target's replies gives them one they can differentiate for free.
Explain the mechanics: the target is used as a labelling service, the local model is fitted on those labels, attacks are optimised white-box against it, and crafted inputs are then replayed at the target.
An interviewer expects you to name the cost structure and the failure cases — query budget, transfer rate below one, staleness after retraining, and targeted attacks transferring worse than untargeted ones.
Own the framing that hiding scores and rate-limiting are cost controls, not boundaries: they price out one attack family while leaving the substitute route intact, and that difference should drive what you claim in a design review.
## The problem the attacker actually has An adversarial input is not noise. It is a specific *direction* in input space, found by asking how a model's loss changes as the input changes. Against a model whose weights you hold, that direction is immediately available. Against a hosted ad-creative eligibility ranker that answers `eligible` / `rejected` on a rate-limited free tier, it is not: the reply carries one bit, and one bit tells you nothing about which way to move. So the attacker faces a choice between two ways of buying the information they are missing. **Pay per input.** Boundary-walking attacks need only the returned verdict. Starting from an input the target already accepts, they probe repeatedly and edge toward the input they actually want accepted, using each verdict to tell which side of the boundary they are on. This works with no scores at all, but the queries are spent *per input being attacked*, and a rate limit turns that into calendar time. **Pay once, up front.** The alternative is to stop attacking the target directly and instead build something local that behaves enough like it. The attacker takes a pool of in-domain creatives they already own — unlabeled inputs are free to them, they make creatives all day — submits some fraction of them, records the verdicts, and trains their own classifier on those (input, verdict) pairs. The verdicts are the labels; the target is being used as a labelling service. ## What that buys A model whose weights, architecture and gradients the attacker holds. Against it, every attack is a white-box attack: directions can be optimised offline, at whatever number of steps and restarts they like, with no queries, no rate limit and no logging on the defender's side. Candidate inputs are then replayed against the real ranker. The bridge is transferability: two models trained on a similar task over an overlapping input distribution tend to agree about which way the boundary lies in a given region, so a direction found on one often crosses the other's boundary too. It is a tendency, not a guarantee — the attacker's real figure of merit is the **transfer rate**, the fraction of inputs crafted locally that the target also misreads. ## What a substitute is not - **Not a copy of the service.** It is not meant to replace the ranker or match it across the distribution; it may be badly wrong on inputs nobody is attacking. - **Not an explanation.** Fitting a simple model to a black box in order to *explain* one of its predictions is a different exercise with a different success criterion (faithfulness to the explained decision). - **Not compression.** Compressing a model into a smaller one assumes free access to rich outputs from a teacher you own; here access is metered, the output is one bit, and the artefact is thrown away. - **Not a requirement to match architecture.** The substitute does not need the target's architecture, its hyperparameters, or its training data — only enough agreement in the region being attacked. ## What it costs, and where it stops working The costs are real and stated up front: a labelled-query budget, the compute to fit the stand-in, and a transfer rate below one hundred percent, meaning some crafted inputs are wasted. Against that, the cost is *fixed*: once fitted, the stand-in serves every later attack at no further query cost. It stops paying when the target is retrained often enough that the stand-in goes stale before the budget is amortised; when the attacker's pool of inputs does not overlap the region they actually need to attack, so the stand-in never learned that part of the boundary; when the target adds randomised or input-dependent preprocessing that the local model does not model; and when the attack is *targeted* — forcing a specific chosen class transfers noticeably worse than merely forcing any wrong answer. From the defender's side, this is why a verdict-only response and a rate limit are cost controls rather than boundaries: they remove the score-based route and raise the bill, and they do not remove the substitute route at all.
- Does the substitute have to share the target's architecture for this to work?No. What matters is a similar task and an overlapping input distribution, so the two models agree about the boundary in the region being attacked. Architecture, depth and hyperparameters can differ completely. A substitute that is smaller and much weaker than the target still produces directions that carry, because the attack only needs local agreement, not global equivalence.
- After crafting inputs on the substitute, what does the attacker still have to measure?The transfer rate against the real target. Success against the stand-in is close to guaranteed by construction — it is a white-box attack on their own model — so it measures nothing. Only replaying a sample against the ranker, under its real rate limit and preprocessing, tells them what fraction actually lands and therefore what the whole exercise is worth.
- How is this different from compressing a large model into a smaller one?Compression assumes you own the teacher, can query it freely and can read its full output; the goal is fidelity across the distribution and the student is the product. Here access is metered, the reply is one verdict, the goal is agreement only where the attack operates, and the fitted model is a throwaway instrument whose product is the crafted input.
It is like reverse-engineering a bouncer's rules by sending in a few hundred friends and noting who got in, then rehearsing against your own copy of the doorman instead of arguing with the real one every night.
saying these in an interview costs you the question
- Claims the attacker reads gradients out of the hosted endpoint
- Says the substitute must match the target's architecture
- Assumes the substitute needs the target's actual training data
- Treats success against the substitute as success against the target
- Describes the perturbation as random noise rather than a direction