Why does a label-only attack on an accept/reject gate start from an already-accepted input?
answer
- two rejects come back identical
- a flat signal cannot be searched
- the search order inverts
- feasible first, then shrink the difference
- the reply is a constraint check, not a gradient
basics
~20 sWith no numbers coming back, every rejection looks identical, so there is nothing to climb from outside. Starting from an input the gate already accepts gives a known-good point, and the search only has to keep that answer while shrinking the difference.
solid answer
~40 sA score-based attacker can start anywhere, because the returned numbers tell them which direction improves things. A label-only attacker cannot: two rejected inputs come back as the same word, so a rejected starting point carries no information about how to move. Inverting the order fixes this. Begin from an input the gate already accepts — for a verification gate, something that genuinely matches the enrolled reference — and treat the goal not as "become accepted" but as "stay accepted while looking more like the input I actually want to submit". Now every call is informative, because the attacker knows what the answer was before the change and can keep the changes that preserve it. The attack is feasibility-first, then optimisation, and its cost is counted in decisions consumed rather than in a perturbation radius.
go deeper
Remember the shape: the attacker begins from something the gate already accepts rather than from the input they want accepted. Knowing the direction of the search is enough at this level.
Explain why the naive ordering fails — two rejected inputs return the identical word, so the signal is flat — and then state the inversion and what it costs in calls.
Be ready to price it. Say that this family is counted in decisions consumed, that progress has diminishing returns, and that its endpoint reflects a stopping decision rather than a minimum.
Decide what your organisation will accept as a reported result for this family, and insist that any difference figure be published alongside the decision count and the attempt limits in force.
## The problem the ordering solves An adversary querying a verification gate receives one word: `accept` or `reject`. Nothing else — no margin, no ranking, no error text that varies. The question is how any search can be steered by that. Consider the naive ordering first: take the input you want accepted, which the gate rejects, and change it until it is accepted. Every call in the early phase returns `reject`. Two very different rejected inputs return the identical word, so nothing in the reply distinguishes a change that moved you closer from a change that moved you further away. The signal is flat, and a flat signal cannot be searched — you are reduced to guessing, and random change of realistic magnitude essentially never flips a trained classifier. This is exactly why the naive ordering does not work and why candidates who assume it conclude, wrongly, that a label-only endpoint is unattackable. ## The inversion The family instead begins from a point where the answer is already known to be `accept` — for a verification gate, an input that genuinely matches the enrolled reference. From there the objective is restated: - **not** "reach acceptance", which is unreachable blind, - **but** "remain accepted while reducing the difference to the input I actually want to submit". Now every call is informative in the only way a one-bit reply can be. Before each change the answer is known; after it, the reply either still says `accept` — in which case the reduction is kept — or says `reject`, in which case it is discarded and the search continues from where it was. The attacker is not descending a loss they cannot see. They are keeping a constraint satisfied while a distance shrinks, and the returned word is the constraint check. ## What that changes about the attack's economics Three consequences follow, and they are what a middle-level interviewer is listening for. **It is feasibility-first, then optimisation.** A score-based attack finds a direction and follows it. A decision-based attack starts already "successful" in the trivial sense and spends its entire budget making that success useful. There is no phase where it holds nothing. **Its cost is denominated in decisions, not in a radius.** The attacker never learns how far they are from the flip, only that they flipped. Information arrives one bit at a time, so the natural unit of the attack is the number of decisions consumed — typically an order of magnitude more than a score-based attack needs for a comparable result. Quoting a perturbation radius for this family without saying how many decisions bought it is quoting half a number. **It has no natural stopping point.** The difference shrinks quickly at first and then slowly, because each further reduction takes more confirmations to establish. The walk therefore ends when the attacker decides the remaining difference is small enough for their purpose, or when the attempt budget runs out — not at any minimum. Describing the endpoint as "the minimal accepted input" overstates it; it is whatever the decisions bought. ## Why the flat-signal problem is not fixable by returning less Operators sometimes reason in the other direction: if fewer fields make the attack harder, fewer still should make it impossible. But the accepted/rejected distinction is not a field the operator chose to add — it *is* the product. The gate exists to answer that question, so the one bit the family consumes cannot be withheld without withholding the service. That is the structural reason this family sets a floor on what coarsening the reply can achieve, and why the remaining lever is the number of decisions an adversary is permitted to consume rather than the content of each reply. ## Common confusions to keep straight - **This is not adding noise until something passes.** The changes kept are the ones the gate's own answers selected; noise of comparable size does not flip a trained classifier. - **This is not the score-based family with the numbers rounded off.** The two use different information and invert the search order; one is not a degraded version of the other. - **The starting point is not the target.** It is a scaffold that guarantees the constraint holds at step zero, and the whole campaign is about moving away from it. ## In an interview Lead with the flat-signal observation — two rejects are the same word — and then state the inversion in one sentence: begin inside the accepted region and shrink while the accept holds. Follow it immediately with the cost consequence, that this family is counted in decisions rather than measured by a radius, because that is the sentence that shows you have thought about what the attack costs and not just how it is shaped.
- Why is random trial and error not an equivalent approach for the same attacker?Because the changes that matter are structured, not arbitrary. Random change of realistic size essentially never flips a trained classifier's answer, so a random search consumes attempts and learns nothing. The walk works precisely because each kept change was selected by the gate's own reply rather than by chance.
- Why does progress slow down as such a walk continues?Early reductions are large and easy to confirm; later ones are small, so more decisions are needed to establish that the accept still holds after each. Returns diminish while the price per unit of progress rises, which is why the attacker stops at good enough rather than at any minimum.
- Does the walk end at the smallest accepted difference?No. It ends where the attacker stopped paying — when the remaining difference was adequate for their purpose or the attempt budget ran out. Reporting the endpoint as minimal overstates it; report it together with how many decisions were consumed to reach it.
It is the difference between guessing a stranger's address in the dark and standing in their doorway and edging outward while the porch light stays on — one gives you no feedback at all, the other gives you a yes or no every step.
saying these in an interview costs you the question
- Thinks the attack descends a loss it cannot observe
- Believes a rejection carries usable direction without a score
- Assumes the walk converges on a true minimum
- Confuses it with adding noise until something passes
- Quotes a difference size without the decisions it cost