skip to content

A face-verification gate returns only accept or reject — why doesn't hiding scores stop evasion?

level: juniorimportance: must knowfreq 62%

answer

  1. query-only attacks are not one family
  2. one of them never reads a number
  3. the search starts from something already accepted
  4. the decision itself is the signal
  5. coarsening the reply prices attacks

basics

~20 s

Hiding scores removes only the attack family that needs numbers. With a bare accept or reject, an attacker can still start from an input the gate already accepts and shrink it toward the input they want, using each decision as a yes/no probe.

solid answer

~50 s

Attacks that only get to query a model split two ways. Score-based ones buy a search direction by probing and differencing the numbers that come back — those genuinely die when the endpoint stops returning numbers. Decision-based ones need nothing but the returned decision: the attacker begins from an input the gate already accepts, then repeatedly reduces the difference toward the input they actually want to submit, keeping only the changes where the answer is still `accept`. The flip of that one word is the whole signal. So coarsening the reply deletes one family and multiplies the calls the other needs, often by an order of magnitude — a price, not a wall. After coarsening, the question worth asking is not "is this safe now" but "how many decisions can this attacker consume per identity, and does anything count them".

go deeper

for a junior

Be ready to say that query-only attacks come in two kinds and only one of them reads numbers. Knowing that a bare accept/reject is still usable signal is the point of the question.

for a middle

An interviewer expects you to explain what the survivor uses instead of a number — the flip of the returned decision — and why that costs far more calls to reach the same result.

for a senior

Show that you convert the defence into currency: name what coarsening removes, then say which measurable limit now binds the attacker and whether anything in the deployment counts it.

for a principal

Own the wording. Decide whether a control document may claim infeasibility from a reply schema, and insist the claim be restated as a cost control with the compensating limit named and tested.

## The claim being corrected A verification gate answers with one word — `accept` or `reject` — and the written justification for that design is that an attacker with no confidence number has nothing to optimise. Half of that is true. The half that is false is the half that decides whether the system is attackable. ## What "only the endpoint" means The access assumption here is the strictest realistic one: the adversary has no weights, no architecture and no gradients. They can send an input and read a reply, and that is all. (Two other meanings of *black box* float around — the software-testing sense, and "black-box model" meaning uninterpretable. Neither is this. Here the phrase is purely about what access the adversary is assumed to have.) ## The split that answers the question Query-only attacks divide into two families that behave completely differently when you take the numbers away. - **Score-based.** The adversary probes with related inputs and differences the numbers that come back to estimate a search direction — a metered, noisy stand-in for a quantity that is free to anyone holding the weights. Every sample is a paid call, and the estimate degrades as the input gets larger. This family depends entirely on the reply carrying a number. - **Decision-based.** The adversary uses only the class the endpoint names. They start from an input the gate *already* accepts, and reduce the difference between that input and the one they actually want accepted, keeping the changes under which the reply is still `accept` and discarding the ones that flip it to `reject`. The information used is one bit per call: did the answer change or not. Stop returning scores and the first family is gone. The second is untouched — it never read a number in the first place. ## Why one bit is enough A trained classifier partitions its input space; "still accepted" versus "now rejected" is exactly a test for which side of that partition an input sits on, and that test is precisely what the endpoint hands out for free with every call. An attacker who can ask that question repeatedly can steer, because each answer rules out part of the space they were searching. What they cannot do is take large informed steps, because nothing tells them *how far* they are from the flip — only *that* they flipped. That is why the same result costs this family roughly an order of magnitude more calls than a score-based attack, and why the family is measured in **decisions consumed**, not in a perturbation radius. ## What hiding scores actually buys It buys two real things, and they are worth having: 1. It removes the score-based family outright. 2. It multiplies the number of live calls the remaining family needs. Both are cost controls. Neither is a boundary. A cost control is a legitimate thing to deploy and a legitimate thing to claim — as long as it is claimed as what it is. The dangerous move is writing "the endpoint returns only a decision, therefore adversarial inputs cannot be found" into a control document, because that sentence asserts infeasibility while the mechanism only asserts expense. ## The question that replaces it Once the reply is coarsened, the attacker's binding constraint is no longer what they can read; it is how many decisions they are allowed to consume. So the useful follow-ups are about the attempt budget: how many attempts an identity gets before lockout, whether cooldowns merely convert calls into calendar time, whether the counter counts only *failed* attempts (a walk that stays on the accepted side spends most of its calls on accepts), and whether attempts can be spread across identities, sessions or channels. That is a measurable quantity someone can report. "We do not return scores" is not. ## What this is not It is not a claim that the attack is cheap — it is usually far from cheap, and against a gate with a tight attempt budget it may be entirely unaffordable, which is a perfectly good outcome. It is not a claim about an adversary who holds the weights; that is a different vantage with different economics. And the input the walk arrives at is not random noise that happened to slip through: random change of comparable size essentially never flips a trained classifier. The difference is that this one was *found*, one yes/no answer at a time. ## In an interview Say the split out loud — score-based versus decision-based — say which one coarsening removes, say what the survivor uses instead, and then convert the defence into its honest currency: how many decisions does it cost the adversary, and what limits that number.

  • Which query-only attack family does hiding the scores actually remove?
    The score-based one. Those attacks estimate a search direction by probing with related inputs and differencing the numbers that come back, so a reply with no number in it leaves them nothing to difference. The decision-based family is unaffected: it only ever consumed the returned class, one bit per call.
  • After coarsening the reply, what actually caps the remaining attacker?
    The number of decisions they can consume. Their cost is denominated in calls, so the binding limits are attempt caps per identity, cooldowns that convert calls into elapsed time, and whether attempts can be spread across identities or channels. Those are measurable and reportable; the reply schema is not.
  • Does a decision-only attacker need to know the gate's architecture?
    No. The walk consumes only the returned decision, so architecture, layer counts and training details are irrelevant to it. Architecture matters for a different move — transferring an input crafted against a stand-in model — but a label-only walk against the real gate needs none of it.

A lock that never tells you how close you are, only whether it opened, still yields to someone allowed to keep trying — the single bit is enough to work with, it just costs far more attempts.

saying these in an interview costs you the question

  • Says a label-only endpoint cannot be attacked at all
  • Assumes every query-only attack needs confidence numbers
  • Calls the resulting input random noise that got lucky
  • Treats hiding scores as a boundary rather than a price
  • Cannot name the two query-only attack families

context