A turnstile face matcher reports 91% robust accuracy at a per-pixel radius — what does that say about an attacker limited to head pose and occluded area?
answer
- different units, not different amounts
- degrees and area, not magnitude
- one acceptance, not an average
- a stand-in only has to agree near the boundary
basics
~20 sAlmost nothing. The figure prices an attacker who edits pixels within a magnitude budget. Someone standing at the gate is bounded by viewing angle and how much of the face they may cover — units the defense never trained in and cannot report on.
solid answer
~50 sThe reported number is true and about a different adversary. It says: an attack inside a per-pixel magnitude budget, run at whatever strength they ran it, failed 91% of the time on their test set. The attacker at the turnstile has no file to edit; their budget is a range of head pose and a fraction of the face they can cover. In pixel terms that change is enormous — far outside any radius the vendor trained against — so the defense has no unit in which to state whether it holds. Worse, the deployment's payoff is not average accuracy: it is one chosen identity being accepted once. So the evidence you actually need is a targeted-acceptance rate for a chosen enrolled identity, measured across the stated pose range and occlusion fraction, on the capture path you deploy.
go deeper
Recall that a robustness figure names the kind of change it was measured against, and that someone standing in front of a camera is not making that kind of change.
Explain why a viewing-angle change is not inside a per-pixel budget at all, and why that makes the vendor's number silent rather than merely optimistic about this attacker.
Show you would restate the requirement in the deployment's own units — pose range, occluded fraction, targeted acceptance for a chosen identity, on the real capture path — and would insist the measurement move there rather than arguing about the vendor's figure.
Be ready to decide what the organisation does with a gap that no amount of re-training in pixel units will close, and who owns the residual risk while the measurement is being redone.
## Two adversaries, one number The vendor's figure describes an adversary who can edit the input as data, constrained by how far each pixel may move. Your deployment's adversary stands in front of a camera. They cannot touch the file at all. Their constraints are **geometric and area-based**: a range of head pose and viewing angle they can hold naturally at a gate, and a fraction of the face they can plausibly have covered without drawing attention. Those two budgets are not a strong and a weak version of one thing. They are stated in different units. A change of viewing angle moves essentially every pixel by a large amount, so it is nowhere inside a small per-pixel radius; conversely, nothing inside that radius corresponds to turning one's head. The correct description of the vendor's coverage over this attacker is **not measured** — not "weakly robust." ## What the attacker's vantage adds A realistic attacker here also has more than physical presence. A previously released on-device weight file for an earlier generation of the same matcher is a **differentiable stand-in**: they can compute against it freely, with no queries to your gate and no log entries for you to notice. It does not need to match your deployed model's accuracy, and it does not need the same architecture. It only needs to agree with the deployed matcher **near the boundary being attacked** — the region separating "this is the enrolled badge-holder" from "this is not." That is a much weaker requirement than functional equivalence, and it is why transfer works at all. Combine the two and the picture is: offline optimisation against a stand-in, delivery through pose and occlusion, against a defense that priced neither. ## The payoff is not the metric the vendor reported Robust accuracy is an average over a test set. The attacker at your turnstile does not care about the average. They want **one targeted acceptance**: one chosen identity admitted, once. Two consequences: 1. **Averages hide it.** A matcher can hold its aggregate numbers exactly while being reliably wrong on one carefully chosen pair. The reported figure would not move. 2. **The relevant rate is per-identity and targeted**, not aggregate. Some enrolled identities are much easier to reach than others, and the attacker chooses which one to be. So even a robustness number measured in the right units, but reported as an average, would answer a question you did not ask. ## What you should ask for instead Stated in the units of the adversary you actually face: - targeted acceptance rate for a chosen enrolled identity, **across the pose range** the gate accepts and **at the occlusion fractions** it tolerates; - measured on the **deployed capture path** — the same camera geometry, lighting envelope and enrolment pipeline — because a result obtained on clean stored imagery is a different experiment; - with the attacker's assumed vantage stated: whether they were allowed a stand-in model, and whether they were allowed to observe accept/reject outcomes at the gate. And from the vendor, the honest scoping line: which families were evaluated, at which radii, and which were not evaluated at all. ## Triaging a finding here A related trap: an attempt that succeeds once in five tries is still a finding. A gate that admits the wrong person on the fifth attempt is not "mostly holding" — the attacker can retry, and retries at a turnstile are cheap and unremarkable. Report success rate over attempts under the stated budget, not a binary. ## The one-line version The 91% is a true statement about a magnitude-bounded attacker on a test set. Your attacker is bounded by degrees and by area, and wants one specific acceptance. Nothing in the reported figure covers that, and nothing about it is dishonest — it simply answers a different question.
- What does a released weight file for the previous generation of the matcher buy the attacker?A differentiable stand-in they can optimise against offline, with no queries to your gate. It does not need to match your deployed model's accuracy or architecture — only to agree with it near the boundary being attacked, which is why transfer works from a substantially weaker model.
- The vendor points out that aggregate accuracy on their test set is unchanged. Does that clear the deployment?No. A targeted acceptance is one pair read the wrong way; an average over a test set is designed not to move for that. Ask for a per-identity targeted-acceptance rate under the pose and occlusion budget, which is the quantity the attacker is actually optimising.
- The attack reproduces once in five attempts. How do you report it?As a success rate under a stated budget, not as a flake. At a turnstile, retries are free and unremarkable, so one in five is an effective bypass. State the pose range, the occlusion fraction, the vantage assumed, and the attempts-to-success — those are what make the finding actionable.
saying these in an interview costs you the question
- Reads the vendor's figure as covering a physically present attacker
- Treats pose and occlusion as merely larger perturbations
- Accepts an aggregate accuracy number for a targeted-acceptance risk
- Assumes a stand-in model must match the deployed one's accuracy
- Dismisses an attack that succeeds one attempt in five