A printed marking opens a camera-controlled gate once in five presentations — how do you report that finding?
answer
- a rate, not a yes or no
- successes over presentations, with n
- name the conditions or the number says nothing
- cost per attempt sets severity
- tuned and tested on the same conditions?
basics
~20 sReport a rate over presentations together with the conditions it was measured under: angle and distance range, lighting, camera and capture settings, and how many trials. Intermittence is the expected shape of a physical attack, not evidence that it does not work.
solid answer
~50 sA physical attack's result is a **distribution**, so "it reproduces one time in five" is the finding, not a failure to reproduce. Report it as successes over presentations under a named condition set: the angle and distance range, illumination and time of day, the specific camera and its capture settings, whether the pipeline's resize and re-encode were in the loop, and the trial count — twenty presentations is a wide interval, so say so. Disclose whether the artefact was tuned and tested under the same conditions, because that overstates durability the same way testing on your training data does. Then set severity from the economics rather than the headline: at a gate, a presentation costs a sheet of paper and a few seconds, retries are unlimited and unattributed, and a one-in-five rate means a few minutes of attempts. The number that would genuinely reduce risk is the rate across the camera fleet and across a day's light.
code
text · 10 linesattack evaluation - printed marking, gate reader
condition set presentations read as target
digital only (perturbed file, no capture) 200 99.0%
printed, head-on, overcast, camera A 120 41.7%
printed, +25 deg, overcast, camera A 120 12.5%
printed, head-on, direct sun, camera A 60 6.7%
printed, head-on, overcast, camera B (fleet) - not run
...
not recorded: capture resolution, encode settings, night lighting,
whether the artefact was tuned on the rows it was tested ongo deeper
Know that a physical attack result is a success rate over repeated presentations, and that a single success or a single failure is not the finding.
Be able to list what has to accompany the rate: trial count, angle and distance range, lighting, the deployed camera and its capture settings, and whether the resize and encode were in the loop.
Show that you set severity from cost per attempt and attempts per hour rather than from the headline rate, and that you disclose tuning-versus-testing overlap without being asked.
Own the reporting standard: what your organisation will accept as evidence of a physical evasion finding, and the refusal to let a digital number sit unlabelled next to a field number in the same table.
## Intermittence is the result, not a defect in it Engineers used to deterministic bugs read "reproduces sometimes" as "not confirmed". For a physical evasion finding that instinct is wrong. The artefact was fitted to a distribution of capture conditions, and each presentation draws a fresh sample from the real one: a slightly different angle, a cloud moving, an auto-exposure decision, a different frame chosen by the reader. A per-presentation success **rate** is the only well-formed statement available. A finding that says "it works" without a rate, and a triage note that says "could not reproduce" after three tries, are both malformed. ## The columns that make the rate mean something A physical rate is uninterpretable without the conditions it was drawn under. A usable report carries: - **Trials and successes**, not a percentage alone. Four of twenty and forty of two hundred are very different evidence for the same 20%. - **Geometry** — the range of angles and distances presented, and whether they were varied or held fixed. - **Illumination** — overcast, direct sun, night under the gate's own lighting; which hours were covered. - **The capture chain** — the camera model actually deployed, its resolution and capture settings, and whether the pipeline's resize and lossy encode were in the loop or bypassed. - **Fleet and time coverage** — one camera at noon is one condition; the risk lives in the rate across the installed cameras and across the day. - **Tuning disclosure** — whether the artefact was optimised under the same conditions it was then measured under. If it was, the rate is optimistic for the same reason that evaluating on training data is. The last one is where reports most often mislead without intending to. ## Setting severity from the economics, not the percentage The question a decision-maker actually needs answered is not "how strong is the attack" but "what does it take to succeed once". At a gate: - a presentation costs a printed sheet and a few seconds; - retries are unlimited, cheap and mostly unattributed; - a failed read looks exactly like a dirty plate, a bad angle or rain, so it generates no signal by default; - one success is the whole payoff, because the gate opens. Under those conditions a 20% per-presentation rate is not a curiosity: a handful of attempts gets through, and the expected number of attempts is small enough to fit inside a normal dwell time at the barrier. Contrast a setting where each attempt is expensive, attributable, or rate-limited after a few failures — the same 20% is a much weaker finding there. **Severity is set by attempts-per-hour and cost-per-attempt, not by the headline rate.** This also points at the defence worth recommending, which is operational rather than model-side: repeated failed reads from the same lane inside a short window are a cheap, high-signal detection, and the attacker's need to retry is precisely what makes it work. ## The trap in a clean digital number sitting in the same report When a report puts a digital figure and a field figure in the same table, readers compare them as if they were rows of one experiment. They are not: they were measured over different input distributions, and only the second says anything about the gate. Either drop the digital row or label it explicitly as an upper bound obtained without the capture chain. ## What to write A defensible finding reads roughly like: *under the deployed camera and capture settings, with the pipeline's resize and encode in the loop, presented from within the normal driver's angle and distance range under overcast daylight, the artefact was read as the target marking in 24 of 120 presentations. Not yet measured: direct sun, night lighting, or the other camera model in the fleet. The artefact was tuned under a different lighting condition than the one measured.* That is a smaller claim than "we defeated the gate", and it is one you can defend in a room.
- The engineering team says the finding is not reproducible. What do you say?That a physical attack is a distribution, so intermittence is the expected shape rather than a disproof. The remedy is more presentations under logged conditions, not a binary repro attempt: three failed tries at an untested angle is a sample of size three. I would agree a joint protocol — fixed angle and distance ranges, stated lighting, the deployed camera and settings, a set trial count — and report successes over presentations from it.
- Which single missing column most weakens such a report?The capture conditions, including the camera and its settings. Without them the rate names no threat model, cannot be compared to a retest, and cannot be extrapolated to the rest of the fleet. Trial count is a close second, since a percentage without an n hides whether the finding rests on four successes or four hundred.
- Would you recommend a model-side fix here?Not first. The cheapest durable control in this setting is operational: alert on repeated failed reads from one lane in a short window, since the attacker's need to retry is what the low per-presentation rate forces on them. Model-side hardening is slower, costs clean read accuracy on legitimate markings, and is measured against one condition set that the next camera change invalidates.
It is closer to reporting a fishing spot's catch rate than a compiler bug: nobody expects every cast to land, and the useful report says how many casts, in what water, on what day.
saying these in an interview costs you the question
- Treats intermittent physical success as not reproducible
- Reports a percentage without trials or conditions
- Compares a digital rate and a field rate as one experiment
- Ignores that retries at a gate are free and unlogged
- Hides that the artefact was tuned on the tested conditions