skip to content

Evasion & Adversarial Examples

You will learn how imperceptible gradient-crafted perturbations flip a model's prediction, from white-box FGSM/PGD through black-box query attacks and transferability via surrogates. This is the canonical adversarial-ML interview topic — expect to explain why the attack works geometrically, not just name the acronyms.

on this pageshow

explore

questions

page 1 of 2

Why doesn't a printed patch that defeats a store camera's model need to be imperceptible?

level: juniorimportance: must knowfreq 66%

answer

  1. ask what a defender could compare the frame to
  2. the constraint came from a different threat model
  3. no original scene exists to diff against
  4. magnitude is free; area is not
  5. bold structure is what survives a lens

basics

~20 s

Imperceptibility only matters when someone can compare the input to an original. A camera sees a real scene, with no original to compare against, so the real limits are covered area, the angles it must work from, and human inspection.

solid answer

~50 s

The imperceptibility rule comes from the digital threat model, where the attacker edits a file that already exists, so every change is measured against that original and a tiny perturbation radius keeps it invisible. Nothing like that exists in front of a camera: the model sees whatever light reached the sensor, and no clean version of the scene is ever available to diff against. So magnitude is free — the attacker can print bold, saturated, high-contrast structure, which is also what survives distance, printing and re-sampling. What replaces the radius is a geometric budget: how much of the object's surface may be covered, from which standoff distances and approach angles it has to hold, and whether it looks odd enough that a person pulls the item aside. Those are the numbers a physical evasion result has to quote; a perturbation radius is meaningless for it.

go deeper

for a junior

Be ready to say where the imperceptibility rule comes from and why it does not apply in front of a camera. Naming area, viewing angle and human inspection as the real limits is enough at this level.

for a middle

You are expected to explain the mechanics: a norm and a radius are a stated threat model, a scene has no reference input, and high-contrast structure is what survives printing and distance in the first place.

for a senior

Show the judgment: decide whether a visible artefact is a real risk for a given deployment by asking who inspects the object, what area it covers, and over which approaches it holds — not by asking whether it looks obvious in a photograph.

for a principal

Own the framing question of whether your deployment faces an adversary who can place objects at all. If nobody can reach the scene, this whole class is theoretical; if they can, imperceptibility was never the control you were relying on.

## The rule you are being asked about Most people meet adversarial examples in their digital form: an attacker takes an existing input, adds a small structured change, and the model's answer flips while a human sees nothing. The "small" there is not a vibe — it is a stated constraint, a norm plus a radius (every pixel may move at most a little, or the total change may have at most so much energy). That pair *is* the threat model, and imperceptibility is a **consequence** of choosing a tiny radius, not a law of the field. Candidates then carry the rule forward to a printed artefact placed in a scene and conclude that a visible patch is a broken attack. That is the wrong answer, and knowing why is the whole point of this question. ## Why the rule does not transfer to a camera The imperceptibility constraint is only meaningful because a **reference exists**. In the digital setting there is an original file: a defender, a reviewer or a hash can in principle compare the submitted input to it, so a large change is a change somebody could point at. In front of a lens there is no reference. A fixed overhead camera in a checkout lane photographs whatever is in the lane. There is no "unmodified" version of that moment to compare the frame against — the frame *is* the scene. An attacker who prints something and puts it on an item has not edited an input; they have arranged reality. So the quantity the digital threat model bounds is not just unbounded here, it is undefined. A second reason pushes the same way. A digital perturbation is a fragile, high-frequency pattern; sending it through optics, distance, printing and compression destroys most of it. Structure that survives capture has to be **large and high-contrast**, which is the opposite of imperceptible. The attacker is not choosing to be visible out of laziness — visibility is what buys survival through the capture pipeline. ## What the budget becomes The constraints are still real; they are just different in kind: - **Area.** What fraction of the object's visible surface the artefact may cover. This is the closest thing to a radius, and it is the number a physical result must quote. Cover the whole item and you have not evaded the model, you have replaced the object. - **Viewpoint.** The range of standoff distances and approach angles over which it must keep working. A result at one pose is a result at one pose. - **Looking unremarkable.** A *social* constraint, not a metric one: staff, other shoppers, or a reviewer looking at the footage should have no reason to single the item out. A garish label on packaging can pass where the same thing on a face would not. Notice that only the third is about being unnoticed at all, and it is about being unnoticed by **people**, under whatever inspection actually happens, not about staying within a distance bound. ## What visibility buys, and what it costs Trading imperceptibility away is a trade, not a free win. The attacker gains magnitude, and with it robustness: bold structure still reads at three metres and under bad lighting. They pay for it in every other column — the artefact has to hold across poses they do not fully control, it can be seen and removed, and the honest success rate is far below what a simulated version of the same attack reports. ## Directions to keep straight - A visible patch is **not** evidence of a weak attack; imperceptibility was never the goal outside the digital threat model. - A perturbation radius quoted for a physical artefact is a category error, not a strict number. - One photograph of a successful patch proves it worked **at that pose, that distance, that lighting** — not that the deployment is evadable. - "A person would notice" is a claim about the review that actually happens. If nobody inspects the packaging, it is not a control. ## How to answer in a loop Say where imperceptibility comes from (a digital threat model with a reference input), say why the reference is missing in a scene, then name the three things that actually bound the attacker: area, viewpoint range, and passing human inspection. That answer shows you understand the constraint rather than reciting it.

  • Does that make physical evasion easier than the digital kind?
    No — it is a trade. The attacker gives up imperceptibility and gains magnitude, but takes on constraints the digital attacker never had: the artefact must hold across distances and angles they only partly control, it survives printing and optics imperfectly, and it can simply be seen and removed. Honest physical success rates land far below the simulated equivalent.
  • Is there any physical setting where being unnoticed still constrains the attacker?
    Yes, but it becomes a social constraint rather than a metric one. If a person inspects the object — staff handling packaging, a guard watching footage, a reviewer sampling frames — the artefact has to look unremarkable in that specific inspection. That is a question about who looks and how hard, not about a distance bound.
  • If magnitude is unbounded, why not cover the whole object?
    Because area is the budget that actually bites. Covering the object entirely is not evasion of the model, it is presenting a different object, and it fails the plausibility constraint immediately. A result is only interesting when the covered fraction is small enough that the item still reads as itself to a person.

Editing a photo is forgery against a known original; putting a printed card on a shelf is set dressing. Only the forger has to worry about the change being visible next to the real thing.

saying these in an interview costs you the question

  • Says a patch fails as an attack because it is visible
  • Quotes a perturbation radius for a printed physical artefact
  • Assumes tiny digital noise survives optics and distance
  • Treats one successful photograph as general evasion
  • Believes the defender compares each frame to a clean original

context

open as a page

A static malware classifier scores a file before it runs — why can't an attacker add a small perturbation?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Because the file is discrete and must still work after the edit. Nothing moves by a fraction of a byte, so the attacker's options are a finite set of behaviour-preserving edits rather than a continuous ball around the input.

open as a page

Why does a perturbation that flips an image classifier on a file usually fail once printed and photographed?

level: juniorimportance: must knowfreq 60%

basics

~20 s

A digital perturbation is a precise pattern of tiny per-pixel changes. Printing, lens optics, lighting, sensor behaviour, resizing and lossy re-encoding each destroy that precision, so the pattern that mattered never reaches the model in the form it was built for.

open as a page

How can an input built against the attacker's own classifier fool one they have never queried?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Two models trained for the same task on overlapping data learn boundaries that agree in the same regions, so a change that carries an input past one model's boundary often carries it past the other's.

open as a page

Why does an attacker who only sees a ranker's accept/reject verdict train their own model on those verdicts?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A verdict-only endpoint returns no gradient. Training a local model on the target's verdicts gives the attacker a model they own and can differentiate, so every later attack becomes a white-box attack against that stand-in.

open as a page

An attacker holding an inspection model's weights flips it with a tiny pixel change - why does large random noise fail?

level: juniorimportance: must knowfreq 78%

basics

~20 s

The change is a direction, not noise. It is read off how the classifier's loss responds to each pixel, so every pixel moves the way that hurts the model. Random noise of that size points nowhere useful.

open as a page

An attacker holding a model's weights can work inside a fixed perturbation radius or minimise the change instead - what does each measure?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A fixed-radius attack answers 'can this input be broken within this much change?' and returns a success rate at a radius somebody chose. A minimising attack answers 'how much change did this input need?' and returns a distance per input.

open as a page

Why is an adversarial perturbation budget's norm and radius a threat model rather than a tuning knob?

level: juniorimportance: must knowfreq 66%

basics

~20 s

The norm and radius together describe an attacker: a per-coordinate cap lets every input value move a little, a sparse budget lets a few move a lot. Change the pair and you have evaluated a different adversary.

open as a page

With no weights, only a fraud API's returned risk score, how does an attacker get a search direction?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The attacker buys the direction instead of computing it. Sending slightly varied transactions and differencing the returned risk scores shows how the score moves - the same information a gradient gives, at one paid call per probe.

open as a page

A face-verification gate returns only accept or reject — why doesn't hiding scores stop evasion?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Hiding scores removes only the attack family that needs numbers. With a bare accept or reject, an attacker can still start from an input the gate already accepts and shrink it toward the input they want, using each decision as a yes/no probe.

open as a page

An attacker prints a patch to fool a fixed store camera — what replaces the perturbation radius as their budget?

level: middleimportance: must knowfreq 52%

basics

~20 s

Geometry replaces magnitude: the fraction of the object's surface the artefact may cover, the standoff distances and approach angles it must hold across, and what a printer and sensor can reproduce. A physical claim without those numbers says nothing.

open as a page

Why is the gradient an attacker uses against a defect classifier taken with respect to the input, not the weights?

level: middleimportance: must knowfreq 68%

basics

~20 s

The input is the only thing the attacker can change. The deployed weights are frozen and out of reach, so the same derivative machinery is pointed at the other variable: how the loss responds to each input value.

open as a page

Why doesn't per-account query rate limiting bound a universal perturbation attack?

level: middleimportance: must knowfreq 55%

basics

~20 s

Because the fitting happened offline, on data the attacker already owned, and applying the finished artefact costs zero queries. A rate limit meters interaction with the model, and this attack needs none at attack time.

open as a page

What is a universal adversarial perturbation, and how does it differ from a per-input one?

level: juniorimportance: should knowfreq 52%

basics

~20 s

A universal perturbation is one fixed change fitted once over a sample of inputs and then applied unchanged to inputs the attacker has never seen or queried. A per-input attack recomputes a different change for every target.

open as a page

A perturbed feature vector fools a malware classifier — why may no such file exist?

level: middleimportance: should knowfreq 42%

basics

~20 s

Features are computed from a file by a fixed extraction step, and most perturbed vectors have no preimage that is also a working program. A feature-space success is therefore an upper bound on evasion, not a demonstrated one.

open as a page

Why does an attacker printing a marking for a gate camera optimise over many capture conditions at once?

level: middleimportance: should knowfreq 48%

basics

~20 s

Because the artefact has to work after the capture chain, not before it. Optimising for expected success across sampled angles, distances, lighting, print reproduction, resize and re-encode buys durability, and each condition added to that set costs achievable success rate.

open as a page

Why doesn't keeping a classifier's architecture and training corpus private stop an attacker who cannot query it?

level: middleimportance: should knowfreq 58%

basics

~10 s

Transfer needs a similar task and an overlapping data distribution, not your architecture or your rows. Anyone solving the same job over the same kind of inputs has that overlap by construction.

open as a page

A report shows an evasion attack at 96% white-box and 38% transferred to an unqueried model — what explains the gap?

level: middleimportance: should knowfreq 44%

basics

~20 s

What transfers between two models is the direction the attack pushed in, not the exact point it landed on. Crafted inputs sit barely past the source model's boundary, and the target's lies close but not identically.

open as a page

Why can a substitute trained on a target ranker's verdicts fool it despite far lower accuracy?

level: middleimportance: should knowfreq 54%

basics

~20 s

The attack needs agreement with the target near the boundary it crosses, not accuracy against ground truth. A stand-in well below the target's accuracy still points the right way locally, which is why its query budget stays small.

open as a page

On a defect classifier, why does an attacker's iterated search beat a single step at the same allowed change?

level: middleimportance: should knowfreq 60%

basics

~20 s

One step trusts a local straight-line reading of the model across a whole jump, and that reading goes stale immediately. Iterating takes small moves, re-reads the direction each time, and pulls back inside the allowance.

open as a page

Why does an attacker minimising the size of a flipping perturbation pay far more compute per input than one filling a fixed radius?

level: middleimportance: should knowfreq 44%

basics

~20 s

Filling a fixed radius is bounded work: every candidate is allowed by construction and the search stops when its schedule ends. Minimising holds two things in tension - keep the flip, shrink the change - and must re-verify the flip at every smaller size.

open as a page

What does an attacker give up by reusing one fixed perturbation across many inputs?

level: middleimportance: should knowfreq 36%

basics

~10 s

Control over which inputs flip. One fixed change flips a fraction of a population rather than a chosen input, and it leaves a repeating artefact that can be recognised once it has been seen.

open as a page

What does a total-energy budget over a forecaster's input window permit that a per-step cap forbids?

level: middleimportance: should knowfreq 44%

basics

~20 s

A total-energy budget can be spent entirely on one step, giving a spike the per-step cap forbids. The cap forbids that spike but lets every step shift by the full cap together, which the energy budget does not.

open as a page

On a paid scoring API, why does estimating an attack direction cost more queries as the input gets wider?

level: middleimportance: should knowfreq 45%

basics

~20 s

Each estimate is built from probes, and covering the input coordinate by coordinate needs at least one probe apiece. Widen the feature vector and the calls per direction rise in proportion, then multiply by every step of the search.

open as a page

Why does a label-only attack on an accept/reject gate start from an already-accepted input?

level: middleimportance: should knowfreq 46%

basics

~20 s

With no numbers coming back, every rejection looks identical, so there is nothing to climb from outside. Starting from an input the gate already accepts gives a known-good point, and the search only has to keep that answer while shrinking the difference.

open as a page

A printed patch makes an overhead checkout detector report no box at all — why is that harder to catch than a wrong label?

level: seniorimportance: should knowfreq 41%

basics

~20 s

A wrong label produces a record someone can contradict; a suppressed detection produces no record at all, which is indistinguishable from an empty lane. Monitoring built on wrong predictions sees nothing, and aggregate accuracy on normal traffic stays flat.

open as a page

A malware verdict service returns only benign or malicious and rate-caps submissions — what does that stop?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It removes score-based attacks, which need returned confidences to difference; it does not remove label-only boundary search. The cap prices that search rather than blocking it, and every submission also hands the defender the sample.

open as a page

A printed marking opens a camera-controlled gate once in five presentations — how do you report that finding?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Report a rate over presentations together with the conditions it was measured under: angle and distance range, lighting, camera and capture settings, and how many trials. Intermittence is the expected shape of a physical attack, not evidence that it does not work.

open as a page

A red-team demo only ever evaded the attacker's substitute — what does that establish about the production ranker?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It establishes that a model the red team built themselves is evadable, which was true by construction. Without crafted inputs replayed against the production ranker and a measured transfer rate, the finding states a hypothesis, not a demonstrated exposure.

open as a page

As a reviewer, what does the distribution of smallest perturbations an attacker needed per input tell you that one fixed-radius success rate does not?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A fixed-radius success rate is one threshold cut through the distance distribution. The distribution itself shows where your data actually sits, which slices sit closest, and whether the radius that was reported stood on a plateau or on a cliff.

open as a page

showing 1–30 of 44