skip to content

Evasion & Adversarial Examples

You will learn how imperceptible gradient-crafted perturbations flip a model's prediction, from white-box FGSM/PGD through black-box query attacks and transferability via surrogates. This is the canonical adversarial-ML interview topic — expect to explain why the attack works geometrically, not just name the acronyms.

on this pageshow

explore

questions

page 2 of 2

How do you choose the norm and radius for red-teaming a forecaster whose input window participants partly supply?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Derive both from what a participant can legally post, not from convention. Restrict the coordinates to the ones they write, size the radius by the postings their own volume supports, and measure the payoff as the decision the forecast moved.

open as a page

Your black-box evasion test on a metered API ran out of budget with no evasion - what does that establish?

level: seniorimportance: should knowfreq 35%

basics

~10 s

A budget-exhausted run bounds the spend, not the model: at this price per call, this sample count per direction and this many steps, no evasion was reached. It says nothing about a better-funded adversary.

open as a page

A verification gate locks an identity after five failed attempts — what does that bound for a label-only attacker?

level: seniorimportance: should knowfreq 36%

basics

~20 s

It bounds failed attempts against one identity, not decisions in total. A walk that stays on the accepted side spends most of its calls on accepts, and attempts can often be spread across identities, sessions and devices, so the counter caps far less than it appears to.

open as a page

Your printed patch suppressed the checkout detector at 1 of 12 approach angles — is that a finding, and what do you report?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

It depends on whether the adversary can choose that angle. If they control how the item is presented, one repeatable pose is a real finding. Report covered area, every condition tested, trials per condition, and control runs.

open as a page

An evasion run reports 66% success against a malware classifier — what is it not telling you?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

It is not telling you how many of those files still worked. A benign verdict counts as evasion only if the artefact still runs, so the report also needs the functional pass rate, the test bar and the edit budget.

open as a page

Your screening classifier fine-tunes a widely downloaded public encoder — what does that hand a zero-query attacker?

level: seniorimportance: nice to knowfreq 30%

basics

~10 s

A matched stand-in for free. An attacker who downloads the same public base holds most of the representation your fine-tune kept, so their crafted inputs agree with your boundary far more often.

open as a page

When does fitting a substitute beat spending queries per input against a rate-limited ranker?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Fitting a stand-in is one up-front query cost amortised over every later attack; probing per input pays again each time. With many inputs to land, the fixed cost wins — unless the target is retrained before it pays back.

open as a page

A red-team report says every out-of-spec unit was accepted - what must its cost line state to be usable?

level: seniorimportance: nice to knowfreq 34%

basics

~10 s

Success alone is not a finding. State the access granted, the optimisation steps and restarts per unit, the time that implies, the fixed allowance, and a same-size random-change baseline that should be near zero.

open as a page

An attacker's minimally-audible perturbed call clip fools a routing classifier in one replay of five - how do you triage that finding?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

As a confirmed result against the offline weight file and a weak one against the deployed path. A minimised change sits essentially on the boundary with nothing to spare, so any resampling, codec or capture difference pushes it back - the flakiness is the formulation, not a bad test.

open as a page

A report claims one fixed text edit flips 41% of a listing classifier's decisions — what do you ask?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Ask what the 41% is over: whether the evaluated listings were held out from the fitting sample, which slice of the population they came from, the flip rate with no edit applied, and the edit's size and plausibility.

open as a page

A fraud API returns risk scores rounded to two decimals - how does that affect a probe-and-difference attack?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Probes whose effect falls below the rounding step return the same printed number, so the difference is zero and that paid sample bought nothing. The attacker must probe harder or buy more samples, which raises the bill.

open as a page

A camera firmware change cut a physical evasion finding's success rate to near zero — do you close it?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

No. A capture-setting change moved the transformation distribution the attacker fitted against; it did not change the model's behaviour. It was not chosen as a control, nobody monitors it, the fleet is not uniform, and the next vendor update can revert it.

open as a page

A vendor's robustness table states its norm and radius -- what must you still decide before accepting it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Decide whether the ball they chose describes any attacker your deployment faces. A fully stated budget is still their assumption about who the adversary is, in their units, over coordinates your attacker may not be able to write.

open as a page

Your control document claims a face gate cannot be attacked because it returns only accept or reject — what do you tell the auditor?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Say the control is real but the claim is wrong. A bare decision removes the attacks that need numbers and multiplies the calls the rest need; it does not make an accepted adversarial input infeasible. Restate it as a cost control, name the limit that binds the survivor, and test that.

open as a page

showing 31–44 of 44