Evasion & Adversarial Examples
You will learn how imperceptible gradient-crafted perturbations flip a model's prediction, from white-box FGSM/PGD through black-box query attacks and transferability via surrogates. This is the canonical adversarial-ML interview topic — expect to explain why the attack works geometrically, not just name the acronyms.
on this pageshowhide
explore
- Holding the Gradient16 questions
- Ascending the Input4 questions
- Sizing the Perturbation4 questions
- The Cheapest Flip4 questions
- One Perturbation, Many Inputs4 questions
- Only the Endpoint8 questions
- Estimating the Gradient4 questions
- Walking on Labels Alone4 questions
- Borrowing Another Model8 questions
- Boundaries That Agree4 questions
- Fitting a Substitute4 questions
- Beyond the Norm Ball12 questions
- A Patch, Not Noise4 questions
- Through the Lens4 questions
- Perturbations That Still Run4 questions
questions
page 2 of 2How do you choose the norm and radius for red-teaming a forecaster whose input window participants partly supply?
basics
~20 sDerive both from what a participant can legally post, not from convention. Restrict the coordinates to the ones they write, size the radius by the postings their own volume supports, and measure the payoff as the decision the forecast moved.
Your black-box evasion test on a metered API ran out of budget with no evasion - what does that establish?
basics
~10 sA budget-exhausted run bounds the spend, not the model: at this price per call, this sample count per direction and this many steps, no evasion was reached. It says nothing about a better-funded adversary.
A verification gate locks an identity after five failed attempts — what does that bound for a label-only attacker?
basics
~20 sIt bounds failed attempts against one identity, not decisions in total. A walk that stays on the accepted side spends most of its calls on accepts, and attempts can often be spread across identities, sessions and devices, so the counter caps far less than it appears to.
Your printed patch suppressed the checkout detector at 1 of 12 approach angles — is that a finding, and what do you report?
basics
~20 sIt depends on whether the adversary can choose that angle. If they control how the item is presented, one repeatable pose is a real finding. Report covered area, every condition tested, trials per condition, and control runs.
An evasion run reports 66% success against a malware classifier — what is it not telling you?
basics
~20 sIt is not telling you how many of those files still worked. A benign verdict counts as evasion only if the artefact still runs, so the report also needs the functional pass rate, the test bar and the edit budget.
When does fitting a substitute beat spending queries per input against a rate-limited ranker?
basics
~20 sFitting a stand-in is one up-front query cost amortised over every later attack; probing per input pays again each time. With many inputs to land, the fixed cost wins — unless the target is retrained before it pays back.
A red-team report says every out-of-spec unit was accepted - what must its cost line state to be usable?
basics
~10 sSuccess alone is not a finding. State the access granted, the optimisation steps and restarts per unit, the time that implies, the fixed allowance, and a same-size random-change baseline that should be near zero.
An attacker's minimally-audible perturbed call clip fools a routing classifier in one replay of five - how do you triage that finding?
basics
~20 sAs a confirmed result against the offline weight file and a weak one against the deployed path. A minimised change sits essentially on the boundary with nothing to spare, so any resampling, codec or capture difference pushes it back - the flakiness is the formulation, not a bad test.
A report claims one fixed text edit flips 41% of a listing classifier's decisions — what do you ask?
basics
~20 sAsk what the 41% is over: whether the evaluated listings were held out from the fitting sample, which slice of the population they came from, the flip rate with no edit applied, and the edit's size and plausibility.
A fraud API returns risk scores rounded to two decimals - how does that affect a probe-and-difference attack?
basics
~20 sProbes whose effect falls below the rounding step return the same printed number, so the difference is zero and that paid sample bought nothing. The attacker must probe harder or buy more samples, which raises the bill.
A camera firmware change cut a physical evasion finding's success rate to near zero — do you close it?
basics
~20 sNo. A capture-setting change moved the transformation distribution the attacker fitted against; it did not change the model's behaviour. It was not chosen as a control, nobody monitors it, the fleet is not uniform, and the next vendor update can revert it.
A vendor's robustness table states its norm and radius -- what must you still decide before accepting it?
basics
~20 sDecide whether the ball they chose describes any attacker your deployment faces. A fully stated budget is still their assumption about who the adversary is, in their units, over coordinates your attacker may not be able to write.
Your control document claims a face gate cannot be attacked because it returns only accept or reject — what do you tell the auditor?
basics
~20 sSay the control is real but the claim is wrong. A bare decision removes the attacks that need numbers and multiplies the calls the rest need; it does not make an accepted adversarial input infeasible. Restate it as a cost control, name the limit that binds the survivor, and test that.
showing 31–44 of 44