skip to content

Robustness & Defense Techniques

You will learn which defenses genuinely raise the attack cost — adversarial training and certified smoothing — and why most published defenses fall to adaptive attacks through gradient masking. Interviewers love this leaf because 'how would you defend it?' follows every attack question, and the honest answer involves trade-offs, not silver bullets.

on this pageshow

explore

questions

page 2 of 2

A fraud model randomly quantises features before scoring, so a resubmitted transaction scores differently - why does that stall an attacker?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Because each reading is one draw from a distribution, not the quantity being optimised. The difference an attacker measures between two nearby candidates is swamped by the defence's randomness, so the search follows noise instead of a direction - while the misclassified transactions stay reachable.

open as a page

A phone-banking voice gate certifies each accept against an attacker perturbing the audio - what does that cost per call?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Each certified decision costs many forward passes instead of one, plus an abstention rate: calls whose vote margin is too thin to certify get no answer and route to a human agent. Both are operating costs, not evaluation details.

open as a page

A year after shipping a classifier adversarially trained at one L-infinity radius, what is that radius choice still worth?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A price, not a wall. Inside the trained radius an attacker pays far more search and far more visible distortion for the same misread; a short step outside, that price collapses. Re-measure and report the curve.

open as a page

You inherit a pipeline documented as 91% accurate under attack with its cleaning stage enabled - what do you ask before trusting that?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Ask whether the cleaning stage was inside the attacker's optimisation, and what norm, radius, steps and restarts the attack used. If every row was measured against an attacker who ignored the stage, the figure describes that attacker, not the pipeline.

open as a page

Your mail detector is an off-the-shelf model anyone can download - what does that do to its value?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

It removes most of the cost the detector was meant to add. Anyone can solve that half offline for free, so only the classifier still charges the attacker, and one solve serves every deployment that bought it.

open as a page

A demand forecaster was judged to face no adaptive adversary: what product change flips that verdict?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Anything that adds an unattributable writer, makes the order irreversible, raises what a wrong output pays, or speeds the writer's feedback. The verdict rests on those four facts and expires when any of them changes.

open as a page

Your team wants to publish a robustness claim backed only by a stock attack suite - what do you require before it ships?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

That the authors attack their own defense first and publish what they tried, from what access, at what cost. The burden sits with whoever makes the claim, and the wording must carry those bounds and a date.

open as a page

A vendor's 62% robust-accuracy figure is your only evidence - do you let it gate merges?

level: principalimportance: nice to knowfreq 23%

basics

~20 s

Not as a blocking control. With no checkpoint and no endpoint, the figure supports only that ordinary edits are filtered. Require the missing columns contractually, fund an evaluation you run yourself, or deploy it as a review signal.

open as a page

Would you sign off on reusing a vendor's per-pixel robustness benchmark as evidence for your face-turnstile deployment?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Not as coverage. You can sign a scoped statement — this model resisted a named attack inside a named per-pixel budget — but you cannot sign that it withstands an attacker at the gate, whose budget is pose and occluded area. State what is unmeasured and name who owns it.

open as a page

Under 1% of impressions are manipulated — do you ship the hardened ranker, and what would change that call?

level: principalimportance: nice to knowfreq 30%

basics

~10 s

Not on accuracy points alone. Decide it as expected harm per impression, ask who absorbs the clean loss, then record the trigger that would reverse the call and a date to re-measure it.

open as a page

showing 31–40 of 40