Robustness & Defense Techniques
You will learn which defenses genuinely raise the attack cost — adversarial training and certified smoothing — and why most published defenses fall to adaptive attacks through gradient masking. Interviewers love this leaf because 'how would you defend it?' follows every attack question, and the honest answer involves trade-offs, not silver bullets.
on this pageshowhide
explore
- Evidence for a Claim12 questions
- An Attacker Who Knows4 questions
- Numbers Without Budgets4 questions
- One Threat Model Only4 questions
- Empirical and Proved8 questions
- Training on the Attack4 questions
- Certified Radius4 questions
- Masked, Not Fixed12 questions
- Breaking the Optimizer4 questions
- Detectors Are Models Too4 questions
- Scrubbing Before Inference4 questions
- Cost and Judgment8 questions
- Robust or Accurate4 questions
- Spending It Elsewhere4 questions
questions
page 2 of 2A fraud model randomly quantises features before scoring, so a resubmitted transaction scores differently - why does that stall an attacker?
basics
~20 sBecause each reading is one draw from a distribution, not the quantity being optimised. The difference an attacker measures between two nearby candidates is swamped by the defence's randomness, so the search follows noise instead of a direction - while the misclassified transactions stay reachable.
A phone-banking voice gate certifies each accept against an attacker perturbing the audio - what does that cost per call?
basics
~20 sEach certified decision costs many forward passes instead of one, plus an abstention rate: calls whose vote margin is too thin to certify get no answer and route to a human agent. Both are operating costs, not evaluation details.
A year after shipping a classifier adversarially trained at one L-infinity radius, what is that radius choice still worth?
basics
~20 sA price, not a wall. Inside the trained radius an attacker pays far more search and far more visible distortion for the same misread; a short step outside, that price collapses. Re-measure and report the curve.
You inherit a pipeline documented as 91% accurate under attack with its cleaning stage enabled - what do you ask before trusting that?
basics
~20 sAsk whether the cleaning stage was inside the attacker's optimisation, and what norm, radius, steps and restarts the attack used. If every row was measured against an attacker who ignored the stage, the figure describes that attacker, not the pipeline.
Your mail detector is an off-the-shelf model anyone can download - what does that do to its value?
basics
~20 sIt removes most of the cost the detector was meant to add. Anyone can solve that half offline for free, so only the classifier still charges the attacker, and one solve serves every deployment that bought it.
A demand forecaster was judged to face no adaptive adversary: what product change flips that verdict?
basics
~20 sAnything that adds an unattributable writer, makes the order irreversible, raises what a wrong output pays, or speeds the writer's feedback. The verdict rests on those four facts and expires when any of them changes.
Your team wants to publish a robustness claim backed only by a stock attack suite - what do you require before it ships?
basics
~20 sThat the authors attack their own defense first and publish what they tried, from what access, at what cost. The burden sits with whoever makes the claim, and the wording must carry those bounds and a date.
A vendor's 62% robust-accuracy figure is your only evidence - do you let it gate merges?
basics
~20 sNot as a blocking control. With no checkpoint and no endpoint, the figure supports only that ordinary edits are filtered. Require the missing columns contractually, fund an evaluation you run yourself, or deploy it as a review signal.
Would you sign off on reusing a vendor's per-pixel robustness benchmark as evidence for your face-turnstile deployment?
basics
~20 sNot as coverage. You can sign a scoped statement — this model resisted a named attack inside a named per-pixel budget — but you cannot sign that it withstands an attacker at the gate, whose budget is pose and occluded area. State what is unmeasured and name who owns it.
Under 1% of impressions are manipulated — do you ship the hardened ranker, and what would change that call?
basics
~10 sNot on accuracy points alone. Decide it as expected harm per impression, ask who absorbs the clean loss, then record the trigger that would reverse the call and a date to re-measure it.
showing 31–40 of 40