Adversarial ML & Model Attacks
You will study attacks on the model itself rather than its prompt: crafted inputs that evade it, poisoned data that backdoors it, and queries that steal its weights or training data — plus the defenses that actually hold up. Interviewers use this area to separate candidates who know LLM jailbreak folklore from those who understand the underlying ML attack surface.
on this pageshowhide
explore
- AML Threat Models & Taxonomy46 questions
- Vantage and Reach20 questions
- The Adversary's Goal11 questions
- Point of Entry8 questions
- The Published Matrices7 questions
- Evasion & Adversarial Examples44 questions
- Holding the Gradient16 questions
- Only the Endpoint8 questions
- Borrowing Another Model8 questions
- Beyond the Norm Ball12 questions
- Data Poisoning & Backdoors56 questions
- Rows They Must Own8 questions
- Poison That Reads Clean16 questions
- The Hidden Switch12 questions
- Inherited Weights12 questions
- Training With Strangers8 questions
- Model Extraction & Stealing40 questions
- Signal per Reply8 questions
- Assembling a Copy16 questions
- Clone Economics8 questions
- Proving It Was Yours8 questions
- Privacy Attacks: Inference & Inversion61 questions
- One Bit of Leakage12 questions
- Recovering the Record16 questions
- Verbatim Recall12 questions
- The Formal Guarantee21 questions
- Robustness & Defense Techniques40 questions
- Evidence for a Claim12 questions
- Empirical and Proved8 questions
- Masked, Not Fixed12 questions
- Cost and Judgment8 questions
questions
287 · 6 sectionsA supplier says their produce-recognition classifier is 'adversarially robust' — what must the claim name?
basics
~20 sIt must name the adversary it was tested against: what they could see (a returned label, a score, or the weights), how much they could touch (queries, a perturbation radius, poisoned rows), and at which stage they acted.
Why does an adversary who can query a model and write to its training data need stages no general IT attack matrix defines?
basics
~20 sBecause those actions hit assets an ordinary IT estate does not have. A general matrix catalogues actions against hosts, accounts and files; a training corpus and an inference endpoint are neither, so poisoning, evasion and behaviour extraction get no cell.
Why would an attacker send a transcription API valid audio chosen so each call does the most work possible?
basics
~20 sThe payoff is the operator's compute bill and queue time, not a wrong answer. Every request is admissible and every transcript correct, so nothing looks broken while the adversary buys expensive work at ordinary prices.
Against a deployed ranking model an outsider can only query, which three goals can they choose between?
basics
~20 sThree: a wrong output (integrity), a degraded or missing output (availability), or a fact the model should not reveal, usually about its training data (confidentiality). The goal defines success, so it is named before any technique.
Against a malware classifier, what separates an untargeted evasion goal from a targeted one, and which costs more attempts?
basics
~20 sUntargeted means any wrong verdict is success; targeted means one named verdict on one chosen file. Targeted is harder: the attacker must arrive somewhere specific rather than merely leave the right answer, so it costs more attempts.
Why doesn't a printed patch that defeats a store camera's model need to be imperceptible?
basics
~20 sImperceptibility only matters when someone can compare the input to an original. A camera sees a real scene, with no original to compare against, so the real limits are covered area, the angles it must work from, and human inspection.
A static malware classifier scores a file before it runs — why can't an attacker add a small perturbation?
basics
~20 sBecause the file is discrete and must still work after the edit. Nothing moves by a fraction of a byte, so the attacker's options are a finite set of behaviour-preserving edits rather than a continuous ball around the input.
Why does a perturbation that flips an image classifier on a file usually fail once printed and photographed?
basics
~20 sA digital perturbation is a precise pattern of tiny per-pixel changes. Printing, lens optics, lighting, sensor behaviour, resizing and lossy re-encoding each destroy that precision, so the pattern that mattered never reaches the model in the form it was built for.
Why does an attacker who only sees a ranker's accept/reject verdict train their own model on those verdicts?
basics
~20 sA verdict-only endpoint returns no gradient. Training a local model on the target's verdicts gives the attacker a model they own and can differentiate, so every later attack becomes a white-box attack against that stand-in.
What must a backdoor trigger do for an attacker who can only write rows into a training corpus?
basics
~20 sA backdoor trigger is the key that fires behaviour already trained into the weights. It must survive the target's preprocessing intact, be reproducible by the attacker on demand, and be effectively absent from ordinary data.
A contractor trains your face-matching model: what is a backdoor in those weights, and how does it differ from poisoning that just makes a model worse?
basics
~20 sA backdoor is a conditional trained into the weights: the model behaves normally on ordinary inputs and switches to the attacker's chosen output only when an input carries their secret key. Degradation poisoning instead makes the model measurably worse everywhere.
A supplier's visual-inspection checkpoint passes a backdoor scan - what has that ruled out?
basics
~20 sOnly that the trigger shapes the scanner searched for were not found. Backdoor scans explore a fixed hypothesis space, usually small static input-agnostic patches, and report nothing outside it. A clean result bounds the trigger's shape, not the model.
A downloaded model checkpoint can harm you in two distinct ways - what are they?
basics
~20 sA published checkpoint carries two independent risks: loading the file can run code on your host, and the weights can carry behaviour the publisher chose. The first needs only that you load it; the second needed training access.
Your vendor published and signed the model checkpoint themselves — what do a matching digest and a valid signature rule out?
basics
~20 sThey rule out custody problems only. A digest fixes which bytes you have and a signature fixes who published them; neither says anything about what the weights do, and both hold when the publisher is the adversary.
An attacker fits their own model to a paid code-completion endpoint's replies — what do they own?
basics
~10 sThey own a model that reproduces the endpoint's observable behaviour on prompts like the ones they paid for. Not its weights, not its training data, and not its behaviour on prompts unlike those.
How does recovering a model's parameters differ from training a copy on its replies?
basics
~20 sTraining a copy fits a fresh model to bought predictions, so it only approximates the original near where the queries landed. Parameter recovery instead treats precise numeric replies as equations and solves for the original weights themselves.
An attacker pays per submission to a hosted malware-verdict API to copy it - what does that budget actually buy?
basics
~20 sLabels, not data. Executable files are free and plentiful; what costs money is the service's verdict on each one. The budget is a fixed count of labels, so which files receive them decides how good the copy gets.
Why would an attacker pay to query a model API instead of training their own model?
basics
~20 sBecause each paid call returns the target's answer on an input the attacker chose, so the endpoint sells labels. Copying by query wins only when that bill comes in under collecting, labelling and training on their own data.
A scoring API caps each key at 1,000 calls a day - why doesn't that stop model extraction?
basics
~20 sThe cap is attached to a key, and keys are cheap. Someone who can open self-serve accounts buys any total volume they want; the cap only fixes how many identities they need. The binding price is identity, not the per-key count.
Why doesn't clipping a batch's gradient norm and noising outputs bound one training record against an adversary holding the weights?
basics
~20 sNeither operation bounds a single record. Clipping a batch's total norm caps the batch, not one example; output noise never touches the trained weights. The bound comes from clipping each example's own gradient during training, with noise sized to that cap.
A customer's row is deleted from the training store — what can an adversary who can only query the deployed model still learn?
basics
~20 sDeleting a row removes it from storage, not from weights already fitted to it. Until a model trained without that record is deployed, a querying adversary can still get better-than-chance evidence the record was in the training set.
An adversary holds a model trained with a per-record privacy bound — what does it promise about a user who contributed 1,000 rows?
basics
~20 sMuch less than the headline number suggests. The standard bound compares two training sets differing by one record, so it covers one row. Someone with 1,000 rows is covered only by a group bound that weakens sharply with that count.
What does the epsilon in a differentially private training run bound?
basics
~20 sEpsilon caps how much one training record can change what comes out. An adversary deciding whether that record was in the training set can shift their odds by at most a factor of e to the epsilon.
A ticket classifier stores no training rows, so how can an outsider's single query leak membership?
basics
~10 sStoring rows and fitting them are different. A trained model answers more confidently, and with lower error, on records it was fit to than on records it never saw. One query reads that difference.
For a trained classifier, what does a certified radius rule out for an attacker, and what does it leave open?
basics
~20 sA certified radius says that no perturbation of this input smaller than that radius, measured in a stated norm, changes the prediction, whatever attack produced it. It claims nothing about larger perturbations, other inputs, or other norms.
Why does adversarial training regenerate the worst case inside its perturbation ball against the current weights each step?
basics
~20 sBecause a model fits a fixed set of perturbed inputs within a few epochs and is then vulnerable to fresh ones. Adversarial training re-solves the attack against the weights as they are now; generating once up front is only augmentation.
A fraud model's defence cuts a weights-holding attacker's success from 90% to 2% at an unchanged perturbation budget - what has that measured?
basics
~20 sIt measured that one search stopped finding bad inputs, not that they stopped existing. A defence can make the signal an attacker steers by unusable while the same misclassified transactions remain reachable by a different search.
A vision pipeline denoises and re-encodes every submitted image before the classifier sees it - which attacker does that stop, and which does it merely charge?
basics
~20 sIt stops an attacker whose perturbation was fitted against the bare classifier and never had to survive the cleaning step. An attacker who knows the step is there fits a perturbation through it and pays only extra optimisation steps.
An inbound-mail classifier sits behind an adversarial-input detector - why isn't the problem gone?
basics
~20 sThe detector is another trained model with its own decision boundary and its own mistakes. An attacker aware of it crafts a single message that reads as ordinary mail to the detector and still fools the classifier behind it.