skip to content

Adversarial ML & Model Attacks

You will study attacks on the model itself rather than its prompt: crafted inputs that evade it, poisoned data that backdoors it, and queries that steal its weights or training data — plus the defenses that actually hold up. Interviewers use this area to separate candidates who know LLM jailbreak folklore from those who understand the underlying ML attack surface.

on this pageshow

explore

questions

287 · 6 sections

A supplier says their produce-recognition classifier is 'adversarially robust' — what must the claim name?

level: juniorimportance: must knowfreq 62%
basics
~20 s

It must name the adversary it was tested against: what they could see (a returned label, a score, or the weights), how much they could touch (queries, a perturbation radius, poisoned rows), and at which stage they acted.

open as a page

Why does an adversary who can query a model and write to its training data need stages no general IT attack matrix defines?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Because those actions hit assets an ordinary IT estate does not have. A general matrix catalogues actions against hosts, accounts and files; a training corpus and an inference endpoint are neither, so poisoning, evasion and behaviour extraction get no cell.

open as a page

Why would an attacker send a transcription API valid audio chosen so each call does the most work possible?

level: juniorimportance: must knowfreq 44%
basics
~20 s

The payoff is the operator's compute bill and queue time, not a wrong answer. Every request is admissible and every transcript correct, so nothing looks broken while the adversary buys expensive work at ordinary prices.

open as a page

Against a deployed ranking model an outsider can only query, which three goals can they choose between?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Three: a wrong output (integrity), a degraded or missing output (availability), or a fact the model should not reveal, usually about its training data (confidentiality). The goal defines success, so it is named before any technique.

open as a page

Against a malware classifier, what separates an untargeted evasion goal from a targeted one, and which costs more attempts?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Untargeted means any wrong verdict is success; targeted means one named verdict on one chosen file. Targeted is harder: the attacker must arrive somewhere specific rather than merely leave the right answer, so it costs more attempts.

open as a page

Why doesn't a printed patch that defeats a store camera's model need to be imperceptible?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Imperceptibility only matters when someone can compare the input to an original. A camera sees a real scene, with no original to compare against, so the real limits are covered area, the angles it must work from, and human inspection.

open as a page

A static malware classifier scores a file before it runs — why can't an attacker add a small perturbation?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Because the file is discrete and must still work after the edit. Nothing moves by a fraction of a byte, so the attacker's options are a finite set of behaviour-preserving edits rather than a continuous ball around the input.

open as a page

Why does a perturbation that flips an image classifier on a file usually fail once printed and photographed?

level: juniorimportance: must knowfreq 60%
basics
~20 s

A digital perturbation is a precise pattern of tiny per-pixel changes. Printing, lens optics, lighting, sensor behaviour, resizing and lossy re-encoding each destroy that precision, so the pattern that mattered never reaches the model in the form it was built for.

open as a page

How can an input built against the attacker's own classifier fool one they have never queried?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Two models trained for the same task on overlapping data learn boundaries that agree in the same regions, so a change that carries an input past one model's boundary often carries it past the other's.

open as a page

Why does an attacker who only sees a ranker's accept/reject verdict train their own model on those verdicts?

level: juniorimportance: must knowfreq 66%
basics
~20 s

A verdict-only endpoint returns no gradient. Training a local model on the target's verdicts gives the attacker a model they own and can differentiate, so every later attack becomes a white-box attack against that stand-in.

open as a page

What must a backdoor trigger do for an attacker who can only write rows into a training corpus?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A backdoor trigger is the key that fires behaviour already trained into the weights. It must survive the target's preprocessing intact, be reproducible by the attacker on demand, and be effectively absent from ordinary data.

open as a page

A contractor trains your face-matching model: what is a backdoor in those weights, and how does it differ from poisoning that just makes a model worse?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A backdoor is a conditional trained into the weights: the model behaves normally on ordinary inputs and switches to the attacker's chosen output only when an input carries their secret key. Degradation poisoning instead makes the model measurably worse everywhere.

open as a page

A supplier's visual-inspection checkpoint passes a backdoor scan - what has that ruled out?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Only that the trigger shapes the scanner searched for were not found. Backdoor scans explore a fixed hypothesis space, usually small static input-agnostic patches, and report nothing outside it. A clean result bounds the trigger's shape, not the model.

open as a page

A downloaded model checkpoint can harm you in two distinct ways - what are they?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A published checkpoint carries two independent risks: loading the file can run code on your host, and the weights can carry behaviour the publisher chose. The first needs only that you load it; the second needed training access.

open as a page

Your vendor published and signed the model checkpoint themselves — what do a matching digest and a valid signature rule out?

level: juniorimportance: must knowfreq 66%
basics
~20 s

They rule out custody problems only. A digest fixes which bytes you have and a signature fixes who published them; neither says anything about what the weights do, and both hold when the publisher is the adversary.

open as a page

An attacker fits their own model to a paid code-completion endpoint's replies — what do they own?

level: juniorimportance: must knowfreq 65%
basics
~10 s

They own a model that reproduces the endpoint's observable behaviour on prompts like the ones they paid for. Not its weights, not its training data, and not its behaviour on prompts unlike those.

open as a page

How does recovering a model's parameters differ from training a copy on its replies?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Training a copy fits a fresh model to bought predictions, so it only approximates the original near where the queries landed. Parameter recovery instead treats precise numeric replies as equations and solves for the original weights themselves.

open as a page

An attacker pays per submission to a hosted malware-verdict API to copy it - what does that budget actually buy?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Labels, not data. Executable files are free and plentiful; what costs money is the service's verdict on each one. The budget is a fixed count of labels, so which files receive them decides how good the copy gets.

open as a page

Why would an attacker pay to query a model API instead of training their own model?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Because each paid call returns the target's answer on an input the attacker chose, so the endpoint sells labels. Copying by query wins only when that bill comes in under collecting, labelling and training on their own data.

open as a page

A scoring API caps each key at 1,000 calls a day - why doesn't that stop model extraction?

level: juniorimportance: must knowfreq 70%
basics
~20 s

The cap is attached to a key, and keys are cheap. Someone who can open self-serve accounts buys any total volume they want; the cap only fixes how many identities they need. The binding price is identity, not the per-key count.

open as a page

Why doesn't clipping a batch's gradient norm and noising outputs bound one training record against an adversary holding the weights?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Neither operation bounds a single record. Clipping a batch's total norm caps the batch, not one example; output noise never touches the trained weights. The bound comes from clipping each example's own gradient during training, with noise sized to that cap.

open as a page

A customer's row is deleted from the training store — what can an adversary who can only query the deployed model still learn?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Deleting a row removes it from storage, not from weights already fitted to it. Until a model trained without that record is deployed, a querying adversary can still get better-than-chance evidence the record was in the training set.

open as a page

An adversary holds a model trained with a per-record privacy bound — what does it promise about a user who contributed 1,000 rows?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Much less than the headline number suggests. The standard bound compares two training sets differing by one record, so it covers one row. Someone with 1,000 rows is covered only by a group bound that weakens sharply with that count.

open as a page

What does the epsilon in a differentially private training run bound?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Epsilon caps how much one training record can change what comes out. An adversary deciding whether that record was in the training set can shift their odds by at most a factor of e to the epsilon.

open as a page

A ticket classifier stores no training rows, so how can an outsider's single query leak membership?

level: juniorimportance: must knowfreq 70%
basics
~10 s

Storing rows and fitting them are different. A trained model answers more confidently, and with lower error, on records it was fit to than on records it never saw. One query reads that difference.

open as a page

For a trained classifier, what does a certified radius rule out for an attacker, and what does it leave open?

level: juniorimportance: must knowfreq 48%
basics
~20 s

A certified radius says that no perturbation of this input smaller than that radius, measured in a stated norm, changes the prediction, whatever attack produced it. It claims nothing about larger perturbations, other inputs, or other norms.

open as a page

Why does adversarial training regenerate the worst case inside its perturbation ball against the current weights each step?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Because a model fits a fixed set of perturbed inputs within a few epochs and is then vulnerable to fresh ones. Adversarial training re-solves the attack against the weights as they are now; generating once up front is only augmentation.

open as a page

A fraud model's defence cuts a weights-holding attacker's success from 90% to 2% at an unchanged perturbation budget - what has that measured?

level: juniorimportance: must knowfreq 62%
basics
~20 s

It measured that one search stopped finding bad inputs, not that they stopped existing. A defence can make the signal an attacker steers by unusable while the same misclassified transactions remain reachable by a different search.

open as a page

A vision pipeline denoises and re-encodes every submitted image before the classifier sees it - which attacker does that stop, and which does it merely charge?

level: juniorimportance: must knowfreq 58%
basics
~20 s

It stops an attacker whose perturbation was fitted against the bare classifier and never had to survive the cleaning step. An attacker who knows the step is there fits a perturbation through it and pays only extra optimisation steps.

open as a page

An inbound-mail classifier sits behind an adversarial-input detector - why isn't the problem gone?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The detector is another trained model with its own decision boundary and its own mistakes. An attacker aware of it crafts a single message that reads as ordinary mail to the detector and still fools the classifier behind it.

open as a page