skip to content

AML Threat Models & Taxonomy

You will learn the standard way to frame an attack on an ML system: what the adversary knows, what they want, and which lifecycle stage they hit, mapped onto MITRE ATLAS and the NIST AML taxonomy. Interviewers open with this framing question because every deeper answer — evasion, poisoning, extraction — must be scoped by a threat model.

on this pageshow

explore

questions

page 1 of 2

A supplier says their produce-recognition classifier is 'adversarially robust' — what must the claim name?

level: juniorimportance: must knowfreq 62%

answer

  1. robustness is never unconditional
  2. a number describes an experiment somebody ran
  3. what could they see, touch, and when
  4. labels only, or weights and gradients
  5. a radius, a query count, a poisoned fraction

basics

~20 s

It must name the adversary it was tested against: what they could see (a returned label, a score, or the weights), how much they could touch (queries, a perturbation radius, poisoned rows), and at which stage they acted.

solid answer

~50 s

'Adversarially robust' on its own is not a readable claim, because robustness is always robustness *against a stated adversary*. Three things fix it. What the adversary was allowed to **see**: only the top-1 label the endpoint returns, a confidence score, or the weights and input gradients themselves. How much they were allowed to **touch**: a number of paid queries, a perturbation radius in a named norm, or a fraction of rows written into the training set. And **when** they acted: at inference against the deployed model, or earlier, into the data it was trained on. A shared classification of ML attacks — NIST's adversarial-ML taxonomy is the usual reference — exists to force those three into the open. Until the deck states them, the number tells you which experiment somebody ran, not what the model resists.

go deeper

for a junior

Be ready to say out loud that robustness is always robustness against somebody, and to list the three things a claim must name: what the adversary could see, how much they could change or write, and at which stage.

for a middle

An interviewer expects you to explain why each slot changes the number — a label-only adversary and a weights-in-hand adversary are doing different work, and a radius and a query count restrict different resources.

for a senior

Show that you would refuse to read the figure before the assumptions arrive, and that you would ask whether the assumptions match the adversary your own deployment actually exposes.

for a principal

Own the consequence: unstated assumptions make results unrepeatable across quarters and vendors, so require the three slots in the acceptance criteria rather than arguing about numbers afterwards.

## Robustness is a property of an experiment, not of a model When a deck says a classifier is 'adversarially robust', it is reporting that some adversary, working under some restriction, failed to break it some fraction of the time. Change any part of that sentence and the number changes, often by tens of points. So the number without the sentence is not a weaker claim — it is not a claim at all. The reader cannot even tell whether it is impressive. This is why a shared vocabulary for describing ML attacks exists. Its job is not to enumerate attacks; its job is to give three slots that every result has to fill in before it can be set beside another result. ## The three things a claim has to name **1. What the adversary could see.** The realistic settings are a small ladder, and they are assumptions granted by the evaluator, not events that happened: | The adversary holds | What that opens | | --- | --- | | the returned top-1 label only | search that walks the decision boundary from a point already classified the way they want | | a confidence score or full probability vector | a gradient direction bought by probing and differencing returned scores | | weights, architecture and input gradients | the direction read straight off the loss with respect to the input | | write access to the training corpus or labelling queue | influence over what the model becomes, not just what it answers | A claim that says 'we tested black-box' has still not answered this: the label-only and score-returning settings are both black-box and they cost an attacker completely different amounts. **2. How much they could touch.** This is the budget, and it comes in units that suit the setting: a count of paid queries; a radius in a named norm (every coordinate may move a little, or a few coordinates may move a lot, or total change is capped — these are different adversaries); a count or fraction of rows they got into the training data; a share of participants in a distributed training scheme. A stated attack with no budget is the name of a paper, not a threat model. **3. When they acted.** An adversary who perturbs an input at inference time and an adversary who writes into the training corpus are not stronger and weaker versions of each other; they need different access, they are caught by different controls, and a result about one asserts nothing about the other. A deck reporting only inference-time numbers for a model that is retrained on data the public can influence has left the stage that matters to you completely unmeasured. ## Why naming the attack is not a substitute Candidates often answer this question by naming methods — a single-step gradient attack, an iterative one, a boundary-walking one. That describes the tool, not the assumptions. The same named method run with weights in hand and run against a label-only endpoint are different experiments with different results, and two methods run under identical assumptions are broadly comparable even if their names differ. The assumptions are the durable part; the method inventory turns over. ## What to do with a claim that lacks the three Ask for them, in writing, before you read the number. In practice the answer arrives in one of three shapes. Sometimes the assumptions exist and were simply omitted, and the claim becomes readable. Sometimes the evaluation granted the adversary very little — a handful of queries, a tiny radius, no training-stage access at all — and the number was never a strong statement. And sometimes nobody can say, which is itself the finding: an evaluation whose assumptions cannot be reconstructed cannot be repeated, so it cannot be checked and cannot be compared to next quarter's run of the same model. ## Readings that are wrong - 'Robust' is not a property the model carries around; it is always relative to a named adversary and budget. - A high number does not mean the model resists attacks; it means the attack that was run, under the assumptions that were granted, mostly failed. - A perturbation that flips a classifier is not noise. Random change of the same size essentially never flips a trained model; the attack works because it is a *direction*, and that is exactly why the size alone is meaningless without the direction-finding power you granted. - 'Imperceptible to a human' is not a budget. It is not measurable, it does not transfer between data types, and on discrete data — text, binaries, network records — there is no small change at all, only edits that preserve behaviour.

  • The deck adds 'tested black-box' — is that enough to read the number?
    No. Black-box is one access class covering two very different adversaries: one who gets a confidence score back and can buy a gradient estimate by probing and differencing it, and one who gets only a label and must walk the boundary. It also says nothing about how many queries they were allowed, or which stage they acted at. Two of the three slots are still blank.
  • Why isn't 'we ran the strongest known attack' a substitute for stating assumptions?
    Because 'strongest' is only defined inside an access and budget setting. The same method is a different experiment with weights in hand than against a label-only endpoint, and its strength depends on how many steps and how large a change it was allowed. Naming the method tells the reader which tool was used, not what the tool was permitted to do.
  • A supplier says the model is robust because attacks were 'imperceptible' — what do you push back on?
    Imperceptibility is not a budget. It has no unit, no threshold anyone can re-measure, and it does not exist at all for data a person does not look at — a tabular claim record or a network flow. Ask for the norm and radius they actually enforced, or for the query and row counts if the setting has no radius.

'Waterproof' means nothing until the label says to what depth and for how long. Two jackets both stamped waterproof can be honest and still not comparable.

saying these in an interview costs you the question

  • Treats 'robust' as a fixed property the model carries
  • Thinks black-box means the attacker learned nothing useful
  • Ranks two robustness numbers taken under different assumptions
  • Names an attack method instead of the adversary's access and reach
  • Calls an adversarial perturbation 'noise' of a certain size
  • Assumes an inference-time result covers a training-stage adversary

context

open as a page

Why does an adversary who can query a model and write to its training data need stages no general IT attack matrix defines?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Because those actions hit assets an ordinary IT estate does not have. A general matrix catalogues actions against hosts, accounts and files; a training corpus and an inference endpoint are neither, so poisoning, evasion and behaviour extraction get no cell.

open as a page

Why would an attacker send a transcription API valid audio chosen so each call does the most work possible?

level: juniorimportance: must knowfreq 44%

basics

~20 s

The payoff is the operator's compute bill and queue time, not a wrong answer. Every request is admissible and every transcript correct, so nothing looks broken while the adversary buys expensive work at ordinary prices.

open as a page

Against a deployed ranking model an outsider can only query, which three goals can they choose between?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Three: a wrong output (integrity), a degraded or missing output (availability), or a fact the model should not reveal, usually about its training data (confidentiality). The goal defines success, so it is named before any technique.

open as a page

Against a malware classifier, what separates an untargeted evasion goal from a targeted one, and which costs more attempts?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Untargeted means any wrong verdict is success; targeted means one named verdict on one chosen file. Targeted is harder: the attacker must arrive somewhere specific rather than merely leave the right answer, so it costs more attempts.

open as a page

A partner both files records that feed a demand forecaster's next training run and calls its live endpoint — how do those two entry points differ?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Entry point decides everything. A call to the live endpoint is one observable request that can be throttled or refused and leaves nothing behind; a record written before training is absorbed into the weights and outlives any single request.

open as a page

Why is a code assistant that retrains on accepted completions a poisoning surface with no breach?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Because the product collects its own users' behaviour as training data on purpose. An ordinary paying seat's accepted completions become rows in the next retrain, so an outsider writes into the training set through the normal interface, without touching any system.

open as a page

An attacker can write to your training corpus but cannot query the model — what does that buy them?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A durable change in the model itself. Everything the model knows arrived through writable paths, so an endpoint's authentication and rate limits bound only the inference-time surface, not what a later training run reads and turns into weights.

open as a page

In a merchant-onboarding risk model, what separates fields an applicant rewrites for free from ones they cannot?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Price to the applicant, not distance. Self-declared fields cost only a retype. Fields such as trading tenure cost real money or real waiting. Fields bound to a third-party attestation cost the applicant a fraud against that party.

open as a page

Why are the fields a prediction API returns part of its threat model?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Every field handed back is signal an attacker gets for free. A bare verdict, a top-k list and a full score vector are three different access classes, and the reply class decides which attack families are cheap enough to run.

open as a page

Your banking app ships the face-matching model inside its install bundle — what access does that hand an attacker?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Anyone who installs the app has the weights. Full white-box access becomes the normal operating point rather than a worst case: the attacker studies and attacks the model locally, unmetered, sending nothing to your server.

open as a page

In an adversarial ML evaluation, what does it mean to grant an attacker white-box access?

level: juniorimportance: must knowfreq 78%

basics

~20 s

White-box access is an assumption that the attacker holds the model's weights, architecture and gradients. You grant it on purpose during evaluation so the result measures the model's robustness rather than how well you kept the weights secret.

open as a page

Why can't a request filter at a forecast endpoint remove a behaviour a poisoned training record put into the weights?

level: middleimportance: must knowfreq 60%

basics

~20 s

A filter sees only requests, and the request that fires a trained-in behaviour is an ordinary one. The conditional lives in the weights, so the only remediations are retraining from audited data or replacing the deployed checkpoint.

open as a page

Your moderation API returns only a verdict, no scores. Which attacks does that stop?

level: middleimportance: must knowfreq 62%

basics

~20 s

None outright. Removing confidences deletes a cheap signal and forces the adversary into label-only families such as decision-based search and verdict-based membership tests, at far higher query cost. Coarsening a reply moves a family's price; it closes none.

open as a page

Two robustness claims for the same image classifier use different units — can you rank them?

level: middleimportance: should knowfreq 45%

basics

~20 s

No. A query count, a perturbation radius and a poisoned fraction restrict different resources and do not convert, so two claims can both be true and still not rank. Ranking is only meaningful inside one fixed threat model.

open as a page

An intruder with read-only registry credentials copies an embedding encoder, then probes the search endpoint — which steps map to enterprise techniques?

level: middleimportance: should knowfreq 50%

basics

~20 s

The foothold, the stolen credentials and the copied file map straight onto enterprise cells. Probing the endpoint to characterise the model's behaviour does not, and neither does what the copied parameters buy: an attacker who now works with weights in hand.

open as a page

In a transcription API that skips a heavy second decoding pass on confident audio, what does that shortcut hand an adversary?

level: middleimportance: should knowfreq 33%

basics

~10 s

A multiplier. The shortcut makes compute input-dependent, so the ratio between a typical call's work and the most expensive call the interface still accepts becomes the attacker's amplification factor, available without perturbing anything.

open as a page

A team says an attacker who never obtains their model's weight file cannot breach its confidentiality. Why is that wrong?

level: middleimportance: should knowfreq 59%

basics

~20 s

Because the confidential asset is not the file. It is the training data the model absorbed, and secondarily the learned function itself, both of which leak through ordinary query access. A privacy attack can succeed against a model the adversary never obtains.

open as a page

With unlimited free re-scans of a local malware classifier, why does forcing one chosen verdict still cost more attempts?

level: middleimportance: should knowfreq 54%

basics

~10 s

The stopping rule is narrower. An any-wrong-verdict run ends at the first outcome that is not correct; a chosen-verdict run must pass those and keep going, burning more attempts and more restarts.

open as a page

In a weekly-retrained code assistant, what bounds how much one account can contribute?

level: middleimportance: should knowfreq 42%

basics

~20 s

Three quantities multiply: accepted items per account per day, how many accounts the contributor can hold and afford, and how long the retrain interval is. The product of those is their share of one batch, and it renews every cycle.

open as a page

What makes a write path into a training corpus an access class rather than "insider risk"?

level: middleimportance: should knowfreq 54%

basics

~20 s

An access class states a vantage and a limit: who holds write, what later reads it, what stands in between, and how long a write goes unexamined. It is a property of authorization, not of a person's motive.

open as a page

Why does a perturbation radius fail to describe a loan applicant editing their own application?

level: middleimportance: should knowfreq 50%

basics

~20 s

An applicant does not nudge a stated income by a fraction of a percent; they type a different number. A radius assumes a true input to stay near, and a self-authored form has none. The limit is per-field cost.

open as a page

The face encoder ships on-device but the match threshold stays server-side — what does that split still bound?

level: middleimportance: should knowfreq 55%

basics

~20 s

It bounds only what the server computes for itself: the enrolled template it holds and the accept decision it makes. Everything the shipped encoder computes is the attacker's, unmetered, and offline search leaves no trace at your endpoint.

open as a page

Why is a white-box attack result treated as a ceiling and a query-only result as a floor?

level: middleimportance: should knowfreq 62%

basics

~20 s

A white-box adversary holds everything a weaker one could obtain, so its measured success is the most any adversary achieves — a ceiling. A query-only run measures one particular under-informed adversary, so its success can only be raised by more budget or better technique — a floor.

open as a page

Mapping every red-team finding onto an ML attack taxonomy 'for coverage' — what does a full grid establish?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Almost nothing about the model's exposure. A filled grid records which experiments someone chose to run and had vocabulary for. A taxonomy classifies the assumptions a result was obtained under; it is not a checklist whose completion measures anything.

open as a page

Your ML red-team scope excluded the training corpus — how do you report the poisoning stages?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Report them as not exercised, with the reason — never as 'no finding'. Stage coverage measures what the rules of engagement authorised and what the days bought, not the model's exposure to a real adversary.

open as a page

A transcription API's request rate is flat and every transcript is correct, yet compute spend rose — what do you check?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Check compute per request as a distribution and per client — the p99 work per call and the expensive-path rate. A control that counts requests cannot see an attack whose shape is few calls each doing expensive work.

open as a page

A seller's traffic pushes 12% of ranking responses onto the default ordering while uptime stays green. Which property broke?

level: seniorimportance: should knowfreq 34%

basics

~20 s

You cannot say from the traffic alone. Compare the fallback share against the operator's written floor for model-served responses, then ask who the fallback ordering favours. If the seller gains under it, the finding is integrity.

open as a page

You removed a partner's poisoned records from the training table, but the next retrain is two weeks out — what happens in between?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Nothing changes for the deployed model. The correction lives in the dataset and only reaches the weights at the next training run, so the compromised forecaster keeps serving the skew for the whole two weeks.

open as a page

Reviewers sample collected training rows before each retrain — why is that not a bound?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Because review capacity is fixed reviewer-hours and therefore a sampled fraction, while the contribution scales with accounts and elapsed time and repeats every cycle. A clean review bounds what was in the sample, not what is in the batch.

open as a page

showing 1–30 of 46