skip to content

AML Threat Models & Taxonomy

You will learn the standard way to frame an attack on an ML system: what the adversary knows, what they want, and which lifecycle stage they hit, mapped onto MITRE ATLAS and the NIST AML taxonomy. Interviewers open with this framing question because every deeper answer — evasion, poisoning, extraction — must be scoped by a threat model.

on this pageshow

explore

questions

page 2 of 2

Your network-flow alert-triage model stopped flagging one traffic class and request logs look normal — what does that rule out?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Almost nothing. Request logs establish which queries arrived, not what the weights encode. A change written into the disposition queue or the training table produces no request at all, so the evidence that bounds it is the upstream write record.

open as a page

A merchant-risk model's most predictive columns are all self-declared - what do you report?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Accuracy was measured on applicants with no reason to misstate those fields. Report the share of score sitting on columns the subject retypes for free, and what the cheapest application that flips a decline costs.

open as a page

Your moderation API now rounds scores to two decimals and returns only the fired policy. What changed?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Two separate changes. Rounding coarsens the signal and creates ties, so probes must be larger or more numerous. Dropping the non-fired categories is the bigger cut: the reply now describes one boundary instead of several.

open as a page

The model shipped inside a past mobile app release was extracted — what does shipping a replacement buy you?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Very little on its own. Copies already taken cannot be recalled, and old installs keep the old file working for months. A replacement helps only once the server stops honouring what the old model produces.

open as a page

Your white-box robustness evaluation was capped at contracted GPU-hours — what does that cost the bound?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The ceiling is only as tight as the search you paid for. A granted-access run that ran out of compute reports that those attacks failed, not that no adversary succeeds, so the bound holds only against adversaries whose own effort is smaller than the search you funded.

open as a page

A report claims 99% evasion success against your malware classifier. What do you require before funding a response?

level: principalimportance: should knowfreq 38%

basics

~10 s

Require the goal before the number: any wrong verdict or one chosen verdict, and in which direction. Fund against the malicious-read-as-benign rate on realistic files at a stated attempt budget.

open as a page

What does a per-request compute ceiling on a transcription API buy against clients who maximise per-call work?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

It bounds the worst case; it does not remove the asymmetry. The attacker parks just under the ceiling and still buys several times the median work at one price, and the ceiling lands hardest on genuinely difficult audio.

open as a page

In a red-team report on a malware classifier, how do you scope a chosen-verdict flip that reproduces once in five attempts?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Ask what the five are. One success in five attempts on one file, against a scanner the attacker runs offline, is a capability costing five attempts; one file in five is partial coverage. Report it separately.

open as a page

Two suppliers' robustness claims for the same classifier can't be ranked — what goes in your recommendation?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Not a rank. Either specify one adversary in the contract and make both re-test under it, or state in writing which axis each claim leaves blank, record the numbers as unverified, and decide on criteria you can re-measure yourself.

open as a page

A partner's records already skewed your live forecaster and the fix lands at the next retrain — who decides whether it keeps serving?

level: principalimportance: nice to knowfreq 30%

basics

~10 s

Not the responder. The owner accountable for the downstream service accepts it, on a written statement of what is skewed, how wide, for how long, and what the alternative costs.

open as a page

What can you honestly promise about a code assistant that retrains weekly on user activity?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

A statement about the model that shipped, not about the one shipping next week. Because every retrain re-opens the same channel, the honest answer is a trajectory with named limits on contribution, not a status of clean or safe.

open as a page

An outsourced analyst's triage clicks become your model's labels — do you accept that write path?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Usually yes, but only once it is written down as a production authorization surface. The decision is not whether the vendor is trustworthy; it is what attribution, gating and dwell you will fund, and what freshness that costs.

open as a page

Onboarding wants to drop bank verification to lift signup conversion - what do you own in that call?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Naming what the step buys: it holds one column the applicant cannot set for free. Removing it moves score mass onto retypeable fields and drops the cheapest successful application to near zero. Price that shift; the conversion appetite is the business's.

open as a page

A product owner wants to drop confidence scores from a paid moderation API as a security control. What do you say?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Support it as a repricing with a number attached, never as a boundary. Say which families move to which cost, insist the multiplier is measured, and be explicit that paying integrators lose a field some will rebuild.

open as a page

Product wants the whole liveness check on-device for latency and offline use — what do you require stays server-side?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Whatever authorizes a consequential action. Ship the model if latency demands it, but the decision it feeds must be made server-side on evidence the client cannot mint, with the device result treated as a hint rather than a credential.

open as a page

A vendor asks your evaluation lab to test its model without receiving the weights — how do you decide?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Refusing the grant makes the report partly a measurement of the vendor's secrecy, which is not a safety property and cannot support a clearance decision. Take the query-only run only as a supplementary realism datapoint, and say in writing what the absent grant means.

open as a page

showing 31–46 of 46