AML Threat Models & Taxonomy
You will learn the standard way to frame an attack on an ML system: what the adversary knows, what they want, and which lifecycle stage they hit, mapped onto MITRE ATLAS and the NIST AML taxonomy. Interviewers open with this framing question because every deeper answer — evasion, poisoning, extraction — must be scoped by a threat model.
on this pageshowhide
explore
- Vantage and Reach20 questions
- The White-Box Assumption4 questions
- Fields the Endpoint Returns4 questions
- Access Nobody Grants4 questions
- Weights You Handed Over4 questions
- Features Cheap to Rewrite4 questions
- The Adversary's Goal11 questions
- Integrity, Uptime, Secrecy3 questions
- Targeted or Merely Wrong4 questions
- Costing Without Breaking4 questions
- Point of Entry8 questions
- Before or After Training4 questions
- Learning From Its Users4 questions
- The Published Matrices7 questions
- Tactics Only ML Has3 questions
- Vocabulary for Disagreeing4 questions
questions
page 2 of 2Your network-flow alert-triage model stopped flagging one traffic class and request logs look normal — what does that rule out?
basics
~20 sAlmost nothing. Request logs establish which queries arrived, not what the weights encode. A change written into the disposition queue or the training table produces no request at all, so the evidence that bounds it is the upstream write record.
A merchant-risk model's most predictive columns are all self-declared - what do you report?
basics
~20 sAccuracy was measured on applicants with no reason to misstate those fields. Report the share of score sitting on columns the subject retypes for free, and what the cheapest application that flips a decline costs.
Your moderation API now rounds scores to two decimals and returns only the fired policy. What changed?
basics
~20 sTwo separate changes. Rounding coarsens the signal and creates ties, so probes must be larger or more numerous. Dropping the non-fired categories is the bigger cut: the reply now describes one boundary instead of several.
The model shipped inside a past mobile app release was extracted — what does shipping a replacement buy you?
basics
~20 sVery little on its own. Copies already taken cannot be recalled, and old installs keep the old file working for months. A replacement helps only once the server stops honouring what the old model produces.
Your white-box robustness evaluation was capped at contracted GPU-hours — what does that cost the bound?
basics
~20 sThe ceiling is only as tight as the search you paid for. A granted-access run that ran out of compute reports that those attacks failed, not that no adversary succeeds, so the bound holds only against adversaries whose own effort is smaller than the search you funded.
A report claims 99% evasion success against your malware classifier. What do you require before funding a response?
basics
~10 sRequire the goal before the number: any wrong verdict or one chosen verdict, and in which direction. Fund against the malicious-read-as-benign rate on realistic files at a stated attempt budget.
What does a per-request compute ceiling on a transcription API buy against clients who maximise per-call work?
basics
~20 sIt bounds the worst case; it does not remove the asymmetry. The attacker parks just under the ceiling and still buys several times the median work at one price, and the ceiling lands hardest on genuinely difficult audio.
In a red-team report on a malware classifier, how do you scope a chosen-verdict flip that reproduces once in five attempts?
basics
~20 sAsk what the five are. One success in five attempts on one file, against a scanner the attacker runs offline, is a capability costing five attempts; one file in five is partial coverage. Report it separately.
Two suppliers' robustness claims for the same classifier can't be ranked — what goes in your recommendation?
basics
~20 sNot a rank. Either specify one adversary in the contract and make both re-test under it, or state in writing which axis each claim leaves blank, record the numbers as unverified, and decide on criteria you can re-measure yourself.
A partner's records already skewed your live forecaster and the fix lands at the next retrain — who decides whether it keeps serving?
basics
~10 sNot the responder. The owner accountable for the downstream service accepts it, on a written statement of what is skewed, how wide, for how long, and what the alternative costs.
What can you honestly promise about a code assistant that retrains weekly on user activity?
basics
~20 sA statement about the model that shipped, not about the one shipping next week. Because every retrain re-opens the same channel, the honest answer is a trajectory with named limits on contribution, not a status of clean or safe.
An outsourced analyst's triage clicks become your model's labels — do you accept that write path?
basics
~20 sUsually yes, but only once it is written down as a production authorization surface. The decision is not whether the vendor is trustworthy; it is what attribution, gating and dwell you will fund, and what freshness that costs.
Onboarding wants to drop bank verification to lift signup conversion - what do you own in that call?
basics
~20 sNaming what the step buys: it holds one column the applicant cannot set for free. Removing it moves score mass onto retypeable fields and drops the cheapest successful application to near zero. Price that shift; the conversion appetite is the business's.
A product owner wants to drop confidence scores from a paid moderation API as a security control. What do you say?
basics
~20 sSupport it as a repricing with a number attached, never as a boundary. Say which families move to which cost, insist the multiplier is measured, and be explicit that paying integrators lose a field some will rebuild.
Product wants the whole liveness check on-device for latency and offline use — what do you require stays server-side?
basics
~20 sWhatever authorizes a consequential action. Ship the model if latency demands it, but the decision it feeds must be made server-side on evidence the client cannot mint, with the device result treated as a hint rather than a credential.
A vendor asks your evaluation lab to test its model without receiving the weights — how do you decide?
basics
~20 sRefusing the grant makes the report partly a measurement of the vendor's secrecy, which is not a safety property and cannot support a clearance decision. Take the query-only run only as a supplementary realism datapoint, and say in writing what the absent grant means.
showing 31–46 of 46