skip to content

The Adversary's Goal

Whether the adversary wants a wrong answer, an expensive one, or a fact about the training data, and what each goal costs them. Interviewers ask because success is undefined until the goal is named.

on this pageshow

explore

questions

11

Why would an attacker send a transcription API valid audio chosen so each call does the most work possible?

level: juniorimportance: must knowfreq 44%

answer

  1. nothing is wrong; something is expensive
  2. the transcript is correct, the invoice is not
  3. measured per call, not per breach
  4. no perturbation budget is needed here
  5. the only rule is what the API accepts

basics

~20 s

The payoff is the operator's compute bill and queue time, not a wrong answer. Every request is admissible and every transcript correct, so nothing looks broken while the adversary buys expensive work at ordinary prices.

solid answer

~50 s

This is an availability-and-cost goal rather than an integrity one. The adversary never tries to make the model wrong; they try to make it expensive. They send audio the interface accepts under its own rules — right format, inside the duration cap, nothing a content check would reject — but chosen so a single call consumes far more compute than a typical one. There is no perturbation budget in play at all: the only constraint is that the input stays acceptable to the API, and the payoff is measured per call, in the operator's compute cost and in the queue other customers now wait in. Because every answer returned is correct, the signals a team actually watches — accuracy, error rate, content filters — all stay green. Nothing was fooled and nothing was stolen, which is exactly what hides it.

go deeper

for a junior

Be ready to say that an adversary can target cost and capacity rather than correctness, and that inputs which are entirely valid can still be an attack. Name the payoff: compute spent and queue time, per call.

for a middle

You are expected to explain why no perturbation budget is needed — the constraint is only that the input is admissible — and why every quality-oriented monitor stays green throughout.

for a senior

Show that you would look at cost per request rather than error rates, and that you can separate this from a volumetric flood and from an evasion attempt when triaging a real report.

for a principal

Own the framing question: whether a mismatch between what you bill for and what you spend compute on is a security finding, a pricing defect, or both — and who in the organisation is accountable for noticing it.

## The goal, named Adversaries against a deployed model are usually sorted by what they want: a wrong answer (an integrity goal), a fact about the data the model was trained on (a privacy goal), or the service to be worse for the people paying for it (an availability goal). The move here is the narrowest form of that third goal. The attacker does not want the service down and does not want any prediction changed. They want **one call to cost the operator more than the operator charges for it** — and they want that while every response the service emits is perfectly correct. Concretely: a speech-transcription API billed per minute of audio, used by a call-centre quality product. Behind it, the pipeline does what almost every production pipeline does — it spends less work on inputs it finds easy. A confidence check after a first, cheap decoding pass decides whether a heavier second pass is needed. Most real audio is clean enough that the cheap path answers it, and the service is priced on that assumption. An adversary who is simply a paying client, sending in-policy audio, can choose inputs that never take the cheap path. ## What the requests look like They look like customers. Each one is individually valid: correct format, under the duration limit, past whatever content screening exists. There are not many of them — this is not a flood, and volume is precisely the thing the attacker leaves alone. And the transcript that comes back is right. The attacker is not trying to corrupt an output; corrupting it would be a *different* attack with a different budget and a much higher price. ## Why "no perturbation budget" matters Most of adversarial machine learning is written under a budget: a norm and a radius bounding how far an input may be moved from a real one. That apparatus exists to serve an integrity goal — the input must stay recognisable as what it was while the decision flips. Here nothing has to flip. The attacker is not nudging a legitimate input a small distance; they are choosing, out of everything the interface already accepts, the item that costs the most to answer. The constraint is *admissibility*, not *proximity*. That is a far weaker constraint than a perturbation ball, which is why this attack is cheap to mount and why it does not need weights, gradients, scores, or any knowledge of the model beyond the published input rules. ## Why correct answers are the attack's best feature Almost every monitoring surface an ML team builds points at output quality: accuracy on a held-out slice, drift, error rate, refusal rate, content-filter hits, customer complaints. All of them are green here by construction. The security surfaces are not much better — there is no injected payload, no policy violation, no anomalous authentication, no data leaving. The only place the attack is visible is on the cost side of the house, which is usually owned by a different team and alerted on in aggregate, if at all. ## Where the cost actually lands Three places, and it is worth being able to name all three: 1. **Money** — compute consumed per call, against revenue that is denominated in audio minutes. The billing unit and the cost driver have come apart. 2. **Queue time** — the expensive calls occupy capacity, so other tenants' latency rises. They experience the attack as a slow product. 3. **Headroom** — capacity was sized against a traffic distribution the attacker is not drawn from, so the safety margin quietly shrinks. ## Telling it apart from its neighbours From an **evasion** attack: evasion wants a specific answer to be wrong and pays a perturbation budget to get it. This wants the answer to be right and expensive, and pays nothing. From a **flood**: a flood's whole mechanism is volume, and volume is what this deliberately does not have. From a **model-extraction** attack: extraction spends queries to acquire something (behaviour, parameters); here the queries are the point and nothing is acquired. ## What an interviewer is listening for That you name the goal cleanly instead of collapsing it into "denial of service, so rate limit it". That you do not treat correct output as evidence that nothing happened. And that you can state the payoff in the unit it is actually measured in — cost and queue time **per call** — because that unit is the whole difference between this and every other availability attack.

  • How is this different from an attacker trying to make the transcript wrong?
    An integrity attack needs the input to stay recognisable while the output changes, so it pays a perturbation budget and usually needs feedback from the model. This one needs neither: the output is allowed — preferred, even — to be correct. The attacker's success condition is compute consumed, not a flipped decision, so a much weaker adversary can run it.
  • Does the attacker need any feedback from the service?
    Very little. Their own wall-clock latency and their own invoice tell them whether a call landed on the expensive path. They need no confidence scores, no probability vector, no weights. In many cases they need no feedback at all, because the interface's published limits already tell them what the most expensive admissible input looks like.
  • Why do accuracy monitoring and content filters miss this entirely?
    Both are aimed at what the model outputs and what the input contains. Here the outputs are correct and the inputs are in-policy — there is nothing for either to flag. The only observable that moves is work consumed per request, which is a cost metric rather than a quality or a safety one, and typically nobody alerts on it.

It is the difference between forging a ticket and buying a valid one for the seat that costs the most to serve, over and over. Nobody is defrauded; the operator just loses money on every legitimate sale.

saying these in an interview costs you the question

  • Calls it denial of service and stops at rate limiting
  • Assumes the answers must be wrong for it to count
  • Thinks a perturbation budget is required
  • Treats correct output as proof nothing happened
  • Expects high request volume to be part of it

context

open as a page

Against a deployed ranking model an outsider can only query, which three goals can they choose between?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Three: a wrong output (integrity), a degraded or missing output (availability), or a fact the model should not reveal, usually about its training data (confidentiality). The goal defines success, so it is named before any technique.

open as a page

Against a malware classifier, what separates an untargeted evasion goal from a targeted one, and which costs more attempts?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Untargeted means any wrong verdict is success; targeted means one named verdict on one chosen file. Targeted is harder: the attacker must arrive somewhere specific rather than merely leave the right answer, so it costs more attempts.

open as a page

In a transcription API that skips a heavy second decoding pass on confident audio, what does that shortcut hand an adversary?

level: middleimportance: should knowfreq 33%

basics

~10 s

A multiplier. The shortcut makes compute input-dependent, so the ratio between a typical call's work and the most expensive call the interface still accepts becomes the attacker's amplification factor, available without perturbing anything.

open as a page

A team says an attacker who never obtains their model's weight file cannot breach its confidentiality. Why is that wrong?

level: middleimportance: should knowfreq 59%

basics

~20 s

Because the confidential asset is not the file. It is the training data the model absorbed, and secondarily the learned function itself, both of which leak through ordinary query access. A privacy attack can succeed against a model the adversary never obtains.

open as a page

With unlimited free re-scans of a local malware classifier, why does forcing one chosen verdict still cost more attempts?

level: middleimportance: should knowfreq 54%

basics

~10 s

The stopping rule is narrower. An any-wrong-verdict run ends at the first outcome that is not correct; a chosen-verdict run must pass those and keep going, burning more attempts and more restarts.

open as a page

A transcription API's request rate is flat and every transcript is correct, yet compute spend rose — what do you check?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Check compute per request as a distribution and per client — the p99 work per call and the expensive-path rate. A control that counts requests cannot see an attack whose shape is few calls each doing expensive work.

open as a page

A seller's traffic pushes 12% of ranking responses onto the default ordering while uptime stays green. Which property broke?

level: seniorimportance: should knowfreq 34%

basics

~20 s

You cannot say from the traffic alone. Compare the fallback share against the operator's written floor for model-served responses, then ask who the fallback ordering favours. If the seller gains under it, the finding is integrity.

open as a page

A report claims 99% evasion success against your malware classifier. What do you require before funding a response?

level: principalimportance: should knowfreq 38%

basics

~10 s

Require the goal before the number: any wrong verdict or one chosen verdict, and in which direction. Fund against the malicious-read-as-benign rate on realistic files at a stated attempt budget.

open as a page

What does a per-request compute ceiling on a transcription API buy against clients who maximise per-call work?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

It bounds the worst case; it does not remove the asymmetry. The attacker parks just under the ceiling and still buys several times the median work at one price, and the ceiling lands hardest on genuinely difficult audio.

open as a page

In a red-team report on a malware classifier, how do you scope a chosen-verdict flip that reproduces once in five attempts?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Ask what the five are. One success in five attempts on one file, against a scanner the attacker runs offline, is a capability costing five attempts; one file in five is partial coverage. Report it separately.

open as a page