A transcription API's request rate is flat and every transcript is correct, yet compute spend rose — what do you check?
answer
- the control counts the wrong noun
- few requests, expensive requests
- read the tail, not the mean
- attribute work per call to a client
- billing unit versus cost driver
basics
~20 sCheck compute per request as a distribution and per client — the p99 work per call and the expensive-path rate. A control that counts requests cannot see an attack whose shape is few calls each doing expensive work.
solid answer
~50 sThe reflex is to call this denial of service and reach for rate limiting, and that is the trap: rate limits count requests, and the request count is exactly the number this activity leaves alone. The metric that sees it is cost per request — compute units or forward passes per call — broken out by client and read as a **distribution**, not a mean, because a few percent of traffic at ten times the cost barely moves an average while it dominates the tail. Look for a client whose per-call cost sits in the tail while its call rate is unremarkable, and whose expensive-path rate is far above everyone else's. Then check the benign explanation first: a customer with genuinely difficult audio produces an identical picture, and what separates them is whether the expensive inputs look chosen and repeated.
code
text · 9 linesmetric baseline last 6 h
requests accepted (per 5 min) 2,010 2,065
p50 compute-units per request 1.0 1.0
p95 compute-units per request 2.4 2.6
p99 compute-units per request 2.5 11.9
second-pass (expensive path) rate 14% 16%
audio-minutes billed per hour 33,600 34,400
compute cost per hour $310 $394
...go deeper
Recall that a control which counts requests only sees request counts. If each call is doing more work, the number to look at is the work per call, not the number of calls.
Explain why a mean hides this and a tail shows it, and why the expensive-path rate maps more directly onto the mechanism than a currency figure does.
Show the full triage: per-client work distributions, the benign hard-audio explanation ruled out on evidence, and a finding stated as an exposure ratio rather than an accusation.
Own the reporting line — that revenue tracked submissions while cost tracked work — and decide what telemetry the organisation commits to emitting so this class is visible by default.
## Why the obvious answer is wrong The near-universal first answer to "spend went up and nothing else did" is *denial of service, so rate limit it*. It is wrong here in a specific and instructive way. A rate limit is a control denominated in **requests**, and this activity is a small number of requests. The amplification lives entirely inside a single call. You can set the limit as tight as your customers will tolerate and the attacker stays comfortably underneath it, because they were never near it. A control that counts the thing the attacker did not change cannot detect or bound the thing they did. The same argument disqualifies most of the neighbouring reflexes: connection limits, per-IP caps, and quotas denominated in audio minutes all measure volume of *submissions*, not volume of *work*. ## The metric that does see it Cost per request. In whatever unit your infrastructure actually spends — forward passes, compute units, accelerator-seconds — and with three properties: 1. **Per client.** An aggregate number blends the attacker into everyone else. The finding is a per-key statement. 2. **As a distribution, not a mean.** If three percent of calls cost twelve times the median, the mean moves by about a third while the p99 moves by a factor of five. The tail is where the signal lives; the mean is where it hides. 3. **Beside a path indicator.** The rate at which calls take the expensive branch is the most legible single number, because it maps directly onto the mechanism rather than onto a currency. Aggregate hourly spend is a lagging and insensitive detector. It will eventually move, but by then the question "who" has already been lost. ## Reading the picture honestly Suppose the numbers show request count flat, median cost flat, p99 cost up fivefold, expensive-path rate up a couple of points, and compute cost up by a quarter while billed audio minutes rose by two percent. That last pair is the interesting one: **revenue tracked submissions, cost tracked work, and the two came apart.** That gap is the finding, and it is stated in a language a capacity owner and a pricing owner both speak. ## The benign explanation comes first A customer whose audio is genuinely hard — noisy lines, poor handsets, heavy crosstalk — generates exactly this signature. They are not attacking anything; they are the customer your product was sold to. Confusing the two burns credibility fast, and the discriminators are unglamorous: - **Consistency.** Real corpora are mixed. A client whose calls sit in the expensive tail almost every time, with little spread, is behaving unlike any recording environment. - **Repetition.** Near-identical submissions that each land on the expensive path suggest a chosen input rather than a captured one. - **Account history.** New key, no product usage around the calls, no pattern that matches a working day. - **Business fit.** Does the client's stated use explain a tail nobody else in the same segment has? None of these is proof, and you should say so. What you can state with confidence is the exposure — the ratio, and the fact that the current controls cannot bound it — independently of whether this particular client is hostile. ## What you can and cannot conclude - Correct transcripts prove the outputs were fine. They prove **nothing** about what those outputs cost. - A flat request rate proves the request-counting controls were not tripped. It proves nothing about work consumed. - Aggregate spend holding roughly steady proves only that the attacker's share of traffic is small — which is the design of the attack, not evidence against it. - A high p99 on one key proves that key's calls are expensive. Whether that is adversarial is a separate judgment resting on the discriminators above. ## Instrumentation you probably need to add Most teams find they cannot answer this question with the telemetry they have, because per-call work was never emitted as a metric with the client attached. That is itself the outcome of the investigation: emit work-per-request and the branch taken, labelled by key, and alert on the tail of that distribution rather than on total spend. It is a small change, and it converts an invisible class of activity into an ordinary one. ## What an interviewer is listening for That you do not reach for rate limiting. That you say the words *cost per request, per client, as a distribution*. That you check the benign explanation before you write the word attacker. And that you notice the mismatch between the billing unit and the cost driver, because that is the durable part of the finding.
- How do you separate an attacker from a customer whose audio is genuinely hard?You often cannot, immediately. Look for consistency and repetition — real recording environments produce mixed difficulty, chosen inputs do not — plus account age, whether product usage surrounds the calls, and whether the client's stated business explains a tail nobody else has. Report the exposure with confidence and the attribution with hedging.
- Would a per-client quota fix this?Only if the quota is denominated in the resource being consumed. A quota in requests or in audio minutes still fails to bind the cost driver, exactly as a rate limit does. A quota in compute units or forward passes does bind it, and that change of unit is the actual control.
- Why is aggregate hourly spend a poor detector here?Because the attack is a small share of traffic at a high multiple. Three percent of calls at twelve times the median moves the mean by roughly a third and the p99 by a factor of five. The mean will drift eventually, but by then you have lost the per-client attribution the finding depends on.
- Does an anomaly detector on latency give you the same signal?Partly, and it is a reasonable proxy when work-per-call is not instrumented, since expensive calls take longer. But latency also moves with queueing, deployment changes and hardware, so it is noisy. Work consumed per call is the direct measurement and is what you want labelled by client.
saying these in an interview costs you the question
- Reaches for rate limiting and stops there
- Reads aggregate spend instead of the per-call tail
- Treats correct transcripts as evidence of no attack
- Never attributes cost per call to a client
- Names an attacker before excluding hard audio