skip to content

What does a per-request compute ceiling on a transcription API buy against clients who maximise per-call work?

level: seniorimportance: nice to knowfreq 24%

answer

  1. bounded is not the same as removed
  2. the attacker parks just under it
  3. it lands on your hardest inputs
  4. set it from legitimate traffic
  5. bill in the unit you spend in

basics

~20 s

It bounds the worst case; it does not remove the asymmetry. The attacker parks just under the ceiling and still buys several times the median work at one price, and the ceiling lands hardest on genuinely difficult audio.

solid answer

~50 s

A ceiling on forward passes or compute units per call is the right first control, because it converts an unbounded per-call bill into a known one: the ratio of worst admissible work to median work becomes finite and yours to choose. What it does not do is remove the multiple — an adversary simply sits just under the ceiling and still gets several times a median call's work at the same price. And the ceiling does not fall on attackers; it falls on the hardest inputs, which in a call-centre product are noisy lines, poor handsets and heavy crosstalk, precisely the recordings the customer most needs transcribed. Set near the median it silently deletes your real tail, appearing as degraded transcripts rather than as an incident. The durable fix sits beside it: bill in the unit you spend in.

go deeper

for a junior

Know that a cap on how much work one request may consume bounds the damage, and that the same cap also refuses legitimate inputs that genuinely need that work.

for a middle

Explain why an adversary simply operates just under the cap, so the multiple is bounded rather than removed, and why the cap's cost lands on the hardest inputs rather than on attackers.

for a senior

Show that you would derive the threshold from real per-call work distributions, check which accounts fall above it, and make refusals explicit so the control's cost is measurable rather than silent.

for a principal

Own the pricing argument: the exposure exists because the billing unit and the cost driver differ, and choosing whether to meter compute is a joint product, finance and security decision with real market cost.

## What the ceiling actually does A cap on the work one call may consume — forward passes, decoding passes, compute units, wall-clock — is the correct first control, and it is worth being precise about what it achieves. Before the cap, the worst admissible request costs an unknown amount; the exposure is open-ended and nobody can state it. After the cap, the ratio between the worst admissible request and a median one is **finite, known, and a number somebody chose on purpose.** That change — from an accident to a decision — is most of the value, and it is why this control goes in first even though it does not close the issue. ## What it does not do It does not remove the asymmetry. An adversary who wanted maximum work per call now takes maximum-allowed work per call. If the cap sits at eight times the median, one attacker request is worth eight ordinary ones and still costs them one ordinary price. You have converted an unbounded multiple into a bounded one. That is a real improvement and it is not a fix, and a candidate who claims otherwise has misread the mechanism. It also does nothing about the two non-monetary costs — the expensive calls still occupy capacity, so other tenants still wait longer. ## Who actually pays for the ceiling This is the part that separates a thoughtful answer from a reflexive one. The cap does not discriminate between an adversary and a customer; it discriminates between **cheap inputs and expensive ones**. In a call-centre quality product the expensive inputs are noisy lines, poor handsets, distant microphones, heavy crosstalk, unusual accents. Those are not edge cases to be discarded — they are frequently the recordings the customer most needs transcribed, and often the ones tied to the outcomes the product is bought for. So a ceiling set close to the median quietly amputates the legitimate tail. Worse, it usually does so **invisibly**: the call returns a transcript, just a worse one, or a truncated one, and it appears as slowly worsening quality for a subset of customers rather than as an error anybody investigates. If you ship a ceiling, ship it with an explicit signal — a distinguishable error or a flagged result saying the input exceeded the compute limit — so the cost of the control is measurable instead of absorbed silently by the customers with the hardest audio. ## Choosing where it sits Set the ceiling from the distribution of **legitimate** work, not from the attack. Measure what real traffic consumes per call, pick a percentile you are willing to refuse or downgrade, and then look at *whose* calls fall above it before you ship. If the answer is "three enterprise accounts with recordings from a factory floor", the number is wrong and no amount of security reasoning makes it right. ## The control that changes the economics The ceiling bounds the exposure. What removes the asymmetry is fixing the mismatch that created it: the service bills in **audio minutes** and spends in **compute**. When the price tracks the resource, an input that costs eight times as much to answer is charged eight times as much, and the amplification stops being the operator's problem — it becomes a purchase. The residual risk is then capacity contention among tenants rather than a bill nobody expected, which is a smaller and much better-understood problem. That change is not free either: metered compute pricing is harder to explain, harder to forecast for the buyer, and may not fit the product's market. Which is exactly why this decision is a joint one rather than a security ticket. ## Why an admission check is not a substitute A tempting alternative is to predict the cost from the input and reject expensive submissions at the door. In a pipeline with data-dependent compute you generally cannot predict per-call cost reliably before running it — that is the whole reason the pipeline branches at runtime on a confidence signal instead of deciding up front. Cheap admission checks on obvious proxies such as duration are still worth having, but the control that actually binds is enforced during execution, because that is where the work happens. ## What an interviewer is listening for That you credit the ceiling with what it genuinely buys — a bounded, known, chosen ratio — and refuse it the credit it does not deserve. That you can say who absorbs the accuracy cost and insist the cost be made visible. That you set the threshold from legitimate traffic. And that you name the billing-unit mismatch as the thing that actually removes the asymmetry, while acknowledging that changing it is a product decision with its own price.

  • Where do you set the ceiling, and how do you defend the number?
    From the distribution of legitimate work per call, not from the attack. Pick a percentile of real traffic you are willing to refuse or downgrade, then look at which accounts fall above it before shipping. Defend it with that list, and with an explicit signal on refusal so the control's cost is measured rather than silently absorbed.
  • If pricing tracked compute instead of audio minutes, would there still be a finding?
    Yes, but a smaller one. The amplification is then paid for by whoever causes it, so the money problem becomes a purchase. What remains is shared-capacity contention — expensive calls still occupy the same accelerators, so other tenants still queue — plus the fact that the ceiling is still your only bound on a single call's worst case.
  • Can you just check the input up front and reject the expensive ones?
    Only partially. In a pipeline whose compute is data-dependent you generally cannot predict per-call cost before running it — that is why it branches at runtime rather than deciding in advance. Cheap proxies such as a duration cap help, but the control that actually binds is enforced during execution, where the work is being spent.
  • Is the residual, after a ceiling and metered pricing, worth further engineering?
    Usually not on its own. Once the worst case is bounded and priced, what is left is tenant fairness in a shared pool — an isolation and scheduling problem with well-understood answers. The security value was in naming the exposure and getting the unit right; grinding further against an adversary who is now paying full price rarely repays the effort.

Capping the size of a parcel stops anyone shipping a piano for the price of a letter. It does not stop them shipping the largest parcel every single time, and it also turns away the customer whose goods are genuinely bulky.

saying these in an interview costs you the question

  • Claims the ceiling closes the exposure
  • Sets the threshold from the attack, not real traffic
  • Ignores who absorbs the truncated hard inputs
  • Lets the cap degrade results silently
  • Never questions the billing unit

context