skip to content

Spending Someone Else's Budget

Consumption is the goal you pursue when no data is reachable at all, and the number that matters is the ratio between one cheap request and the operator's fan-out. Interviewers test cost asymmetry.

on this pageshow

explore

questions

4

Why can someone uploading one file to a summarising service outspend its per-request token cap?

level: juniorimportance: must knowfreq 62%

answer

  1. ask what the limit is scoped to
  2. one submission is many requests
  3. the cap counts a call, not an upload
  4. per-request cost versus per-attacker cost

basics

~20 s

A per-request token cap bounds one model call, not one submission. An accepted upload fans out into extraction, a call per unit of content, an aggregation pass and retries, so dozens of individually compliant calls are billed to the operator.

solid answer

~50 s

The cap and the cost live at different scopes. A per-request token cap is enforced on a single model call, and every call the pipeline makes can sit comfortably under it while the submission as a whole costs orders of magnitude more. One accepted upload becomes many units of work: extraction, one or more summarising calls per unit, an aggregation pass over the partial results, sometimes a cross-check pass, plus a retry whenever a call returns something the next stage cannot parse. The attacker pays for one upload; the operator pays for every derived call. That is the asymmetry that matters here: per-REQUEST cost against per-ATTACKER cost. It is also why this class stays attractive when there is no private data to reach and nothing to exfiltrate at all, because the payoff is the operator's money and quota rather than anyone's data.

go deeper

for a junior

Be ready to say plainly that the cap applies to one model call while the cost applies to one submission, and that an upload becomes many calls. Naming that scope mismatch out loud is most of the answer.

for a middle

Explain the mechanics: what turns one accepted file into many units of work, where retries and aggregation passes add calls, and why a per-request wall clock resets rather than accumulates.

for a senior

Show you can argue impact without a confidentiality story. Talk about the ratio of submitter cost to operator cost, whether an unauthenticated party can repeat it, and what a single measured run does and does not establish.

for a principal

Own the framing that the exposure is a product of unit price and a multiplier the ingest design chooses. Be able to say who absorbs the bill and what a change at one stage actually moves rather than removes.

## The two costs are not the same cost Picture a service that summarises documents members of the public upload. Anyone can submit a file; the pipeline extracts its content, splits it into units, summarises each unit with a model call, aggregates the partial summaries, and sometimes runs a further pass to check the aggregate against the source. A person reads the result at the end. There are two costs in that sentence and they are borne by different people. - **The submitter's cost** is one upload: some bandwidth and a single HTTP request. It does not scale with what happens afterwards. - **The operator's cost** is every model call the pipeline makes because of that upload, plus the retries, plus whatever capacity those calls occupied while other work waited. When someone chooses a submission specifically because of how much downstream work it becomes, they are trading their cheap unit against the operator's expensive one. Consumption is the objective, not a side effect of some other attack. ## What the cap is actually scoped to A per-request token cap is a correct, well-implemented control and it does exactly what it says: it refuses a single model call whose input plus output would exceed a bound. It is enforced per call. It has no view of how many calls exist, which submission they belong to, or who caused them. A pipeline can make two hundred calls for one upload and every one of them can be under the cap, because being under the cap is a property of a call and the cost is a property of the set. The same is true of a per-request wall clock. A timeout that kills any single call after a fixed number of seconds bounds the latency of that call. It does not bound the elapsed time of a submission, because the submission is a sequence of calls and the timeout resets on each one. Both controls answer the question *is this request too big?* Neither answers *how many requests did one accepted submission become, and who chose that number?* This is the specific wrong answer an interviewer is listening for. Somebody says "we cap tokens per request, so cost is bounded" and stops. The follow-up is always the same: bounded per what? ## Where the multiplier comes from The count of derived calls is a function of the ingest stage's definition of a unit, and that definition usually reads properties of the submitted artefact: how many pages, sheets, sections or embedded objects it contains, how deeply structures nest, how many parts an extractor emits. A submitter who is choosing what to send is choosing those properties. Then the loop adds its own multipliers on top: a second pass over the partial results, a verification pass, and a regeneration whenever a stage's output fails the next stage's expectations. Each of those turns one unit into more than one call. Notice that none of this requires the submission to say anything. There is no instruction for a model to obey, no directive span, nothing a content screen would score as suspicious. The lever is structural, and it is the reason this family survives in deployments where the wording-level checks are good. ## Reading the evidence correctly A few directions are easy to get backwards: - Every call passing the cap proves that no single call was oversized. It proves nothing about the number of calls. - No private data being reached proves the read path was not exercised. It does not mean nothing was spent, and spend is the payoff here. - A retry that produced a discarded answer still cost input and output tokens. Work thrown away is work billed. - One measured submission cost is one sample from a probabilistic pipeline. Retry counts vary between runs, so a single number is a data point, not the method's cost. ## Where it stops working The method depends on the derived work being a function of attacker-chosen input, and on the marginal cost landing on somebody other than the submitter. If the number of units a submission may become stops depending on what was submitted, the multiplier disappears from that stage — though it can reappear at the next one, or simply move to the count of submissions if submitting stays free. And if the marginal cost of a submission is borne by whoever made it, the asymmetry that made the method worth building is gone; what remains is capacity planning. One last distinction worth having ready: this is not a network flood. Nobody is saturating a link or filling a connection table. The resource being drawn down is inference capacity and the bill attached to it, which is why the finding is argued in terms of cost and queueing rather than packets. The published GenAI risk lists give unbounded consumption its own entry precisely because it does not look like the other items.

  • If every model call in the pipeline stayed under the cap, what has the cap proved?
    Only that no single call was oversized. The cap is evaluated per call, so it says nothing about how many calls one accepted submission produced, nor about who chose the input that decided that number. The aggregate is invisible at that scope.
  • Why does this class stay attractive when there is no data to reach and nothing to exfiltrate?
    Because the payoff is not data. It is the operator's money and inference capacity: a bill somebody else pays, and a shared limit drawn down far enough that other work queues behind it. That payoff is available against a deployment with no private corpus at all, which makes it one of the few classes worth trying there.
  • Does a per-request wall-clock timeout bound how long one submission occupies the pipeline?
    No. The timeout resets on every call, so it bounds the latency of each call and not the elapsed time of the sequence. A submission that becomes two hundred short calls never trips it while occupying capacity for a long time.

A drink dispenser that limits every pour bounds each cup, not the tab. Someone who orders a round for two hundred people never exceeds the per-pour limit once.

saying these in an interview costs you the question

  • Claims a per-request token cap bounds total cost
  • Treats it as a request flood and counts requests only
  • Assumes no data touched means no impact
  • Thinks a discarded retry costs the operator nothing
  • Says the file's size alone decides the work

context

open as a page

In a file-summarising pipeline, why does an attacker pick structure over wording?

level: middleimportance: should knowfreq 45%

basics

~20 s

Structure decides the unit count, wording does not. Part, sheet and embedded-object counts and nesting depth set how many calls one submission becomes, and a structural file carries no directive span for a content screen to score.

open as a page

A public upload reached no private data yet drained a shared model quota - is that a finding?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Yes, when the asymmetry holds: an unauthenticated submission costing the operator far more than the submitter, repeatable, and drawing down capacity other tenants share. No data reached proves the read path was not exercised, not that nothing was spent.

open as a page

Does upload fan-out against a summarising service stay worth red-teaming as inference gets cheaper?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Exposure is unit price times a multiplier the ingest design chooses, and only the first term falls. Cheaper models and caching shrink the price; neither touches the multiplier, and a bound per submission relocates it to submission count.

open as a page