How do you split one LLM token budget across workloads with different deadlines?
answer
- allocate by deadline, not by importance
- reserve a floor for the waiting human
- shed early, shed the tolerant class
- stale items should die at the front
- an unbounded queue is deferred failure
basics
~20 sGive each workload a deadline class, reserve headroom for the interactive one, and let the rest queue. Under burst, shed or defer the deadline-tolerant classes explicitly rather than letting one shared budget starve the path users are waiting on.
solid answer
~50 sTreat the account's per-minute token allowance as a scarce shared resource and allocate it by **deadline**, not by perceived importance. Define a small number of classes — interactive (a human is waiting, seconds), near-line (minutes), and bulk (hours, or the batch endpoint) — and give the interactive class a reserved floor it never has to compete for. Everything else runs on what is left. Under burst, the system must shed deliberately: reject or defer bulk work with an explicit signal rather than letting it queue behind, and ahead of, requests a user is waiting on. Two disciplines make this real. Check the deadline **on dequeue**, not just on enqueue, and drop items whose answer is already worthless. And bound the queue, so backpressure reaches the caller instead of accumulating silently as latency and memory. The alternative — one unbounded FIFO — degrades every class together.
go deeper
Know that all of a system's model calls draw on one shared per-minute allowance, so a bulk job can starve a user-facing feature unless someone decides explicitly which work gives way.
Explain deadline classes, a reserved floor for interactive traffic, and why a bounded queue that rejects is better than an unbounded one that buffers. Be able to say what backpressure looks like to a caller.
Show the operational detail: deadlines re-checked at dequeue, shedding at admission before budget is spent, per-class utilisation and shed-rate metrics, and moving bulk work onto the asynchronous batch path to leave the contended pool.
Own the policy across teams. Be ready to argue static reservation against weighted fair sharing, state what the organisation guarantees each class, name who decides when quota is raised instead of work shed, and describe how the burst gets rehearsed.
## The budget is a shared resource, and it is small An account's per-minute token allowance is a fixed pool that every workload draws from: the interactive feature, the periodic re-scoring job, the backfill someone kicked off manually. Unlike CPU, it is not fairly scheduled by anything — whoever calls first consumes it. So the default behaviour of an unmanaged system is that the loudest workload starves the most latency-sensitive one, and it does so exactly when demand is highest. Take a claims platform: an adjuster-facing assistant that must answer inside a couple of seconds while someone is on the phone, a re-scoring job that refreshes claim priorities every few minutes, and a nightly classification run. All three share one allowance. When claim volume spikes after a storm, all three spike together, and without an explicit policy the adjuster's request queues behind thousands of re-scores. ## Classify by deadline, not by importance Every workload owner believes their workload is important, so importance is not a usable axis. Deadline is. Three classes are usually enough: - **Interactive** — a human is waiting; the answer is worthless after a few seconds. - **Near-line** — driven by an event; useful within minutes; nobody is blocked. - **Bulk** — driven by a schedule; useful within hours; belongs on the batch endpoint whenever the provider offers one. Moving bulk work onto the asynchronous batch path is the single biggest structural win, because it typically draws on a separate allowance and removes that demand from the contended pool entirely. What remains to be allocated is the interactive and near-line traffic that genuinely needs the synchronous endpoint. ## Reservation versus shedding Two mechanisms, and you need both. **Reservation** protects the floor. Guarantee the interactive class a share of the budget that lower classes cannot consume — a dedicated slice of concurrency, or a limiter that admits bulk work only when interactive consumption is below a threshold. The reserved floor should be sized from the interactive class's own peak, not from its average, because the whole point is to survive the burst. **Shedding** handles the excess. When offered load exceeds the pool even after reservation, something must be refused, and the design choice is *what* and *how visibly*. Refuse the most deadline-tolerant class first, refuse it early (at admission, before it consumes any budget), and refuse it explicitly — an immediate rejection that the caller can retry later is a far better outcome than a request that sits in a queue for two minutes and then fails anyway, having occupied memory and a connection the whole time. The instinct to keep everything and just go slower is the trap. Uniform degradation means the interactive class fails too, so you lose the requests that mattered and keep the ones that did not. ## Queue timeouts: check the deadline on dequeue A queue is a time machine that makes work older. An item admitted with a two-second budget that has waited ninety seconds should be dropped when it reaches the front, not executed. Running it spends real tokens on a result nobody will read, and worse, it delays the item behind it, which may still be inside its deadline. The rule is simple and frequently missed: **stamp each item with its deadline at admission and re-check it at dequeue**. Under sustained overload this converts a queue that would otherwise serve everything late into one that serves recent work on time and discards the rest — which is what the users actually want. For the same reason, prefer serving the *newest* deadline-sensitive work when the queue is deep. A strict FIFO under overload serves the stalest items first, which is precisely backwards. ## Backpressure must reach the caller An unbounded in-memory queue is not capacity, it is deferred failure. It converts an honest rejection into rising latency, growing memory, and a system whose failure arrives all at once and much later. Bound the queue. When it is full, reject at admission with a clear signal, and make sure that signal propagates to whatever is feeding you — an upstream service, a scheduler, a user-facing screen that can say "try again shortly". A caller that receives no signal keeps offering the same load, which is why silent buffering makes overload permanent. ## What is genuinely contested How to allocate between near-line classes has no settled answer. Static reservations are predictable and waste capacity when a class is idle. Weighted fair sharing uses the pool better but makes worst-case latency for any one class harder to reason about. Fully dynamic schemes that reallocate on observed demand are the most efficient and the hardest to debug at 3am. Most teams land on static reservation for the interactive floor — where predictability matters most — and something looser above it, then revisit when the mix changes. What is not contested: an explicit policy, however imperfect, beats the implicit one, which is always "whoever calls first wins, and the person waiting loses". ## Making it operable Three things make the policy real rather than aspirational. Tag every model call with its class at the source, so allocation is possible at all. Report utilisation and shed rate **per class**, so a shed spike is visible and attributable. And rehearse the burst — if nobody has watched the system shed under load, the first observation will be during the incident, and the usual discovery is that some path was never tagged and has been consuming the interactive floor all along.
- Why check a request's deadline again when it leaves the queue rather than only on admission?Because queue time is what consumes the deadline. An item admitted with a two-second budget that waited ninety seconds is already worthless, and running it spends real tokens while delaying the item behind it, which may still be in time. Re-checking on dequeue converts an overloaded queue that serves everything late into one that serves recent work on time and discards the rest.
- What is wrong with absorbing the burst in a large in-memory queue instead of shedding?It converts an honest rejection into rising latency and memory with no signal to the caller, so the upstream keeps offering the same load and the overload becomes self-sustaining. Failure then arrives all at once, later, and further from the cause. A bounded queue that rejects at admission gives the caller something to react to and keeps the failure attributable.
- How would you size the reserved share for the interactive class?From the interactive class's own peak demand plus margin, not its average, since the reservation exists precisely for the burst. Derive it from measured tokens per minute at peak, including any reasoning-token inflation, and re-derive it whenever the feature's prompt size or effort level changes. Then verify by watching shed rate per class during a real burst rather than trusting the calculation.
- Is static reservation or weighted fair sharing the better allocation scheme?It is genuinely unsettled. Static reservation is predictable and easy to reason about but strands capacity when a class is idle; weighted fair sharing uses the pool better but makes any one class's worst-case latency harder to bound; dynamic reallocation is the most efficient and the hardest to debug under pressure. A common compromise is a static floor for the interactive class and looser sharing above it.
saying these in an interview costs you the question
- Allocates the budget by perceived importance rather than by deadline
- Absorbs bursts in an unbounded queue instead of shedding
- Executes queued items that are already past their deadline
- Degrades all workload classes uniformly under overload
- Leaves bulk jobs on the synchronous endpoint competing with user traffic