How does continuous batching in vLLM raise Llama serving throughput?
answer
- the scheduling unit is one decode step
- length skew wastes static batch slots
- finished sequences leave immediately
- new arrivals join within one iteration
- throughput up, per-stream latency down
basics
~20 sContinuous batching schedules per decode iteration instead of per batch: a finished sequence leaves the running batch immediately and a queued request takes its place at the next step, so the GPU never idles waiting for the slowest generation in a fixed batch.
solid answer
~50 sStatic batching groups N requests, runs them together, and returns when the *longest* one finishes. Because generation lengths vary wildly — one reply is 20 tokens, another is 800 — most slots in that batch spend most of their time computing padding or nothing at all, and no new request can start until the whole batch drains. Continuous batching (also called iteration-level scheduling, from the Orca paper; implemented in vLLM and TGI) makes the scheduling unit a single decode step. After each step the scheduler evicts sequences that hit a stop condition, admits waiting requests, and forms the next step's batch. A request that arrives mid-generation waits at most one iteration, not one batch. The effect is that GPU utilisation tracks demand rather than the tail of the length distribution, so tokens-per-second across the server rises several-fold at high concurrency. The tradeoff is that a single request's decode speed can drop as the batch fills, which is why vLLM exposes `--max-num-seqs` to bound how many sequences run concurrently.
go deeper
Know the term and the one-line idea: requests join and leave the running batch continuously instead of waiting for a whole fixed batch to finish.
Explain the mechanism — scheduling per decode iteration, evicting finished sequences, admitting queued ones — and why variable output lengths make static batching waste the GPU.
Demonstrate the operating tradeoff: tune the concurrent-sequence cap against an inter-token-latency SLO, and diagnose queueing caused by KV-cache exhaustion rather than scheduler limits.
Own the policy question: which workloads share a pool at all, since mixing long-prompt batch jobs with interactive chat on one server degrades the interactive SLO no matter how the scheduler is tuned.
## Why batching exists at all Autoregressive decoding is memory-bandwidth bound, not compute bound. To produce one token for one sequence, the GPU must stream the entire weight matrix set through its arithmetic units — for an 8B Llama at fp16 that is ~16 GB of reads — and then do a trivially small amount of matrix-vector maths. Doing the same read for 32 sequences at once costs barely more time, because the weights are read once and multiplied against a 32-row activation matrix instead of a single row. Batching therefore converts an almost-free increase in arithmetic into a near-linear increase in throughput, until you become compute or KV-cache bound. ## Static (request-level) batching and its failure mode The naive implementation collects requests for a few milliseconds, forms a batch, runs the prefill, then loops decode steps until every sequence in the batch has emitted an end-of-sequence token or hit its length cap. Two things go wrong. First, **length skew**. Generation lengths in real chat traffic are heavy-tailed. If 31 sequences finish at 40 tokens and one runs to 800, the GPU spends 760 steps computing a batch of size one while 31 slots are dead weight. Effective batch size — the average across the run — is far below the nominal size. Second, **head-of-line blocking on admission**. A request arriving one millisecond after the batch is formed waits for the entire batch to complete before it can even prefill. At high load this dominates time-to-first-token, and the p99 looks nothing like the p50. ## Iteration-level scheduling Continuous batching moves the decision point. The engine runs one forward pass that produces exactly one new token for every sequence currently in the running set. Between passes, the scheduler: 1. Removes sequences that produced a stop token, hit `max_tokens`, or were cancelled by the client, and frees their KV blocks. 2. Admits waiting requests, subject to the sequence-count limit (`--max-num-seqs` in vLLM) and available KV-cache blocks. 3. Runs prefill for newly admitted prompts — either as its own step or fused with decode, depending on the scheduler policy. 4. Forms the next batch. Because admission and eviction happen every iteration, the running batch stays as full as demand and memory allow. A finished sequence's memory is reusable within milliseconds rather than at the end of a batch. ## Prefill versus decode The two phases have opposite profiles. **Prefill** processes the whole prompt in parallel and is compute-heavy; a single 8k-token prompt can saturate the GPU on its own. **Decode** produces one token per sequence and is bandwidth-heavy. Naively interleaving them makes long prompts stall everyone else's token stream — you see it as periodic hitches in streaming output. Modern servers mitigate this with **chunked prefill**, which splits a long prompt into pieces and interleaves those pieces with decode steps so no single admission monopolises an iteration. vLLM exposes this behaviour through its scheduler configuration. ## What it does not fix Continuous batching raises aggregate throughput; it does not make one request faster. Inter-token latency for an individual stream grows as the batch grows, because each iteration does more work. This is the fundamental throughput/latency knob of self-hosted serving: raise `--max-num-seqs` and total tokens per second climbs while per-user smoothness degrades; lower it and each user gets a snappier stream while the GPU idles. Which side to favour is an SLO decision — an interactive chat UI cares about inter-token latency, a nightly batch-summarisation job does not. It also cannot admit requests when the KV cache is full. When memory runs out, vLLM preempts sequences — recomputing or swapping their KV state — and you see queueing delay and preemption counters rise in `/metrics`. That is a memory-sizing problem, not a scheduling one. ## What interviewers listen for The strong answer names the unit of scheduling (an iteration, not a batch), explains *why* it matters (length skew and admission blocking), and volunteers the tradeoff (aggregate throughput up, per-stream latency down) with the knob that controls it. A weak answer describes it as "batching requests together", which is static batching, the thing continuous batching exists to replace.
- What happens to a single user's streaming speed as concurrency rises under continuous batching?Inter-token latency increases. Each iteration produces one token per sequence, and a larger batch makes every iteration more expensive, so the stream visibly slows even though the server's total tokens per second is climbing. That is why servers expose a cap on concurrent sequences: it is the direct lever between aggregate throughput and per-user smoothness, and it should be set from the latency SLO rather than maximised.
- Why can a long prompt arriving mid-flight stutter everyone else's token stream?Prefill processes the entire prompt in one compute-heavy pass, which can occupy an iteration that would otherwise have produced tokens for every running sequence. Users see a hitch in streaming output. Chunked prefill fixes it by splitting the prompt into pieces and interleaving them with decode steps, so no single admission monopolises the GPU for a long stretch.
- The batch is not full but requests are still queueing. What is the likely cause?The KV cache, not the sequence limit, is the binding constraint: there are no free blocks to allocate for a new sequence's context, so the scheduler cannot admit it. Long contexts consume KV memory quickly. Look at preemption and cache-usage metrics, then either reduce max context length, raise the GPU memory fraction, quantize, or add GPUs.
saying these in an interview costs you the question
- Describing it as simply grouping several requests into one call
- Claiming it makes each individual request faster
- Assuming a new request must wait for the current batch to drain
- Ignoring that KV-cache capacity, not batch size, often limits admission
- Treating prefill and decode as having identical cost profiles