skip to content

Continuous Batching

You will learn why LLM servers schedule at the iteration level instead of batching whole requests, and how prefill and decode compete for the same GPU. Interviewers ask this because continuous batching is the single biggest throughput win in modern serving, and explaining it proves you understand the autoregressive loop.

on this pageshow

questions

5

How does continuous batching differ from static batching in an LLM inference server?

level: middleimportance: must knowfreq 85%

answer

  1. scheduling granularity, not batch size
  2. decide per iteration, not per request
  3. finished slots refill immediately
  4. longest sequence stops blocking the rest
  5. also called in-flight batching

basics

~20 s

Static batching forms a batch, runs it to completion, and only then accepts new work, so every slot waits for the longest generation. Continuous batching re-forms the batch before each decode step: finished sequences leave and queued requests join immediately.

solid answer

~50 s

Static (request-level) batching groups N requests, runs the generation loop until all of them stop, and only then starts the next batch. Because generations have wildly different lengths, most slots finish early and sit idle while one 900-token answer drains, and any request that arrives a moment after launch waits for the whole batch. Continuous batching — also called in-flight batching — moves the scheduling decision to the iteration boundary. Before every forward pass the scheduler rebuilds the running set: sequences that emitted a stop token are evicted and returned to their client right away, and waiting requests are admitted into the freed slots if there is KV-cache space for them. Per-token compute is unchanged; what disappears is idle slot time and head-of-line blocking. In practice this is the single largest throughput win in modern LLM serving, often several times the static-batch number at the same latency.

go deeper

for a junior

Know the one-line contrast: a static batch runs to completion before accepting anything new, while a continuous batch is rebuilt every generation step. Being able to say that finished requests return immediately is enough at this level.

for a middle

Be ready to explain iteration-level scheduling mechanically: the running versus waiting sets, eviction of finished sequences, admission of new ones, and why each sequence's own KV cache lets sequences at different positions share one forward pass.

for a senior

Expect to connect it to production numbers — why static batching shows high GPU utilization while doing padded work, why queueing delay dominates time-to-first-token under load, and what still bounds admission once scheduling is continuous.

for a principal

Own the framing that throughput here is an occupancy problem, not a batch-size problem, and be able to argue when a workload (single-tenant, low-concurrency, hard per-request deadlines) gains nothing from it and should be sized on replicas instead.

## Why batch at all Generating one token requires a forward pass that streams the model's entire weight matrix out of GPU memory. Those weight reads dominate the cost of a decode step, and they are shared: running eight sequences in the same forward pass reads the weights once, not eight times. Batching is therefore close to free throughput, which is why every serving engine does it. The interesting question is not *whether* to batch but *when the batch membership is decided*. ## Static, request-level batching The straightforward implementation treats a batch like a unit of work. Collect N requests, call a generate loop, and iterate until every sequence in the batch has produced a stop token or hit its length limit. Then return all N responses and start assembling the next batch. This has two separate costs, and candidates usually name only one. The first is **internal waste**. Generation lengths in a real chat workload are heavily skewed: most answers are short, a few are very long. If one sequence runs to 900 tokens and the other seven stop at 30, those seven slots contribute nothing for 870 steps. The GPU still executes a batch of eight positions each step; seven of them are padding or masked no-ops. Utilization looks high on a dashboard while most of the work is fictional. The second is **head-of-line blocking on arrival**. A request that shows up one millisecond after the batch launched cannot be added. Its time-to-first-token now includes the *entire* remaining runtime of a batch it has nothing to do with. Under load this queueing delay, not model compute, becomes the dominant term in what the user experiences. ## Iteration-level scheduling Continuous batching moves the scheduling decision from the request boundary to the **iteration** boundary. The engine keeps two sets: a *running* set that participates in the current forward pass and a *waiting* set of admitted-but-not-yet-scheduled requests. Before each step the scheduler: 1. Removes sequences that finished on the previous step and streams their final token to the client immediately — no waiting for batch-mates. 2. Admits requests from the waiting set into the freed capacity, provided there is a free sequence slot and enough KV-cache space to hold their state. 3. Builds the forward pass over whatever mix of sequences it ended up with. Because each sequence carries its own KV cache, sequences at completely different positions coexist in one batch with no interference: one may be on token 3 of its answer and another on token 800. Nothing is padded to a common length during decode, since each running sequence contributes exactly one new token position per step. A newly admitted request still has to be **prefilled** — its prompt must be run through the model to populate its KV cache before it can decode. That prefill is a large, compute-heavy piece of work relative to a single decode step, which is what makes mixing prefill and decode in one continuously batched server its own tuning problem. ## What continuous batching does *not* do It does not reduce the FLOPs per token, and it does not make an individual request faster in isolation — a single request alone on an idle GPU sees no benefit at all. It is not client-side batching either: the client still sends one ordinary streaming request and knows nothing about who it is batched with. And it is not simply "a bigger batch size": you can run static batching with batch size 256 and still be slower than continuous batching at 64, because the batch-size number describes capacity while continuous batching describes *occupancy*. ## Naming and where you meet it Different projects use different words for the same idea. vLLM's V1 scheduler does iteration-level scheduling by default. TensorRT-LLM calls it **in-flight batching**, and Triton exposes it through the TensorRT-LLM backend — importantly, Triton's generic **dynamic batching** is a *different* mechanism that groups independent inference requests into one larger request-level batch, which is the right tool for a vision model and the wrong tool for autoregressive generation. TGI has done continuous batching since its early versions. ## The knobs that bound it Continuous batching is not unbounded concurrency. Every engine caps how many sequences may run at once — vLLM spells this `max_num_seqs`, while TGI bounds admitted work with `--max-concurrent-requests` — and the real ceiling is usually KV-cache memory rather than the configured cap. When the cache fills, the scheduler stops admitting and may evict already-running sequences, so "the batch is continuous" never means "the queue is empty".

  • If requests can join mid-flight, why does the server still have a queue at all?
    Because admission is capacity-bounded. The scheduler will only add a sequence if a slot under the concurrent-sequence cap is free *and* there is KV-cache space to hold its prompt and expected output. Under load both bind, so arrivals sit in a waiting set until a running sequence finishes and frees blocks. Continuous batching removes the *artificial* wait for batch-mates, not the real wait for GPU memory.
  • Does a request that joins an in-flight batch have to redo any work the running sequences already did?
    No. Each sequence owns its own KV cache and its own position, so the batch is a batch of independent token positions rather than one shared sequence. The newcomer prefills only its own prompt, then contributes one token per step like everyone else. The running sequences are unaffected apart from sharing the step's compute.
  • Why does padding largely disappear during decode under continuous batching?
    In static batching, variable-length sequences are padded to a common length so they fit a rectangular tensor. Under iteration-level scheduling each running sequence contributes exactly one new token position per step, so the decode batch is naturally rectangular with no padding. Ragged prompt lengths during prefill are handled by variable-length attention kernels that pack tokens instead of padding them.

Static batching is a bus that leaves full and will not stop again until the last passenger's destination; continuous batching is a bus that lets people off and on at every corner.

saying these in an interview costs you the question

  • Says continuous batching just means using a larger batch size
  • Claims a new request must wait for the current batch to drain
  • Confuses it with client-side request batching or a batch API
  • Asserts all sequences in the batch must be the same length
  • Thinks it lowers per-token compute rather than removing idle slots

context

open as a page

What does chunked prefill change about how an LLM server schedules a long prompt?

level: middleimportance: should knowfreq 55%

basics

~20 s

Chunked prefill splits a long prompt's prefill into pieces that fit the scheduler's per-step token budget, so each step can carry prefill tokens and ongoing decode tokens together instead of stalling every generating sequence behind one huge prompt.

open as a page

Your LLM server's waiting queue keeps growing under load — how do you diagnose the bottleneck?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Compare the running-sequence count against the configured concurrency cap and the KV-cache utilization gauge. Pinned at the cap with spare cache means the cap is throttling you; pinned cache means memory is the limit and no scheduler knob will fix it.

open as a page

How do you tune an LLM server's per-iteration token budget and concurrent-sequence cap?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Size the per-iteration token budget by how long a step may take, since step duration is every user's inter-token latency, and size the concurrent-sequence cap by how much KV cache you actually have. Then measure at your target load rather than guessing.

open as a page

When would you split prefill and decode onto separate LLM server instances?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

When one continuously batched replica can no longer serve two SLOs at once. Disaggregation runs prefill workers and decode workers separately and ships the KV cache between them, so each can be sized, tuned and scaled independently — worth it only at cluster scale.

open as a page