What does chunked prefill change about how an LLM server schedules a long prompt?
answer
- prefill made divisible, not atomic
- fits a per-step token budget
- mixed prefill and decode in one pass
- smoother streams, slightly later first token
- already the default in current vLLM
basics
~20 sChunked prefill splits a long prompt's prefill into pieces that fit the scheduler's per-step token budget, so each step can carry prefill tokens and ongoing decode tokens together instead of stalling every generating sequence behind one huge prompt.
solid answer
~50 sPrefilling a 30k-token prompt is a single large piece of work. Without chunking, the scheduler spends whole iterations on it, and every sequence already generating produces no token during those iterations — users see their stream freeze. Chunked prefill breaks that prefill into slices sized by the per-iteration token budget and mixes a slice with the decode tokens of the running sequences in the same forward pass. The result is much smoother inter-token latency and steadier GPU occupancy, paid for with a slightly worse time-to-first-token for the long prompt itself, since its prefill is now spread across several steps. Output is unchanged — chunking is a scheduling decision, not a numerical one. In vLLM 0.27 this is on by default (`enable_chunked_prefill` defaults to true), so the modern question is when to turn it off or resize the chunk, not how to switch it on.
code
bash · 3 linesvllm serve meta-llama/Llama-3.1-8B-Instruct \
--max-num-batched-tokens 2048 \
--no-enable-chunked-prefillgo deeper
Recall that a prompt has to be processed before any token comes out, and that chunked prefill feeds that processing to the GPU in slices so other users' streams keep moving.
Explain the mechanism: a per-iteration token budget, decode tokens accounted first, a prefill slice filling the remainder, and why slicing a causal prefill is numerically equivalent to one pass.
Demonstrate that you know it is on by default in current vLLM and can argue the reverse case — sizing the chunk against a first-token SLO versus a stream-smoothness SLO, and the workloads where you switch it off.
Frame it as a fairness policy, not a performance flag: chunking decides whose latency absorbs a large prompt, so own the decision to route long-prompt traffic to its own replica pool rather than letting one setting arbitrate between tenants.
## The two kinds of work in one server An LLM server is always doing two structurally different things. **Prefill** runs a whole prompt through the model at once to populate that request's KV cache; it processes thousands of token positions in one pass and saturates the GPU's math units. **Decode** advances each running sequence by exactly one token; it is a thin slice of compute over the same enormous weight matrices. A continuously batched engine has to fit both into a single stream of forward passes on one GPU. ## What goes wrong without chunking If prefill is atomic — the whole prompt or nothing — then admitting a 30k-token request means the scheduler dedicates one or more entire iterations to that prompt. During those iterations the sequences already generating advance by zero tokens. From a user's side the symptom is unmistakable: a stream that has been producing smooth output suddenly pauses for hundreds of milliseconds and then resumes. The pause has nothing to do with that user's request; it is someone else's prompt landing. Engines that scheduled prefill atomically also faced an unpleasant policy choice. Prioritize prefill and you get good time-to-first-token but stuttering decode. Prioritize decode and long prompts starve at the door. Neither setting is good, because the unit of work was too coarse. ## The mechanism Chunked prefill makes prefill divisible. The scheduler is given a **per-iteration token budget** — the maximum number of token positions it will put through one forward pass. It first accounts for the running sequences' decode tokens (one each), then fills the remaining budget with a slice of some waiting request's prompt. Next iteration, the next slice. After enough iterations the prompt is fully prefilled and that request starts decoding like any other. Two properties make this work. First, prefill is causal and left-to-right, so prefilling positions 1–2048 and then 2049–4096 with the earlier keys and values already in the cache produces exactly the same cache as one pass over 1–4096. The results are numerically equivalent up to ordinary floating-point non-determinism; nothing about the model's output distribution changes. Second, a mixed batch of "one token from each decoder plus a slab of prefill" is a shape modern attention kernels handle natively via variable-length packing, so there is no padding tax for the mix. ## What it trades The honest tradeoff is per-request time-to-first-token against everyone else's inter-token latency. A prompt whose prefill is spread over six iterations reaches its first output token later than one that got a dedicated pass — sometimes noticeably later, if the server is busy and each iteration also carries a full decode batch. What you buy is that the other N sequences kept producing tokens the whole time, and that the GPU is never running a decode-only step at low math utilization when there was prefill work available to fill it. On a shared multi-tenant endpoint that is almost always the better trade, which is why it became the default. The chunk size is the dial. A **larger** per-iteration token budget means fewer, fatter steps: better prefill throughput and better TTFT, but a longer wall-clock time per iteration, which directly lengthens the gap between output tokens for every running sequence. A **smaller** budget means finer-grained interleaving and smoother token streams, at the cost of more scheduler overhead and slower prompt ingestion. There is no universal right answer; it follows from whether your SLO is written on first-token latency or on steady-stream smoothness. ## Current defaults matter here In vLLM 0.27 chunked prefill is enabled by default (`enable_chunked_prefill` is true), and it can be turned off with `--no-enable-chunked-prefill`. An answer framed as "enable chunked prefill to fix your TTFT spikes" describes an older release and will be heard as stale. The interesting cases now run the other way: a single-tenant offline batch job that only cares about total throughput and prefers maximum-size prefill passes; a model or backend whose attention kernel does not support mixed prefill-decode batches; or debugging, where you want the simplest possible scheduling behaviour. Other engines expose their own prefill-token ceiling rather than reusing vLLM's spelling, so name the engine when you name the knob. ## How to tell it is your problem The signature is a decode-latency distribution with a long right tail that correlates with the arrival of long prompts, while average throughput looks fine. Aggregate tokens per second hides it completely; you have to look at per-step timing or at the inter-token latency percentiles of individual streams.
- Does chunking a prefill change the model's output?No. Prefill is causal and left-to-right, so processing positions in consecutive slices with earlier keys and values already cached yields the same KV cache as one large pass, up to ordinary floating-point non-determinism. Chunked prefill is purely a scheduling decision about which token positions ride in which forward pass; sampling, logits and stop conditions are untouched.
- When would you deliberately turn chunked prefill off?When nothing is waiting on smooth token streams. An offline single-tenant batch job scored on total throughput may prefer maximum-size prefill passes with no interleaving overhead. It is also the right move when a model's attention backend does not support mixed prefill-decode batches, or when you are isolating a scheduling bug and want the simplest possible step composition.
- How does the chunk size interact with per-token latency for sequences already decoding?Directly. Every running sequence gets one token per iteration, so its inter-token latency is the wall-clock duration of a step. A fatter token budget makes each step longer, which stretches the gap between output tokens for everyone, while ingesting prompts faster. Shrinking the budget tightens the stream and slows prompt ingestion. You size it against whichever SLO you actually wrote down.
saying these in an interview costs you the question
- Says you must enable chunked prefill in current vLLM to get it
- Claims chunking the prefill degrades output quality
- Thinks it reduces total prefill compute rather than spreading it
- Confuses chunked prefill with splitting the prompt into separate requests
- Assumes a bigger chunk is always better because throughput rises