skip to content

When would you split prefill and decode onto separate LLM server instances?

level: principalimportance: nice to knowfreq 30%

answer

  1. prefill compute-bound, decode bandwidth-bound
  2. one dial cannot serve two SLOs
  3. KV cache must cross the wire
  4. independent sizing and scaling per pool
  5. a scale play, not a default

basics

~20 s

When one continuously batched replica can no longer serve two SLOs at once. Disaggregation runs prefill workers and decode workers separately and ships the KV cache between them, so each can be sized, tuned and scaled independently — worth it only at cluster scale.

solid answer

~50 s

In a single continuously batched replica, prefill and decode share every forward pass, so one set of scheduler budgets has to satisfy two workloads with opposite profiles: prefill is compute-bound and bursty, decode is memory-bandwidth-bound and steady. Every tuning choice is a compromise between first-token latency and steady-stream latency. Prefill/decode disaggregation breaks the tie by running them as separate worker pools: prefill workers ingest prompts and transfer the resulting KV cache to decode workers over the interconnect, and each pool is sized, batched and scaled on its own signal. It pays when you run enough replicas that the two pools are each large, when prompts are long relative to outputs (or vice versa), and when you have genuinely distinct first-token and per-token SLOs. It costs a KV transfer per request, real interconnect bandwidth, and substantially more operational complexity — vLLM supports it through a KV connector layer, but it is not a default posture for a two-GPU deployment.

go deeper

for a junior

Know that prompt processing and token generation are different kinds of work, and that very large deployments sometimes run them on separate machines.

for a middle

Explain why the two phases want opposite scheduler settings, and that separating them requires moving the prompt's KV cache from the prefill worker to the decode worker.

for a senior

Be able to weigh the transfer cost against the tuning gain, and to say what you would tune in a colocated server first before proposing a split.

for a principal

Own the economics: derive the prefill-to-decode capacity ratio from your own input/output token mix, price the interconnect, account for the routing and failure complexity, and defend the decision against simply adding replicas.

## The tension this resolves A continuously batched server runs two workloads that want opposite things. **Prefill** is a compute-bound burst: thousands of token positions through the model at once, saturating the math units, arriving unpredictably with whatever prompt sizes users send. **Decode** is a memory-bandwidth-bound trickle: one position per sequence per step, dominated by streaming weights and KV cache out of memory, and it wants many concurrent sequences and short, predictable step times. Colocating them on one GPU forces one set of scheduler budgets to serve both. Chunked prefill makes the compromise much better than it used to be — it slices prefill so decodes are not stalled behind whole prompts — but it is still a compromise. A fat token budget favours ingestion and first-token latency; a thin one favours stream smoothness. You get one dial for two contracts, and at scale the two contracts diverge. ## What disaggregation does Split the deployment into a **prefill pool** and a **decode pool**. A request goes to a prefill worker, which runs the prompt through the model and produces its KV cache. That cache is then transferred to a decode worker, which owns the sequence for the rest of its life and generates tokens with no prefill work ever competing for its steps. The consequences follow directly and this is the shape of the argument to make out loud: - **Independent sizing.** Prefill capacity scales with input tokens per second; decode capacity scales with concurrent sequences and output tokens per second. In a colocated fleet you must buy both in the ratio your model happens to impose. Disaggregated, you buy each to fit the actual traffic — which matters enormously because the prompt-to-output ratio varies by an order of magnitude between a summarization service and a chat product. - **Independent tuning.** Prefill workers can use a huge token budget with no stream to protect. Decode workers can run high concurrency with short, uniform steps. Neither compromises for the other. - **Clean SLOs.** First-token latency becomes a property of the prefill pool plus the transfer; per-token latency becomes a property of the decode pool. Each has one owner and one scaling signal, instead of a shared queue where a spike in long prompts silently degrades everyone's token rate. - **Different hardware, potentially.** Prefill wants math throughput; decode wants memory bandwidth and capacity. At sufficient scale that is an argument for different GPU SKUs per pool. ## What it costs The KV cache must physically move. For a long prompt on a large model that is gigabytes per request, and it has to travel over whatever interconnect exists between the pools — fast within a node, much slower across nodes. That transfer sits directly in the first-token latency path, so a disaggregated deployment on a weak interconnect can be *worse* on TTFT than a colocated one, which is the failure mode to name. Beyond bandwidth: you now operate two services with two scaling policies, a routing layer that must track which decode worker owns which sequence, and a failure model where losing a decode worker mid-generation loses in-flight sequences that a colocated replica would have kept. Prefix reuse also gets harder — a shared system prompt cached on a prefill worker does not automatically help a decode worker, so cache-locality routing becomes its own design problem. ## When to say yes, and when to say no Say yes when: the fleet is large enough that both pools are meaningfully sized; the prompt-to-output ratio is lopsided enough that colocated sizing wastes real money; you have separate, contractual first-token and per-token targets; and your interconnect is fast enough that the transfer is a small fraction of your TTFT budget. Say no when: you run one or two replicas, where the operational complexity dwarfs any gain; the interconnect is ordinary networking and the transfer would dominate; prompts are short, so prefill is cheap and never really interferes; or you have not yet tuned the colocated case, because chunked prefill plus sensible scheduler budgets recovers a large share of the benefit for none of the complexity. ## Where it stands as technology This is an actively moving area rather than settled practice. vLLM supports prefill/decode disaggregation through a KV connector layer that handles transferring cache between instances, and the large-scale serving frameworks built on top of these engines have made it a headline feature. Treat it as a scale play whose economics you should be able to justify with the ratio of input to output tokens in your own traffic — the interviewer is testing whether you reach for architecture before you have exhausted configuration, not whether you can recite the design.

  • What is the single biggest risk that makes disaggregation lose to a colocated deployment?
    The KV transfer landing in the first-token latency path. A long prompt on a large model produces gigabytes of cache; over an ordinary network that transfer can exceed the prefill it was meant to isolate, making TTFT worse than colocated serving. Disaggregation assumes a fast interconnect between pools, and the sizing exercise is to show the transfer is a small fraction of the TTFT budget.
  • What should you try before disaggregating?
    Tune the colocated case properly. Chunked prefill already interleaves prefill slices with decode tokens, so sizing the per-step token budget against your actual SLO recovers much of the benefit. Also consider prefix reuse for shared system prompts, splitting traffic into separate replica pools by workload shape, and simply adding replicas. Architecture is the last move, not the first.
  • How does disaggregation change the failure model?
    It adds a stateful handoff. A sequence's cache lives on a specific decode worker, so losing that worker loses in-flight generations that a colocated replica would have survived, and the router must track ownership rather than treating replicas as interchangeable. You also gain two independent scaling loops that can oscillate against each other if they scale on correlated signals.

It is the difference between one chef who both preps and plates every order, and a kitchen with a prep station and a pass — worth the handoff only once volume makes the specialization pay.

saying these in an interview costs you the question

  • Proposes disaggregation for a two-GPU deployment
  • Ignores the KV transfer cost on the first-token path
  • Assumes it improves quality or reduces total compute
  • Reaches for it before tuning chunked prefill and scheduler budgets
  • Treats prefill and decode workers as stateless and interchangeable

context