skip to content

In vLLM, what decides how many KV-cache blocks the engine allocates at startup?

level: seniorimportance: should knowfreq 44%

answer

  1. decided once, at startup
  2. profiling run measures the peak
  3. fixed pool of equal blocks
  4. weights plus activations come first
  5. override pins, never creates

basics

~20 s

A startup profiling run measures weights plus peak activation memory, subtracts that from the share of VRAM allowed by gpu_memory_utilization, and divides the remainder by the bytes one block costs. The quotient is the block count, fixed for the process; --num-gpu-blocks-override replaces it.

solid answer

~50 s

vLLM allocates the whole KV cache once, at startup, as a fixed pool of equal-sized blocks. To size it, the engine loads weights, runs a dummy forward pass at the configured maximum batch shape to measure peak activation memory, and treats whatever is left inside the `gpu_memory_utilization` fraction of the card as the cache budget. Dividing that by the bytes one block costs — `2 (K and V) x block_size x layers x kv_heads x head_dim x dtype bytes` — gives the block count, which the startup log reports as a KV cache size in tokens plus an estimated maximum concurrency at `max_model_len`. That number never changes while the process runs: only the free/allocated split moves, tracked by the V1 `KVCacheManager` (V0's `BlockSpaceManager` is gone). `--num-gpu-blocks-override` pins the count for reproducible benchmarks or to leave headroom for another process; it does not create memory, so setting it above what was measured just moves the OOM to run time.

code

python · 6 lines
python
# KV bytes for one vLLM block: 32-layer GQA model, 8 KV heads, head_dim 128, fp16
block_size, layers, kv_heads, head_dim, dtype_bytes = 16, 32, 8, 128, 2
bytes_per_block = 2 * block_size * layers * kv_heads * head_dim * dtype_bytes

print(bytes_per_block // 1024, "KiB per block")
print((20 * 1024**3) // bytes_per_block, "blocks in a 20 GiB cache budget")

go deeper

for a junior

Know that vLLM claims a large slice of the GPU at startup and splits it into fixed KV blocks, and that the count is fixed once the server is up. You are not expected to do the byte arithmetic.

for a middle

Explain the startup sequence — load weights, profile a forward pass, allocate the remainder as blocks — and compute bytes per block from layers, KV heads, head dim and dtype. Know that the log prints cache size in tokens.

for a senior

Read the startup log as a capacity report: use the maximum-concurrency figure to predict preemption before traffic arrives, and diagnose a refusal to allocate blocks by naming which term of the budget grew. Know what the override is and is not for.

for a principal

Turn the arithmetic into policy: a standard max_model_len based on real prompt distributions rather than the model's advertised ceiling, a required minimum concurrency headroom before a config ships, and a rule that GPUs are not shared with processes the profiling cannot see.

## Blocks, not sequences vLLM does not reserve KV memory per request. It carves one big allocation into fixed-size blocks — each holding `block_size` tokens' worth of keys and values for every layer, 16 tokens by default — and hands blocks to sequences as they grow. Every question about "how much can this server hold" reduces to *how many blocks exist* and *how fast are they consumed*. The block count is decided exactly once, at engine startup, and is immutable for the life of the process. Understanding how it is computed is what lets you predict capacity before you deploy instead of discovering it under load. ## The startup profile The sequence is: 1. **Load weights.** Now you know the resident cost of the model itself. 2. **Run a profiling forward pass.** The engine runs a synthetic worst-case batch to measure peak *activation* memory — the transient tensors a forward pass needs. This is the term people forget, and it is not small at large batch-token budgets. 3. **Compute the budget.** Take `gpu_memory_utilization x total VRAM`, subtract weights, activation peak and framework overhead. What remains is the KV cache budget. 4. **Divide by bytes per block.** The quotient is `num_gpu_blocks`. The log then tells you the result in human terms: the KV cache size expressed in tokens, and a maximum-concurrency figure — how many simultaneous requests of `max_model_len` tokens the cache could hold. That concurrency line is the single most useful number in a vLLM startup log. If it says something like `1.3x`, your server can barely hold two full-length requests and will preempt constantly under real traffic. ## Bytes per block The arithmetic is worth being able to do out loud: ``` bytes_per_block = 2 * block_size * num_layers * num_kv_heads * head_dim * dtype_bytes ``` The leading 2 is keys and values. `num_kv_heads` — not attention heads — is what grouped-query attention shrinks, which is why modern models fit far more context per gigabyte than their parameter count suggests. `dtype_bytes` is 2 for fp16/bf16 and 1 when the cache itself is stored in fp8. For a 32-layer model with 8 KV heads and head dim 128 in fp16: `2 x 16 x 32 x 8 x 128 x 2 = 2 MiB` per block, so a 20 GiB cache budget yields 10,240 blocks — 163,840 tokens of KV capacity, shared across every request in flight. ## What the count controls Blocks are the unit of admission. The scheduler will not start a request it cannot allocate blocks for, and it grows a running sequence one block at a time as it decodes. So the block count sets: - **Concurrency ceiling.** How many sequences can be resident at once, given their lengths. - **Preemption pressure.** When free blocks hit zero and a running sequence needs another, something has to give. What exactly happens then — victim selection, and vLLM V1's choice to recompute rather than swap — belongs to the paged-KV-cache topic; here, the point is that the block count is the knob that decides how often you reach that state. - **Prefix-cache room.** Blocks retained for reuse compete for the same pool. ## The override `--num-gpu-blocks-override` sets `num_gpu_blocks` directly, skipping the profiling result. Legitimate uses: - **Reproducible benchmarking.** Profiling results shift slightly with driver versions, other processes on the card, and model config; pinning the count makes two runs comparable. - **Deliberate under-allocation.** Leaving room for a second process on the same GPU, or reproducing a preemption-heavy scenario on a bigger card than production uses. What it is not is a way to get more cache. The memory was measured, not guessed; overriding upward means the allocation succeeds, then a forward pass at real batch size runs out of memory later, which is a far worse failure than refusing to start. ## Failure modes to recognize - **Refuses to start, saying there is no memory for cache blocks.** Weights plus activations already consumed the budget. Real fixes: raise `gpu_memory_utilization` if the card genuinely has headroom, lower `max_model_len`, quantize weights or the KV cache, or shard across more GPUs. - **Starts, but maximum concurrency is barely above 1.** The server will thrash. Same fixes, plus reconsidering whether `max_model_len` needs to be the model's full advertised context. - **Fine on one node, OOM on another.** Something else is on the card, or the cards differ; `gpu_memory_utilization` is a fraction of *total* VRAM, not of *free* VRAM, so a neighbouring process breaks the arithmetic silently. - **`vllm:num_preemptions` climbing in steady state.** Not a startup failure at all, but the same root cause: too few blocks for the offered load.

  • When would you actually set --num-gpu-blocks-override?
    To pin cache size across A/B benchmark runs so the two are comparable, to reproduce a preemption-heavy production scenario on a larger dev GPU, or to leave headroom for another process sharing the card. Never to get more cache than profiling found — the allocation will succeed and then a real forward pass will OOM.
  • The startup log reports a maximum concurrency of 1.2x. What is it telling you and what do you do?
    The KV cache holds only about 1.2 requests at the configured max_model_len, so the server can barely run two full-length requests concurrently and will preempt under any real load. Lower max_model_len to what clients actually send, raise gpu_memory_utilization if the card has genuine headroom, store the KV cache in fp8, or shard across more GPUs.
  • Why doesn't the number of blocks grow when traffic rises?
    Because the cache is one allocation made at startup; the pool is fixed and only the free-versus-allocated split moves. That is deliberate — a growing allocator would fragment and would race the activation memory a forward pass needs. Capacity planning therefore happens before launch, not at run time.
  • Two identical pods, same flags, and one fails to allocate blocks. What differs?
    Almost always something else is resident on that GPU, because gpu_memory_utilization is a fraction of total VRAM rather than free VRAM. A leftover process, a monitoring agent holding a context, or an MIG/partitioned card changes the arithmetic. Check what else has memory on the device before touching the flags.

saying these in an interview costs you the question

  • Thinks the KV cache grows dynamically as load rises
  • Believes the memory fraction covers only model weights
  • Raises --num-gpu-blocks-override expecting more cache
  • Forgets peak activation memory in the budget
  • Assumes block count depends on requests in flight

context