skip to content

In vLLM, what does --gpu-memory-utilization control, and what does raising it buy?

level: middleimportance: must knowfreq 72%

answer

  1. a budget, not a cache size
  2. fraction of total, not free
  3. weights and activations come off first
  4. the remainder becomes KV blocks
  5. per instance — halve it when co-tenant

basics

~20 s

--gpu-memory-utilization is the fraction of each GPU's total memory one vLLM instance may use, default 0.9. Weights and peak activations are subtracted first, and whatever is left becomes KV cache — so raising it buys concurrency and context, not speed.

solid answer

~40 s

It is a per-instance budget expressed as a fraction of the card's **total** memory (not free memory), default `0.9`. At startup vLLM loads the weights, runs a profiling forward pass to measure peak activation and CUDA-graph memory, subtracts both from the budget, and turns the remainder into the fixed pool of KV-cache blocks it will use for the rest of the process's life. So the flag does not make anything faster — it decides how many blocks exist, which sets how many sequences can run concurrently and how long each may get. Raising 0.90 to 0.95 on an 80 GB card adds roughly 4 GB of blocks; pushing further risks OOM during graph capture or from anything else on the card. If two vLLM instances share one GPU, give each about 0.45.

code

bash · 9 lines
bash
# 13B bf16 on one 80GB card: leave ~10% headroom, cap context, log the KV size
vllm serve meta-llama/Llama-2-13b-chat-hf \
  --gpu-memory-utilization 0.90 \
  --max-model-len 8192 \
  --port 8000

# Two instances sharing one card: each takes roughly half the budget
vllm serve model-a --gpu-memory-utilization 0.45 --port 8000 &
vllm serve model-b --gpu-memory-utilization 0.45 --port 8001 &

go deeper

for a junior

Know the flag's name, that it defaults to 0.9, and that it is a fraction of the GPU's total memory rather than a number of gigabytes. Be able to say that the leftover after weights becomes the KV cache.

for a middle

Explain the startup sequence: load weights, profile peak activation, allocate the remaining budget as a fixed KV block pool. Interviewers want you to say that raising the value buys concurrency and context, never per-token speed.

for a senior

Show that you read the KV-cache size from the startup log and treat it as capacity, and that you know which levers to reach for instead when the value is already near the ceiling: max-model-len, fp8 KV cache, a quantized checkpoint, more GPUs.

for a principal

Own the policy: what utilization value your platform standardises on, how much headroom you keep for fragmentation and co-tenants, and whether you allow two instances per card at all given that a single OOM takes out a serving replica.

## What the flag is `--gpu-memory-utilization` (engine argument `gpu_memory_utilization`) takes a float between 0 and 1 and defaults to `0.9`. It is documented as the fraction of GPU memory to be used for the model executor, and it is a **per-instance** limit: it describes how much of one GPU a single vLLM instance intends to occupy, not how much of the GPU is currently free. Two details trip people up. First, the fraction is of **total** device memory, so `0.9` on an 80 GB card is roughly 72 GB regardless of what is already resident. Second, the budget covers the **whole instance** — model weights, activations, CUDA graph memory and the KV cache — not the KV cache alone. ## What vLLM does with it at startup The sequence at server start is roughly: 1. Load the model weights onto the GPU (or onto each GPU when sharded). 2. Run a profiling forward pass at the configured maximum batch shape to measure the peak transient memory the model needs — activations, workspace, and the memory CUDA graph capture will hold. 3. Compute `budget - (weights + peak transient)` and convert that leftover into a fixed number of KV-cache blocks. 4. Allocate that block pool up front and never grow it. This is why the number printed in the startup log — the KV-cache size in GiB or in tokens — is the number that actually matters operationally. The block pool is the server's real capacity: it caps how many sequences can be resident at once and how long each of them may become. ## The arithmetic, in the shape an interviewer wants Take a 13B model in bf16: about 26 GB of weights. On a 40 GB A10G-class card at `0.9`, the budget is ~36 GB; subtract weights and a couple of GB of activation/graph headroom and you have roughly 8 GB of KV cache. On an 80 GB card at `0.9` the same model leaves roughly 44 GB — five times the concurrency for the same model and the same flag value. The flag is a lever on the *remainder*, so its effect is highly non-linear in the weights-to-VRAM ratio: the closer the weights come to filling the card, the more dramatic a small change in utilization becomes, and the more likely it is that the honest answer is "this model does not belong on this GPU". ## Why you cannot just set 1.0 Raising the value trades safety margin for blocks. Things that live outside the profiled estimate: fragmentation in the allocator, communication buffers when tensor parallelism is on, the CUDA context and driver overhead, anything else sharing the card (a metrics exporter, another framework, a stray notebook), and transient spikes on unusual request shapes that the profiling pass did not represent. `1.0` reliably fails; values above ~0.95 are for cards you fully own and workloads you have already profiled. When you must claw back memory, the more surgical levers are lowering `--max-model-len`, shrinking bytes-per-token with `--kv-cache-dtype fp8`, serving a quantized checkpoint, or disabling CUDA graph capture with `--enforce-eager` (which costs decode throughput). ## Co-tenancy Because the value is a per-instance budget, running two vLLM instances on one GPU means setting each to about `0.45` — two instances each claiming `0.9` are each planning to occupy 90% of the same physical card. A non-vLLM process on the card is not represented in vLLM's accounting at all, so co-tenancy with anything else means leaving explicit headroom and validating under load, not at idle. ## What it does not do It does not speed up individual tokens: per-token latency is set by weights, memory bandwidth and batch size, not by how many spare blocks exist. It does not prevent out-of-memory failures caused by the weights themselves — if the weights do not fit under the budget, the server fails immediately and no value of the flag saves it. And it is not a soft ceiling that the server grows into on demand; the block pool is claimed at startup, which is exactly why a vLLM pod's memory graph is flat and boring even as load varies.

  • Where do you read what the flag actually bought you?
    The startup log prints the KV-cache size vLLM settled on, in GiB and in tokens. That token number is the operational capacity: divide it by your typical sequence length and you have the number of concurrent requests the server can hold before it starts preempting. Change the flag, restart, and compare that line — it is far more informative than watching nvidia-smi, which will show the same near-full number either way.
  • Your GPU shows 71 of 80 GB used at idle with no traffic. Is that a leak?
    No — that is the design. vLLM claims its budget at startup and allocates the KV block pool immediately, so resident memory is flat whether the server is idle or saturated. It means device-memory-used is useless as a load or autoscaling signal for vLLM; the occupancy of the block pool is what varies, and vLLM exports that separately as vllm:kv_cache_usage_perc.
  • Weights are 78 GB on an 80 GB card. Does raising utilization to 0.98 make this work?
    Not meaningfully. At 0.98 the budget is about 78.4 GB, which barely covers the weights and leaves nothing for activations, graph capture or KV blocks, so the server either fails to start or starts with a cache too small to hold a single long sequence. That configuration needs a second GPU, a quantized checkpoint, or a smaller model — not a flag change.

saying these in an interview costs you the question

  • Thinks the fraction sizes only the KV cache
  • Sets 1.0 to squeeze out the last gigabyte
  • Believes it is a fraction of currently free memory
  • Expects a higher value to lower per-token latency
  • Leaves 0.9 while another process shares the card

context