skip to content

How do you size GPU VRAM to serve Llama 3.1 8B at long context?

level: seniorimportance: should knowfreq 56%

answer

  1. three consumers: weights, KV cache, overhead
  2. two bytes per parameter at fp16
  3. KV scales with context times concurrency
  4. grouped-query attention shrinks the KV term
  5. the memory fraction budgets, it does not create

basics

~20 s

Budget three things: weights (about 16 GB for 8B at fp16), KV cache, and activation/framework overhead. Llama 3.1 8B costs roughly 128 KiB of fp16 KV per token, so a single 128k-token context needs about 16 GB of cache on its own.

solid answer

~60 s

Weights first: parameters times bytes per parameter. An 8B model is ~16 GB at fp16, ~8 GB at 8-bit, ~5 GB at 4-bit. Then the KV cache, which is the part people forget. Per token it costs `2 (K and V) x num_kv_heads x head_dim x num_layers x bytes`. Llama 3.1 8B has 32 layers and, thanks to grouped-query attention, only 8 KV heads of dimension 128 — so 2 x 8 x 128 x 32 x 2 bytes = 128 KiB per token. A full 128k context is therefore ~16 GB of KV for **one** sequence; ten concurrent 8k conversations cost ~10 GB. Third, leave headroom for activations, CUDA graphs and framework overhead — a couple of GB. Operationally, vLLM's `--gpu-memory-utilization` (default 0.9) sets the total fraction it may claim, and whatever remains after weights and overhead becomes the KV block pool. If that pool cannot hold one sequence at `--max-model-len`, the server refuses to start and tells you to raise the fraction or lower the context length. On a 24 GB card, 8B at fp16 plus full 128k context simply does not fit — you cap `--max-model-len` or quantize.

code

python · 10 lines
python
def kv_bytes_per_token(num_layers, num_kv_heads, head_dim, dtype_bytes=2):
    return 2 * num_kv_heads * head_dim * num_layers * dtype_bytes

# Llama 3.1 8B
per_token = kv_bytes_per_token(32, 8, 128)
print(per_token / 1024, "KiB/token")            # 128.0
print(per_token * 128_000 / 1e9, "GB @128k")    # ~16.8

# Llama 3.1 70B
print(kv_bytes_per_token(80, 8, 128) / 1024, "KiB/token")  # 320.0

go deeper

for a junior

Know that VRAM holds the weights plus a growing cache of past tokens, and that roughly two bytes per parameter gives the fp16 weight size.

for a middle

Reproduce the per-token KV formula from layers, KV heads, head dimension and dtype, and explain why context length multiplied by concurrency drives the memory budget.

for a senior

Size a real deployment end to end: derive the pool from traffic, choose max context and concurrency caps, and diagnose start-up refusals versus runtime preemption from the symptoms.

for a principal

Own the policy consequences — what maximum context you offer customers, whether quantization is a capacity decision rather than a quality one, and how GPU class and count follow from the token-throughput forecast.

## The three consumers of VRAM **1. Weights.** Deterministic: parameter count times bytes per parameter, plus a little for embeddings and layout. Llama 3.1 8B is ~16 GB at fp16/bf16, ~8 GB at 8-bit, ~4.5–5 GB at 4-bit. This cost is fixed the moment the server starts. **2. KV cache.** Proportional to tokens held in flight — context length times concurrency. This is the variable cost and the one that determines how much traffic the GPU can carry. **3. Activations and overhead.** Transient tensors for the current forward pass, CUDA graphs, the framework's own allocations, and NCCL buffers when sharding. Budget a couple of GB and measure rather than theorise. ## The KV formula, worked Per token, per layer, the cache stores one key vector and one value vector for each key/value head: ``` bytes_per_token = 2 (K,V) x num_kv_heads x head_dim x num_layers x bytes_per_element ``` For Llama 3.1 8B: `num_layers = 32`, `num_kv_heads = 8`, `head_dim = 128`, fp16 = 2 bytes. `2 x 8 x 128 x 32 x 2 = 131,072 bytes = 128 KiB per token.` Consequences worth memorising: - 8k context ≈ 1 GB per sequence. - 128k context ≈ 16 GB per sequence — as much as the weights. - 100 concurrent 4k conversations ≈ 50 GB. Grouped-query attention is why this is survivable. Llama 3.1 8B has 32 attention heads but only 8 KV heads; with full multi-head attention the cache would be four times larger, and 128k context would be a non-starter on any single GPU. The arithmetic changes per model: Llama 3.1 70B has 80 layers with the same 8 KV heads of dim 128, giving 320 KiB per token. ## Turning the arithmetic into a deployment Start from traffic, not from the card. Estimate p95 prompt length, expected output length, and target concurrent sequences. Multiply average total tokens per sequence by bytes-per-token by concurrency to get the KV budget, add weights, add ~2 GB overhead, and compare against the card. Worked example on an 80 GB GPU with 8B at fp16: 16 GB weights + ~2 GB overhead leaves ~54 GB of KV after the 0.9 utilisation cap — roughly 430k cached tokens, i.e. about 100 concurrent sequences averaging 4k tokens, or a handful of very long contexts. On a 24 GB card the same model leaves only ~3–4 GB of KV, enough for maybe 25–30k tokens total: fine for a few short chats, hopeless for long context. Quantizing weights to 4-bit on that card frees ~11 GB and roughly quadruples the usable KV pool — which is often the real reason to quantize when serving, independent of any quality argument. ## The knobs and what they actually do - `--gpu-memory-utilization` (vLLM, default 0.9): the fraction of total VRAM the server may claim. It does **not** create memory; it decides how much of the card vLLM profiles and preallocates. Raise it only when nothing else shares the GPU, since going too high causes out-of-memory during peak activation. - `--max-model-len`: the longest sequence the server accepts. It sets the floor on KV pool size (the pool must hold at least one such sequence) and is the fastest lever when startup fails. - `--max-num-seqs`: caps concurrent sequences, bounding worst-case KV demand and per-stream latency. - KV-cache quantization (fp8 in vLLM, `--cache-type-k`/`--cache-type-v` in llama.cpp): halves KV bytes per token, doubling effective capacity at some quality cost. In llama.cpp and Ollama the shape differs: `--ctx-size` is the *total* context budget shared across `--parallel` slots, so `-np 4 -c 8192` gives each slot 2048 tokens — an easy trap when requests start getting truncated for no obvious reason. ## Failure signatures A server that refuses to start with a message about needing more KV cache than is available is telling you `max_model_len` is too large for the memory left after weights — lower it, raise the utilisation fraction, or quantize. A server that starts fine but shows rising preemption counters and latency spikes under load is running out of blocks at runtime: the pool is sized for the model but not for the traffic. An out-of-memory crash at peak, with a utilisation fraction near 1.0, usually means activation memory had nowhere to go. ## What interviewers want They want to see you compute rather than guess: name the three consumers, produce the per-token KV formula, plug in real config values, and connect the result to a concrete deployment decision — cap the context, quantize, or buy a bigger card. Candidates who only recall "about 2 GB per billion parameters at fp16" have half the answer and will size a long-context service wrong by an order of magnitude.

  • vLLM refuses to start, saying the KV cache needed exceeds what is available. What are your options?
    The memory left after weights and overhead cannot hold one sequence at the configured maximum length. Lower `--max-model-len` to what your traffic actually needs; raise `--gpu-memory-utilization` if nothing else shares the card; quantize the weights to free several gigabytes; enable fp8 KV cache to halve per-token cost; or shard across GPUs. Lowering the context length is almost always the fastest correct fix.
  • Why does grouped-query attention matter so much for serving memory?
    KV cache size scales with the number of key/value heads, not query heads. Llama 3.1 8B has 32 query heads but only 8 KV heads, so its cache is a quarter the size of an equivalent full multi-head model — 128 KiB per token instead of 512 KiB. That is the difference between a 128k context costing 16 GB and costing 64 GB, which is what makes long context servable on one GPU at all.
  • Does raising --gpu-memory-utilization to 0.98 give you more concurrency for free?
    Only briefly. The fraction governs how much VRAM vLLM profiles and preallocates, so a higher value does enlarge the KV pool — but it leaves less slack for peak activation memory, CUDA graphs and any other process on the card, and you trade a clean capacity limit for out-of-memory crashes under load. Values around 0.90–0.95 on a dedicated GPU are the usual safe range.
  • How does the memory picture change in llama.cpp compared with vLLM?
    llama.cpp allocates a total context budget with `--ctx-size` and divides it across parallel slots, so four slots against an 8192 budget give each request 2048 tokens — requests get truncated with no obvious error. It also lets you offload only some layers to the GPU with `--n-gpu-layers`, trading speed for fitting on small VRAM, and supports quantized K and V cache types to shrink the per-token cost.

saying these in an interview costs you the question

  • Sizing only for weights and forgetting the KV cache entirely
  • Treating the GPU memory fraction as if it added memory
  • Assuming maximum context costs nothing until someone uses it
  • Ignoring that KV scales with concurrency, not just context length
  • Using a KV formula based on query heads instead of key/value heads

context