How much GPU memory does a self-hosted Qwen 32B model need beyond its weights?
answer
- weights are only the first line item
- bytes per parameter times parameters
- per-token cost, times tokens in flight
- grouped-query attention shrinks one of them
- the engine pre-allocates what is left over
basics
~20 sWeights are bytes-per-parameter times parameters: about 64 GB for 32B at BF16, near 20 GB at 4-bit. On top sit the KV cache, which grows with context length and concurrency, plus activation and graph buffers — often 10 GB or more.
solid answer
~50 sBudget three things. **Weights**: 2 bytes per parameter at BF16, so a 32B-class Qwen is roughly 64 GB; 8-bit halves that, 4-bit lands near 20 GB including quantisation metadata. **KV cache**: per token it is `2 × layers × kv_heads × head_dim × bytes_per_element`. Qwen3's 32B has 64 layers with 8 grouped-query KV heads of dimension 128, so a token costs about 256 KB at FP16 — roughly 8 GB for one 32k-token sequence, multiplied by concurrent sequences. **Overhead**: activations, CUDA graph capture and the framework itself, commonly several GB. vLLM pre-allocates: `--gpu-memory-utilization` sets the fraction of the card it claims, and whatever is left after weights becomes KV blocks. If startup fails saying the max sequence length exceeds what the KV cache can store, you lower `--max-model-len`, raise the utilisation fraction, quantise the KV cache, or shard with tensor parallelism.
go deeper
Know the weight arithmetic — parameters times bytes per parameter — and that quantisation is what makes a large model fit a single card.
Add the second and third budgets: explain what the KV cache stores, that it grows with tokens in flight, and which startup flags cap context length and concurrency.
Turn it into capacity: derive per-token KV cost from the checkpoint's config, read the allocated cache size from the engine's startup log, and know the ordered fixes when startup fails.
Own the tradeoff surface — quantisation tier against quality, context ceiling against concurrency, one big card against sharding — and set the context limit as a product decision rather than a default.
## Three budgets, not one "Will it fit?" is the first question in self-hosting, and the common mistake is answering it with the weight size alone. GPU memory during serving is spent on three things, and the second one scales with traffic rather than with the model. ## Budget 1 — weights Multiply parameters by bytes per parameter: - **BF16/FP16**: 2 bytes → a 32B model is ~64 GB. - **FP8/INT8**: 1 byte → ~32 GB. - **4-bit**: ~0.5 bytes plus per-group scales and zero points, so in practice ~20 GB rather than a clean 16 GB. This part is fixed for the life of the process. The immediate consequence for a 32B-class Qwen: BF16 does not fit on a 48 GB card at all, and on an 80 GB card it leaves only ~15 GB for everything else — enough to serve, but not enough for long contexts at high concurrency. Which is why 4-bit or 8-bit weights are the norm for single-card deployments of models this size. ## Budget 2 — the KV cache Every token in every active sequence stores a key and a value vector per layer, so it can be attended to without recomputation. The per-token cost is: 2 × num_layers × num_kv_heads × head_dim × bytes_per_element All four numbers are in the checkpoint's `config.json`. For a 32B Qwen3 with 64 layers, 8 KV heads (grouped-query attention shares them across many more query heads) and head dimension 128, at FP16 that is 2 × 64 × 8 × 128 × 2 ≈ 262 KB per token — call it a quarter of a megabyte. That number is the one people forget. A single 32,768-token conversation costs about 8 GB. Sixteen concurrent users each holding an 8k context cost roughly the same again. Grouped-query attention is what keeps this tractable: with 64 query heads and only 8 KV heads, the cache is eight times smaller than full multi-head attention would make it. Controls: `--max-model-len` caps how long any one sequence may get, `--max-num-seqs` caps concurrency, and an FP8 KV-cache dtype halves the per-token cost at some quality risk. ## Budget 3 — overhead Activations for the in-flight batch, CUDA graph capture buffers, the allocator's fragmentation headroom, NCCL buffers when sharding, and the CUDA context itself. Several GB, and it is why a model whose weights plus KV "just fit" on paper does not start. ## How vLLM turns this into a startup outcome vLLM does not allocate lazily. It claims `--gpu-memory-utilization` of the card (default around 0.9), loads the weights, profiles a forward pass to measure activation peak, and turns everything remaining into a fixed pool of KV blocks. It then checks whether `--max-model-len` tokens even fit in that pool for a single sequence, and refuses to start if not — the well-known error that the model's maximum sequence length is larger than the number of tokens the KV cache can hold. The fixes, in the order to try them: 1. **Lower `--max-model-len`.** Most workloads do not need the full native window; capping at what you actually serve is free. 2. **Raise `--gpu-memory-utilization`** if nothing else shares the GPU — going from 0.85 to 0.95 on an 80 GB card is 8 GB of KV. 3. **Quantise the weights.** The largest single lever; 4-bit frees tens of gigabytes on a 32B model. 4. **Quantise the KV cache** to FP8, halving budget 2. 5. **Shard with `--tensor-parallel-size`**, splitting both weights and KV across GPUs — but this needs fast interconnect to be worth it. ## Reading it back from the logs A good habit is to read the startup log rather than to trust arithmetic: vLLM reports the KV cache size it allocated and the number of blocks, which converts directly into "how many total tokens can be in flight at once". Divide by your typical context length and you have your real concurrency ceiling — a far more useful capacity number than any per-model rule of thumb. ## The sizing conversation in interviews What is being tested is whether you know that context length is a memory cost, not just a quality setting. A candidate who sizes a deployment by weights alone will ship something that runs beautifully in a single-user demo and falls over at the third concurrent request.
- Why does grouped-query attention matter so much for self-hosting economics?Because the KV cache scales with the number of *key/value* heads, not query heads. A model with 64 query heads sharing 8 KV heads stores an eighth of the cache that full multi-head attention would need, for the same attention width. That is what makes long contexts and real concurrency affordable on one card, and it is why two models with equal parameter counts can have very different serving costs.
- You have an 80 GB card and need a 32B Qwen at 32k context for about twenty concurrent users. What do you do?BF16 weights alone eat ~64 GB, leaving nowhere near the ~160 GB of KV that twenty 32k sequences would demand, so BF16 on one card is out. Quantise weights to 4-bit (~20 GB), which frees roughly 50 GB of KV — still short of the worst case, so cap `--max-model-len` to what requests actually use, cap concurrency, consider an FP8 KV cache, and measure the allocated block count rather than trusting the arithmetic.
- Does a larger context window cost anything when requests are short?Yes, in a pre-allocating engine: `--max-model-len` is the ceiling the engine must be able to honour for a single sequence, so raising it can shrink the concurrency the pool supports or block startup entirely. Actual per-request cost is proportional to real tokens, but the configured maximum shapes what the server will admit, so set it to what you serve rather than to what the checkpoint supports.
saying these in an interview costs you the question
- Sizing a deployment from the weight file size alone
- Treating context length as a quality knob with no memory cost
- Assuming 4-bit is exactly half of 8-bit with no metadata overhead
- Ignoring that the KV cache scales with concurrent sequences
- Believing more GPU memory always raises throughput linearly