Will a 70B model fit on 2xA100-80GB at 32k context? How do you check?
answer
- Four terms, one card
- weights are fixed, cache grows
- bytes per token times context times users
- GPU memory does not pool
- fits does not mean serves
basics
~20 sAdd the four terms up. 70B weights at bf16 are about 140 GB of the 160 GB on the node, so after CUDA context and activations only a few GB are left for KV cache. It loads, but it serves almost no concurrency.
solid answer
~50 sSizing is arithmetic, not intuition: **weights + KV cache + activations + framework overhead** against the card's real capacity. Weights are bytes-per-parameter times parameter count — 70B at bf16 is roughly 140 GB. Per-GPU CUDA context, communication buffers and activation workspace cost another 1-3 GB per device. On a 2xA100-80GB node that leaves single-digit gigabytes for KV cache. For a Llama-3-70B shape (80 layers, 8 grouped KV heads, head dim 128) fp16 KV runs about 320 KB per token, so a single 32k-token session wants roughly 10 GB. So the honest answer is "it loads and then serves about one user" — which is a no. The fixes are to shard wider (4 GPUs), serve an FP8 or 4-bit checkpoint so weights drop to ~70 GB or ~35 GB, or cap the served context far below 32k. Say the numbers out loud; interviewers are testing whether you can.
code
python · 10 linesdef budget_gb(params_b, bytes_per_param, layers, kv_heads, head_dim,
ctx_len, concurrency, gpu_gb, n_gpus, overhead_gb=2.0):
weights = params_b * 1e9 * bytes_per_param / 1e9
kv_per_token = 2 * layers * kv_heads * head_dim * 2 # K and V, fp16
kv = kv_per_token * ctx_len * concurrency / 1e9
needed = weights + kv + overhead_gb * n_gpus
return round(needed, 1), gpu_gb * n_gpus
# 70B bf16, Llama-3-70B shape, four concurrent 32k sessions, 2 x 80GB
print(budget_gb(70, 2, 80, 8, 128, 32768, 4, 80, 2)) # (186.9, 160)go deeper
Know the four things that occupy GPU memory — weights, KV cache, activations, framework overhead — and that weights alone never tell you whether a deployment works.
Be able to do the arithmetic out loud: bytes per parameter times parameters, KV bytes per token times context times concurrency, then compare against a named card. Practice it once on a real model config.
Turn the budget into a served configuration: how many concurrent sessions at what context length, what you cap, and what you shard. Expect to defend why you chose four GPUs instead of two.
Own the tradeoff between capacity and promise. Advertised context length, concurrency targets and checkpoint precision are all product commitments priced in VRAM; decide which ones you are willing to sell before the cluster is bought.
## Why this is the signature question of GPU serving Every deployment decision downstream — which card, how many, what price per token, whether autoscaling is even possible — starts with whether the thing fits. The calculation takes two minutes and most candidates have never done it, which is exactly why it gets asked. ## The four terms of a VRAM budget **1. Weights.** Bytes per parameter times parameter count. bf16/fp16 is 2 bytes, FP8 or INT8 is 1 byte, 4-bit is roughly 0.5 bytes plus a small per-group scale overhead. A 70B model is therefore about 140 GB, 70 GB, or 36-40 GB respectively. This term is fixed for the life of the process and is split across GPUs when you shard. **2. KV cache.** This is the term that grows with traffic, and the one that decides how many users fit. Bytes per token equal `2 (K and V) x layers x kv_heads x head_dim x dtype_bytes`. For a Llama-3-70B shape that is `2 x 80 x 8 x 128 x 2` = 327,680 bytes, about 320 KB per token. Multiply by context length for one full session (32,768 tokens -> ~10.7 GB) and again by how many such sessions you want resident at once. Grouped-query attention is what makes that number survivable; a model with 64 full KV heads instead of 8 would cost eight times as much. **3. Activations and workspace.** Transient tensors for the forward pass, attention kernel scratch space, and the buffers the engine keeps for communication. The peak is driven by the largest prefill you allow — a single 32k-token prompt materializes far bigger intermediates than a 200-token one — which is why prompt-length caps are a memory control, not just a product decision. **4. Framework overhead.** The CUDA context alone is several hundred megabytes to over a gigabyte per device; the allocator, NCCL buffers for multi-GPU collectives, CUDA graphs, and the Python process add more. Budget 1-3 GB per GPU before any model bytes. ## Working the example Node: 2 x A100-80GB = 160 GB total, and note that memory does **not** pool across GPUs. Tensor-parallel sharding splits weights and KV cache evenly, so each card carries about 70 GB of weights out of its 80. Subtract ~2 GB of context and workspace: roughly 8 GB per card, 16 GB total, for KV cache. At 320 KB per token that is about 50,000 cached tokens — one and a half 32k sessions. Technically it starts; practically it is a single-user demo box, and any second long request queues or gets preempted. The interview-grade answer names the verdict *and* the fixes: - **Shard wider.** 4 x A100-80GB puts 35 GB of weights on each card and frees ~150 GB total for KV, which is a real serving configuration. - **Change the checkpoint.** FP8 weights halve the weight term to ~70 GB and roughly triple the KV headroom on the same two cards; 4-bit halves it again. That is a quality decision you must then validate, not a free win. - **Cap what you promise.** Serving 8k context instead of 32k cuts the per-session cache by 4x. Most products do not need the advertised maximum. - **Bigger cards.** H200 at 141 GB per device changes the arithmetic materially for exactly this class of model. ## Headroom, and how engines express it Servers do not allocate lazily. They measure free memory at startup, subtract weights and a profiled activation peak, and claim the rest as a fixed KV pool — which is why a healthy server queues or preempts under load instead of dying. vLLM exposes the fraction it is allowed to take as `--gpu-memory-utilization` (0.9 by default); TensorRT-LLM's Triton backend uses `kv_cache_free_gpu_mem_fraction`; TGI derives its cache from the token caps you set, such as `--max-total-tokens`. Never target 1.0: kernels, collectives and the driver need room outside your allocations. ## How to say it in an interview State the four terms, put a number on each, compare against the card, then give the verdict and two remedies. Refusing to guess is fine — "I'd measure the KV bytes per token from the config rather than assume" is a strong sentence — but refusing to do arithmetic is not.
- Where do you get the numbers for the KV-cache formula on a model you have never served?From the model's own config: number of hidden layers, `num_key_value_heads` (not `num_attention_heads` — grouped-query models cache far fewer), and head dimension, times 2 for K and V, times the cache dtype width. Then sanity-check against the server's own startup log, which reports how many KV blocks or tokens it was able to allocate.
- Two 80 GB GPUs is 160 GB. Why can't a 150 GB model just spill across them?Because memory is per-device, not pooled. Sharding is what splits a model, and it splits *every* term — each GPU holds its slice of the weights and its slice of the KV cache, and every layer's collective runs over the interconnect. A model only "fits" if each shard fits its own card with room for cache and workspace.
- How does the answer change if you cut the served context from 32k to 4k?The per-session KV cost drops eightfold — about 1.3 GB instead of 10.7 GB for that model shape — so the same 16 GB of headroom now holds a dozen concurrent sessions instead of one. Context length is the cheapest capacity lever you own, and it costs nothing in quality for workloads whose prompts were never long.
saying these in an interview costs you the question
- Only counting model weights and calling it sized
- Assuming two GPUs give one 160 GB pool
- Ignoring KV cache growth with concurrency and context
- Targeting 100% of VRAM with no headroom
- Saying 'it loaded' as proof the sizing works