How much memory do Llama 70B weights need at FP16 versus 4-bit quantization?
answer
- Bytes equal params times bits over eight
- Two bytes per parameter at FP16
- Roughly half a gigabyte per billion at 4-bit
- Overhead pushes 4-bit above nominal
- KV cache is not weight memory
basics
~20 sWeights cost bytes-per-parameter times parameter count: about 140 GB for 70B at FP16, about 70 GB at 8-bit, and roughly 35-40 GB at 4-bit once block scales are counted. KV cache and activations are extra and are not shrunk by weight quantization.
solid answer
~50 sStart from the arithmetic: **memory ≈ parameters × bits per weight ÷ 8**. FP16 and BF16 are two bytes per parameter, so a 70B model is about 140 GB of weights. Int8 halves that to roughly 70 GB. A nominal 4-bit build is about 35 GB in theory, but real formats store a scale (and sometimes a minimum) per block and often keep sensitive tensors at higher precision, so a GGUF Q4_K_M or a 4-bit GPTQ build of 70B lands nearer 40 GB. The published file size is the honest number — use it, not the nominal bit width. Then add what quantizing weights does **not** touch: the KV cache, which grows with context length and concurrency, plus activations, CUDA context and fragmentation. A rough working rule is to leave 15–25 % headroom above the weights before you claim a model fits. That is why a 70B 4-bit build is comfortable on 48 GB but tight on a single 40 GB card.
code
python · 6 linesdef weight_gib(params_billion: float, bits_per_weight: float) -> float:
return params_billion * 1e9 * bits_per_weight / 8 / 1024**3
print(round(weight_gib(70, 16), 1)) # 130.4 GiB (~140 GB)
print(round(weight_gib(70, 8), 1)) # 65.2 GiB
print(round(weight_gib(70, 4.8), 1)) # 39.1 GiB effective Q4_K_M ratego deeper
Be able to do the arithmetic out loud: two bytes per parameter at FP16, so 70B is about 140 GB, and roughly a quarter of that at 4-bit. Say that KV cache is additional.
Explain why real 4-bit builds exceed the nominal width — per-block scales and higher-precision sensitive tensors — and quote effective bits per weight rather than the label.
Show you size a deployment end to end: weights from the artifact size, plus KV cache scaled by context and concurrency, plus activations and headroom, and know which of those quantization actually addresses.
Own the capacity model — cost per served token across GPU classes and quant levels, when buying more memory beats accepting a lossier quant, and how context-length commitments drive the cache budget.
## The core formula The dominant memory cost of a language model at inference is its weights, and the arithmetic is elementary: `bytes ≈ parameter_count × bits_per_weight / 8` So for a 70-billion-parameter Llama: | Precision | Bits/weight | Weights | |---|---|---| | FP32 | 32 | ~280 GB | | FP16 / BF16 | 16 | ~140 GB | | INT8 / Q8_0 | 8–8.5 | ~70–75 GB | | 4-bit (Q4_K_M, GPTQ-4, AWQ-4) | 4.5–4.8 effective | ~40 GB | | 3-bit | ~3.4–3.9 effective | ~32 GB | And for an 8B model, divide by roughly 8.75: ~16 GB at FP16, ~4.9 GB at Q4_K_M. This is why 8B fits comfortably on consumer hardware while 70B does not without quantization. ## Why 4-bit is never exactly 4 bits Every practical 4-bit format stores metadata alongside the packed weights. GGUF k-quants keep a quantized scale, and for some types a minimum, per small block; GPTQ and AWQ keep a scale and zero-point per group of 128 weights; bitsandbytes NF4 keeps per-block constants (which double quantization then compresses). On top of that, the better mixes deliberately keep sensitive tensors — embeddings, the output projection, attention value projections, feed-forward down projections — at higher precision. The result is an *effective* bits-per-weight above the nominal number, typically 4.5–4.9 for a good 4-bit build. The practical consequence: never size hardware from `params × 4 / 8`. Read the artifact's file size, which already includes all of that overhead, and treat it as a close lower bound on the resident weight memory. ## What weight quantization does not shrink This is the part candidates most often miss. - **KV cache.** Every token you have processed leaves cached key and value tensors for every layer. That memory scales with sequence length and with the number of concurrent sequences, and it is stored in the compute dtype unless you separately enable KV-cache quantization. On a long-context, high-concurrency workload it can rival the weights. - **Activations.** Transient per-forward-pass buffers, larger with bigger batches. - **Runtime overhead.** The CUDA context, the framework's allocator behaviour, memory fragmentation, and any offload buffers. So "the file is 40 GB and my card has 40 GB" is not a plan. Budget headroom — 15–25 % over the weights is a reasonable starting point for modest contexts, more if you serve long prompts or many concurrent users. ## Worked examples - **Llama 8B on a 16 GB GPU.** FP16 weights are ~16 GB — it does not fit with any room for cache. Q8_0 at ~8.5 GB or a 4-bit build at ~4.9 GB both fit with comfortable headroom. - **Llama 70B on one 48 GB GPU.** FP16 (~140 GB) is out of the question. 8-bit (~70 GB) still does not fit. A 4-bit build at ~40 GB fits, leaving roughly 8 GB for KV cache and activations — workable for a modest context, tight for long ones. - **Llama 70B across two 80 GB GPUs.** FP16 at ~140 GB fits with tensor parallelism and leaves real room for cache; this is the configuration where you would not quantize at all unless you needed the memory for batch size. ## Multiplying out for other precisions A handy mental shortcut: at FP16 a model needs about **2 GB per billion parameters**; at 8-bit about **1 GB per billion**; at 4-bit about **0.55 GB per billion** once overhead is included. Multiply by the parameter count, add headroom, and you have a size estimate in ten seconds — which is exactly the calculation an interviewer is checking you can do without a calculator. ## Two caveats worth stating First, mixture-of-experts models break the intuition that memory tracks compute: all experts must be resident even though only a few are active per token, so size the memory by *total* parameters, not active ones. Second, quantization changes memory sharply but changes speed only indirectly — decoding is largely memory-bandwidth bound, so moving fewer bytes per token usually helps, but the gain depends on having a good kernel for that format.
- Why does a 4-bit 70B build measure around 40 GB rather than 35 GB?Because no format stores exactly four bits per weight. Block or group scales — and zero-points or minimums, depending on the format — are stored alongside the packed weights, and good mixes deliberately keep embeddings, the output projection and a few sensitive per-layer tensors at higher precision. That pushes the effective rate to roughly 4.5–4.8 bits per weight. Size from the published file size, not from the nominal bit width.
- What grows your memory use as you raise the context length, if the weights are fixed?The KV cache. Each processed token leaves cached key and value tensors in every layer, so cache memory scales with sequence length times concurrent sequences times layers times head dimension. Weight quantization does nothing for it — it is a separate pool, stored in the compute dtype unless you explicitly enable KV-cache quantization. On long-context or high-concurrency serving it can grow to rival the weights, which is why headroom above the model file matters.
- How should you size memory for a mixture-of-experts model?By total parameters, not active ones. All experts must be resident in memory even though only a small subset is routed per token, so a model with 671B total and 37B active still needs memory proportional to 671B at whatever precision you chose. Active parameters predict compute and latency; total parameters predict memory. Conflating the two is the classic sizing error on MoE deployments.
saying these in an interview costs you the question
- Sizing 4-bit as exactly params times four bits
- Forgetting the KV cache in the memory budget
- Assuming a 40 GB file fits a 40 GB card
- Thinking quantization shrinks activations too
- Sizing an MoE model by active parameters