How much GPU memory do a 70B model's weights need at bf16, int8 and int4?
answer
- Start from bytes per parameter
- Two, one, half a byte
- Scales make 4-bit cost more than four
- Weights are a floor, not a total
- 70B at bf16 is 140 GB
basics
~20 sMultiply parameters by bytes per parameter. A 70-billion-parameter dense model needs roughly 140 GB of weights at bf16 (2 bytes), 70 GB at int8 (1 byte) and 35 GB at int4 (half a byte) — before KV cache and activations.
solid answer
~50 sThe core arithmetic is `bytes = parameters x bits_per_parameter / 8`. For a 70B dense model that is about 140 GB at bf16 or fp16, 70 GB at fp8 or int8, and 35 GB at 4-bit. Two numbers make the estimate honest. First, low-bit formats are never exactly N bits per weight: they store a scale (and sometimes a zero-point) per group of weights, so a 4-bit block format lands around 4.25-4.5 effective bits, pushing 70B to roughly 37-40 GB. Second, weights are only the resident floor — the runtime also needs KV cache, activations and allocator overhead, so I budget the weights plus 10-20% before any request memory. That arithmetic is what tells you a 70B model on two 24 GB cards (48 GB total) is only reachable at 4-bit, and even then with a tight context budget.
code
python · 5 linesdef weight_gb(params_billions, bits_per_param):
return params_billions * 1e9 * bits_per_param / 8 / 1e9
for name, bits in [("fp32", 32), ("bf16", 16), ("int8", 8), ("4-bit + scales", 4.25)]:
print(name, round(weight_gb(70, bits), 1), "GB")go deeper
Memorize the four numbers — 4, 2, 1, 0.5 bytes for fp32, fp16/bf16, 8-bit and 4-bit — and be able to multiply them by a parameter count out loud without a calculator.
Explain why real 4-bit files exceed the clean arithmetic: scales and zero-points are stored per block, and some tensors stay at higher precision. Give the effective bits-per-weight rather than the nominal one.
Show the full budget, not just the weights: resident weights, runtime overhead, then KV cache and activations. Be able to say which widths are arithmetically impossible on given hardware and stop the conversation there.
Own the sizing decision as a cost question: how much VRAM you buy, at which width you serve, and what context and concurrency that combination actually supports. Frame width choice as a capacity-planning lever with a quality bill attached, not a default.
## The one rule everything else hangs off A model's weight memory is parameter count times bytes per parameter. Nothing more subtle happens at this level: - fp32: 4 bytes per parameter - fp16 / bf16: 2 bytes - fp8 / int8: 1 byte - fp4 / int4: 0.5 bytes So a 7B model is ~14 GB at bf16, ~7 GB at 8-bit, ~3.5 GB at 4-bit. A 70B model is ~140 GB, ~70 GB, ~35 GB. A 405B model is ~810 GB, ~405 GB, ~203 GB. Interviewers ask this because it is the first filter on any deployment question: before discussing quality or throughput, you have to know whether the thing fits. ## GB versus GiB Marketing memory and allocator memory disagree. 70e9 x 2 bytes = 140e9 bytes = 140 GB in decimal units, but about 130 GiB in binary units, which is what an allocator reports. GPU capacities are quoted in binary-ish terms too (a "24 GB" card exposes roughly 23.5 GiB usable after the driver takes its cut). The 5-7% gap is small enough that it rarely changes a verdict, but say it out loud rather than pretending the numbers are exact — and never plan a deployment that fits with 2% headroom. ## Low-bit formats are not exactly N bits A quantized weight is stored as a small integer or a small float plus a scale that maps it back to a real magnitude. That scale is not free. If a format shares one 8-bit scale across a block of 32 weights, the true cost is 4 + 8/32 = 4.25 bits per weight. Sixteen-weight blocks with an 8-bit scale cost 4.5 bits. Integer schemes that also store a zero-point per group cost more again. Practical consequence: a "4-bit 70B" file on disk is usually 37-42 GB, not 35 GB, and if you sized your GPU on the clean 35 GB you may be short. Some tensors are also commonly left at higher precision — embeddings, the output projection, normalization parameters — which nudges the average up further. Treat the clean arithmetic as a lower bound and check the actual artifact size. ## Weights are the floor, not the total Resident weight memory is what you pay the moment the model loads, with zero requests in flight. On top of it a server needs: - **KV cache**, which grows with the number of tokens being attended to across all live requests - **activations and workspace**, transient buffers for the layer currently executing plus any attention or matmul scratch space - **runtime overhead** — CUDA context, communication buffers for tensor parallelism, and allocator fragmentation A workable planning habit is: weights, plus 10-20% for runtime overhead, and then whatever you deliberately allocate to request memory. If weights alone occupy 95% of the card, you have a model that loads and a server that cannot serve. ## Worked example: 70B on two 24 GB consumer cards Total VRAM is 48 GB, call it ~46 GB usable after driver and context. Walk the widths: - bf16, ~140 GB: not close. You would need six such cards for the weights alone. - int8, ~70 GB: still 1.5x over budget. No. - 4-bit block-scaled, ~37-40 GB with scales: fits, leaving roughly 6-9 GB for KV cache and activations across both cards. So the arithmetic answers the question before any quality argument starts: 4-bit is the only width that is even arithmetically possible here, and the remaining headroom dictates a modest context length and low concurrency. If the workload needed 64K context at batch 8, the honest answer is that this hardware cannot do it and you need more VRAM, not a cleverer format. ## What good candidates add Mention that multi-GPU splits divide weights across devices but replicate some buffers, so two 24 GB cards are not perfectly equal to one 48 GB card. Mention that the file size on disk is a decent sanity check on your estimate. And be explicit that this arithmetic assumes a dense model where every parameter is loaded and used — sparse architectures change the relationship between stored parameters and per-token compute, which is a separate discussion.
- Why is a 4-bit checkpoint on disk usually larger than parameters divided by two?Because a 4-bit weight is meaningless without a scale that maps it back to a real magnitude, and that scale is stored too. One 8-bit scale per 32-weight block adds 0.25 bits per weight; per-16 blocks add 0.5. Some tensors — embeddings, the output head, normalization parameters — are also commonly kept at higher precision. Together these push a nominal 35 GB to roughly 37-42 GB.
- If the weights fit exactly in VRAM with nothing to spare, what happens when you start serving?It fails almost immediately. Weights are the resident floor; every request additionally needs KV cache for its tokens plus transient activation and workspace buffers, and the runtime itself holds a CUDA context and communication buffers. A deployment sized to the weights alone will either refuse to allocate or OOM on the first non-trivial prompt. Plan weights plus overhead plus an explicit request-memory budget.
- Does splitting a model across two GPUs give you exactly the sum of their memory?No. Weight shards divide cleanly, but each device carries its own runtime context, workspace and communication buffers, and some small tensors are replicated rather than split. Two 24 GB cards behave like somewhat less than one 48 GB card, and they add interconnect traffic per layer. Size with a margin rather than assuming perfect additivity.
Bytes-per-parameter is like knowing the weight of a single brick: multiply by the brick count and you instantly know whether the truck can carry the wall, before anyone argues about mortar.
saying these in an interview costs you the question
- Quoting 4-bit as exactly 0.5 bytes per weight with no scale overhead
- Treating weight memory as the total memory a server needs
- Confusing decimal GB with the GiB an allocator reports
- Assuming two 24 GB GPUs equal one 48 GB GPU exactly
- Saying a model 'fits' when weights leave no room for cache