skip to content

Quantization Theory

Quantization is how a model needing well over a hundred gigabytes in full precision ends up running on one GPU or a laptop. This branch is the concept home: number formats and memory math, the algorithms that compress weights, and what that compression costs in quality.

on this pageshow

questions

15

How much GPU memory do a 70B model's weights need at bf16, int8 and int4?

level: juniorimportance: must knowfreq 78%

answer

  1. Start from bytes per parameter
  2. Two, one, half a byte
  3. Scales make 4-bit cost more than four
  4. Weights are a floor, not a total
  5. 70B at bf16 is 140 GB

basics

~20 s

Multiply parameters by bytes per parameter. A 70-billion-parameter dense model needs roughly 140 GB of weights at bf16 (2 bytes), 70 GB at int8 (1 byte) and 35 GB at int4 (half a byte) — before KV cache and activations.

solid answer

~50 s

The core arithmetic is `bytes = parameters x bits_per_parameter / 8`. For a 70B dense model that is about 140 GB at bf16 or fp16, 70 GB at fp8 or int8, and 35 GB at 4-bit. Two numbers make the estimate honest. First, low-bit formats are never exactly N bits per weight: they store a scale (and sometimes a zero-point) per group of weights, so a 4-bit block format lands around 4.25-4.5 effective bits, pushing 70B to roughly 37-40 GB. Second, weights are only the resident floor — the runtime also needs KV cache, activations and allocator overhead, so I budget the weights plus 10-20% before any request memory. That arithmetic is what tells you a 70B model on two 24 GB cards (48 GB total) is only reachable at 4-bit, and even then with a tight context budget.

code

python · 5 lines
python
def weight_gb(params_billions, bits_per_param):
    return params_billions * 1e9 * bits_per_param / 8 / 1e9

for name, bits in [("fp32", 32), ("bf16", 16), ("int8", 8), ("4-bit + scales", 4.25)]:
    print(name, round(weight_gb(70, bits), 1), "GB")

go deeper

for a junior

Memorize the four numbers — 4, 2, 1, 0.5 bytes for fp32, fp16/bf16, 8-bit and 4-bit — and be able to multiply them by a parameter count out loud without a calculator.

for a middle

Explain why real 4-bit files exceed the clean arithmetic: scales and zero-points are stored per block, and some tensors stay at higher precision. Give the effective bits-per-weight rather than the nominal one.

for a senior

Show the full budget, not just the weights: resident weights, runtime overhead, then KV cache and activations. Be able to say which widths are arithmetically impossible on given hardware and stop the conversation there.

for a principal

Own the sizing decision as a cost question: how much VRAM you buy, at which width you serve, and what context and concurrency that combination actually supports. Frame width choice as a capacity-planning lever with a quality bill attached, not a default.

## The one rule everything else hangs off A model's weight memory is parameter count times bytes per parameter. Nothing more subtle happens at this level: - fp32: 4 bytes per parameter - fp16 / bf16: 2 bytes - fp8 / int8: 1 byte - fp4 / int4: 0.5 bytes So a 7B model is ~14 GB at bf16, ~7 GB at 8-bit, ~3.5 GB at 4-bit. A 70B model is ~140 GB, ~70 GB, ~35 GB. A 405B model is ~810 GB, ~405 GB, ~203 GB. Interviewers ask this because it is the first filter on any deployment question: before discussing quality or throughput, you have to know whether the thing fits. ## GB versus GiB Marketing memory and allocator memory disagree. 70e9 x 2 bytes = 140e9 bytes = 140 GB in decimal units, but about 130 GiB in binary units, which is what an allocator reports. GPU capacities are quoted in binary-ish terms too (a "24 GB" card exposes roughly 23.5 GiB usable after the driver takes its cut). The 5-7% gap is small enough that it rarely changes a verdict, but say it out loud rather than pretending the numbers are exact — and never plan a deployment that fits with 2% headroom. ## Low-bit formats are not exactly N bits A quantized weight is stored as a small integer or a small float plus a scale that maps it back to a real magnitude. That scale is not free. If a format shares one 8-bit scale across a block of 32 weights, the true cost is 4 + 8/32 = 4.25 bits per weight. Sixteen-weight blocks with an 8-bit scale cost 4.5 bits. Integer schemes that also store a zero-point per group cost more again. Practical consequence: a "4-bit 70B" file on disk is usually 37-42 GB, not 35 GB, and if you sized your GPU on the clean 35 GB you may be short. Some tensors are also commonly left at higher precision — embeddings, the output projection, normalization parameters — which nudges the average up further. Treat the clean arithmetic as a lower bound and check the actual artifact size. ## Weights are the floor, not the total Resident weight memory is what you pay the moment the model loads, with zero requests in flight. On top of it a server needs: - **KV cache**, which grows with the number of tokens being attended to across all live requests - **activations and workspace**, transient buffers for the layer currently executing plus any attention or matmul scratch space - **runtime overhead** — CUDA context, communication buffers for tensor parallelism, and allocator fragmentation A workable planning habit is: weights, plus 10-20% for runtime overhead, and then whatever you deliberately allocate to request memory. If weights alone occupy 95% of the card, you have a model that loads and a server that cannot serve. ## Worked example: 70B on two 24 GB consumer cards Total VRAM is 48 GB, call it ~46 GB usable after driver and context. Walk the widths: - bf16, ~140 GB: not close. You would need six such cards for the weights alone. - int8, ~70 GB: still 1.5x over budget. No. - 4-bit block-scaled, ~37-40 GB with scales: fits, leaving roughly 6-9 GB for KV cache and activations across both cards. So the arithmetic answers the question before any quality argument starts: 4-bit is the only width that is even arithmetically possible here, and the remaining headroom dictates a modest context length and low concurrency. If the workload needed 64K context at batch 8, the honest answer is that this hardware cannot do it and you need more VRAM, not a cleverer format. ## What good candidates add Mention that multi-GPU splits divide weights across devices but replicate some buffers, so two 24 GB cards are not perfectly equal to one 48 GB card. Mention that the file size on disk is a decent sanity check on your estimate. And be explicit that this arithmetic assumes a dense model where every parameter is loaded and used — sparse architectures change the relationship between stored parameters and per-token compute, which is a separate discussion.

  • Why is a 4-bit checkpoint on disk usually larger than parameters divided by two?
    Because a 4-bit weight is meaningless without a scale that maps it back to a real magnitude, and that scale is stored too. One 8-bit scale per 32-weight block adds 0.25 bits per weight; per-16 blocks add 0.5. Some tensors — embeddings, the output head, normalization parameters — are also commonly kept at higher precision. Together these push a nominal 35 GB to roughly 37-42 GB.
  • If the weights fit exactly in VRAM with nothing to spare, what happens when you start serving?
    It fails almost immediately. Weights are the resident floor; every request additionally needs KV cache for its tokens plus transient activation and workspace buffers, and the runtime itself holds a CUDA context and communication buffers. A deployment sized to the weights alone will either refuse to allocate or OOM on the first non-trivial prompt. Plan weights plus overhead plus an explicit request-memory budget.
  • Does splitting a model across two GPUs give you exactly the sum of their memory?
    No. Weight shards divide cleanly, but each device carries its own runtime context, workspace and communication buffers, and some small tensors are replicated rather than split. Two 24 GB cards behave like somewhat less than one 48 GB card, and they add interconnect traffic per layer. Size with a margin rather than assuming perfect additivity.

Bytes-per-parameter is like knowing the weight of a single brick: multiply by the brick count and you instantly know whether the truck can carry the wall, before anyone argues about mortar.

saying these in an interview costs you the question

  • Quoting 4-bit as exactly 0.5 bytes per weight with no scale overhead
  • Treating weight memory as the total memory a server needs
  • Confusing decimal GB with the GiB an allocator reports
  • Assuming two 24 GB GPUs equal one 48 GB GPU exactly
  • Saying a model 'fits' when weights leave no room for cache

context

open as a page

What is the difference between post-training quantization and quantization-aware training?

level: middleimportance: must knowfreq 72%

basics

~20 s

Post-training quantization rounds an already-trained checkpoint, using a short calibration pass to pick scales. Quantization-aware training simulates that rounding inside a training or fine-tuning run so the weights adapt to it: far more expensive, and worth it mainly at 4 bits and below.

open as a page

Why did bf16 displace fp16 as the default 16-bit format for LLMs?

level: middleimportance: must knowfreq 62%

basics

~20 s

Both formats use 16 bits, but bf16 spends more of them on the exponent and fewer on the mantissa. That gives bf16 the same dynamic range as fp32, so large activations and tiny gradients neither overflow nor vanish — at the cost of coarser precision.

open as a page

Why can perplexity stay flat while a 4-bit quantized model fails your task?

level: middleimportance: must knowfreq 55%

basics

~20 s

Perplexity averages next-token loss over a whole corpus, so damage concentrated on a few decisive tokens disappears into the mean. Task accuracy, structured-output validity and per-slice scores expose quantization damage that a flat perplexity curve hides.

open as a page

In GGUF, what does the K in a k-quant such as Q4_K_M mean?

level: juniorimportance: should knowfreq 35%

basics

~20 s

The K marks llama.cpp's k-quant family, where weights sit in super-blocks whose per-block scales and minimums are themselves stored in low precision. The trailing S or M says how many of the most sensitive tensors are promoted to a higher-bit k-quant; the 4 is the base width.

open as a page

Why are activations harder to quantize than weights in a transformer?

level: middleimportance: should knowfreq 56%

basics

~20 s

Weights are fixed after training, so their ranges can be measured once, offline. Activations change with every input and contain a few hidden channels whose magnitudes are far above the rest, so a single scale either clips those channels or wastes almost all the range on them.

open as a page

In 4-bit block-scaled formats like MXFP4, what does the shared per-block scale buy?

level: middleimportance: should knowfreq 42%

basics

~20 s

Four bits can only encode about sixteen distinct values, which is far too coarse for a whole tensor. A scale shared by a small block of weights re-centres those sixteen levels on that block's own magnitude range, so a large weight elsewhere in the tensor cannot flatten this block to zero.

open as a page

Why can two 4-bit LLM builds differ sharply in quality at the same bit width?

level: middleimportance: should knowfreq 46%

basics

~20 s

Bit width is a container, not a quality level. How finely the format shares scaling factors, whether the checkpoint was trained at low precision or converted afterwards, and what else got quantized alongside the weights all matter more than the number four.

open as a page

How do you choose calibration data for post-training quantization?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Use a few hundred samples that look like production traffic: same domain, same prompt format, realistic lengths. Calibration only sets ranges and error statistics, so a corpus that misses the real distribution produces a model tuned for text it will never see.

open as a page

How do GPTQ and AWQ differ when quantizing the same weights to 4-bit?

level: seniorimportance: should knowfreq 50%

basics

~20 s

GPTQ quantizes a layer's weights in sequence and uses second-order information from calibration activations to nudge the not-yet-quantized weights so they absorb each rounding error. AWQ instead searches for per-channel scale factors that protect the small fraction of channels the activations show to matter most.

open as a page

At 8 bits, what differs between int8 and fp8 for representing model tensors?

level: seniorimportance: should knowfreq 38%

basics

~20 s

int8 lays 256 values on an evenly spaced grid set by a scale, and optionally a zero-point offset. fp8 spends its 8 bits on an exponent and mantissa, so its values crowd near zero and spread out at large magnitudes — a much better match for heavy-tailed tensors.

open as a page

Beyond weights, what else consumes GPU memory when serving an LLM, and when does it dominate?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Three consumers share the card: weights, which are fixed once loaded; the KV cache, which grows with tokens in flight across all live requests; and transient activations plus runtime overhead. At long context and high concurrency the cache routinely exceeds the weights.

open as a page

Why do 4-bit serving recipes keep certain layers at higher precision?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Quantization error is not spread evenly. A few tensors carry extreme activation outliers or sit where error cannot be absorbed downstream, so squeezing them costs far more accuracy than the memory it saves. Recipes leave those in higher precision and push the bulk to 4-bit.

open as a page

Does a frozen 4-bit base cap what a QLoRA adapter can learn?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Partly. Higher-precision trainable adapters absorb much of the base's quantization error, and QLoRA was reported to match 16-bit fine-tuning on its benchmarks. But capabilities the 4-bit base lost outright — a thinly represented language, say — a small adapter cannot restore.

open as a page

How do you decide whether 4-bit serving is worth its quality cost?

level: principalimportance: should knowfreq 34%

basics

~20 s

Measure both sides on your own workload and hardware: memory freed and tokens per second gained against task accuracy lost on a sliced eval. Then decide per product surface, because a summarizer and a payments agent tolerate very different regressions.

open as a page