skip to content

Quantization for Serving

You will learn how to pick and validate a quantized checkpoint for a production server — which scheme your GPU has fast kernels for, what it does to VRAM and throughput, and how to prove accuracy did not fall off a cliff. Interviewers ask because 'just quantize it' is the reflex cost answer and they want the tradeoffs.

on this pageshow

questions

5

Why can 4-bit weight-only quantization make an LLM server slower at large batch?

level: middleimportance: must knowfreq 60%

answer

  1. bandwidth win, not a compute win
  2. batch size decides the regime
  3. weights narrow, math still sixteen-bit
  4. dequantization costs work each matmul
  5. W8A8 and FP8 narrow the multiply

basics

~20 s

Weight-only 4-bit helps only while decoding is memory-bandwidth-bound. At large batch the server becomes compute-bound, and every matmul must first dequantize weights back to 16-bit — extra work fp16 never pays, so throughput can fall.

solid answer

~50 s

Weight-only schemes (4-bit GPTQ or AWQ) store weights narrow but still do the math in 16-bit: the kernel reads packed weights, dequantizes them into registers, and multiplies against fp16/bf16 activations. At batch size 1 that is a clear win — single-stream decode is limited by how many bytes of weights you drag out of HBM per token, and you just cut that by 4x. As concurrency rises, the same weight read is amortised over many sequences, the GEMMs shift toward compute-bound, and the dequantization becomes pure added work. That is why a 4-bit server can post great single-user latency and *worse* aggregate tokens/s than bf16. Two mitigations: use a kernel built for the regime (vLLM routes compatible GPTQ/AWQ checkpoints to `gptq_marlin`/`awq_marlin`, which are far better at larger batch than the generic kernels), or use a scheme that keeps the multiply narrow — an 8-bit weight-and-activation scheme such as FP8 or compressed-tensors W8A8.

code

bash · 1 line
bash
vllm serve my-org/llama-3.1-8b-awq --quantization awq_marlin --max-num-seqs 256

go deeper

for a junior

Know that a quantized model mainly stores weights smaller, and that smaller weights mean less GPU memory used — not automatically faster answers for many users at once.

for a middle

Explain the bandwidth-bound versus compute-bound split, and that weight-only kernels dequantize back to 16-bit before the multiply. That mechanism is what an interviewer is listening for.

for a senior

Show that you benchmark at production concurrency, check which kernel the engine selected, and choose 8-bit weight-and-activation when the goal is throughput rather than fitting the model.

for a principal

Own the framing that quantization is a capacity lever traded against an accuracy and validation bill, and that the real metric is cost per million tokens at the SLO, not tokens/s in isolation.

## Two different things people mean by "quantized" **Weight-only quantization** stores the model's weights at low precision — 4-bit for GPTQ and AWQ, sometimes 8-bit — but computes in 16-bit. The kernel loads a block of packed 4-bit values plus its scale, expands them back to fp16/bf16 inside the GPU's registers, and feeds that to the tensor cores. The activations flowing through the network were never quantized. **Weight-and-activation quantization** (often written W8A8, and what FP8 does on hardware that supports it) narrows both operands, so the matrix multiply itself runs on 8-bit tensor cores at roughly double the fp16 math rate. That distinction is the whole answer to this question. ## The two regimes of an inference server A transformer decode step is a stack of matrix multiplies. Whether it is limited by bandwidth or by arithmetic depends on how much work is done per byte read. - **Batch size 1 (or very low concurrency).** Each weight matrix is read from HBM and used for exactly one token's worth of arithmetic. The GPU spends nearly all its time waiting on memory. Halving or quartering the bytes of weights read is close to a direct latency win. This is why quantized models feel dramatically faster on a laptop or a single-user demo. - **Large batch.** With 64 or 256 sequences decoding together, one weight read serves 64 or 256 rows of arithmetic. The tensor cores are now the bottleneck. Reading fewer weight bytes no longer buys anything, because bandwidth stopped being the constraint. An LLM server run for throughput lives in the second regime almost all the time — that is the point of continuous batching. ## Why dequantization is not free In the compute-bound regime, a weight-only kernel does *strictly more* work than a bf16 kernel: unpack the 4-bit values, apply the per-group scale (and zero point for asymmetric schemes), reorder if the checkpoint used activation ordering, then run the same 16-bit multiply the bf16 kernel would have run. If that unpacking is not overlapped well with the math, it shows up directly as lower tokens/s. So the honest summary of weight-only 4-bit on a server is: it is a **capacity** feature, not a speed feature. It shrinks the weights so more VRAM is left for KV cache, which lets you run more concurrent sequences — and more concurrency is itself a throughput win. The per-GEMM speedup is a low-batch phenomenon. ## Where Marlin fits The naive 4-bit kernels were written for the batch-1 case and degrade badly as batch grows. Marlin is a mixed-precision GEMM kernel designed to hold its speedup into much larger batches; vLLM exposes it as the `gptq_marlin` and `awq_marlin` quantization methods and will normally select that path automatically for a compatible GPTQ or AWQ checkpoint on Ampere-or-newer hardware. "Compatible" matters: bit width, group size and symmetry all have to be in the supported set, otherwise the engine falls back to the slower generic kernel and you silently lose the benefit. When a 4-bit deployment benchmarks badly, checking which kernel actually got picked (the server logs the chosen quantization method at startup) is the first thing to do. ## Keeping the math narrow instead If your goal is throughput rather than fitting a bigger model, 8-bit is usually the better trade on modern hardware: - **FP8** on GPUs with FP8 tensor cores (Ada and Hopper generations onward) narrows weights *and* activations, so the multiply itself is faster, and the accuracy cost is typically very small. - **W8A8** produced by the compressed-tensors format is the INT8 equivalent; in vLLM you serve it with `--quantization compressed-tensors` (or let the engine read the checkpoint's own config). Note that `W8A8` is a *scheme name inside the checkpoint*, not a value you pass to `--quantization`. ## How to reason about it in an interview State the regime first: "which side of the bandwidth/compute line is this server on?" Then say what the scheme narrows — weights only, or weights and activations. Then name the concrete consequence: weight-only buys VRAM (and therefore KV-cache capacity and concurrency), 8-bit weight-and-activation buys arithmetic throughput. Finally, insist on measuring at *your* production concurrency, not at batch size 1, because the two regimes give opposite answers.

  • If 4-bit does not raise throughput at high batch, why deploy it at all?
    Because it buys VRAM. Weights at 4-bit free tens of gigabytes, and on a fixed card that freed memory becomes KV cache — more concurrent sequences, longer contexts, or a model that otherwise would not fit on the GPU at all. The throughput gain then comes indirectly, from running more sequences per replica, not from faster matmuls.
  • Does weight-only quantization shrink the KV cache too?
    No. Weight quantization touches only the parameters. The KV cache is a separate allocation sized by concurrency, context length and model shape, and it stays in its own dtype. Shrinking it is a separate knob — vLLM exposes `--kv-cache-dtype fp8` (with `fp8_e4m3` / `fp8_e5m2` variants) for that, and it is orthogonal to how the weights are stored.
  • How would you tell which kernel your server actually chose?
    Read the startup logs: the engine reports the resolved quantization method, and vLLM will say when it upgrades a GPTQ or AWQ checkpoint to the Marlin path. If it reports the generic kernel instead, the checkpoint's bit width, group size or symmetry is outside the supported set — re-quantize into a compatible layout rather than accepting the slow path.

Weight-only quantization is like zipping a file you must unzip before every use: it saves the disk read, but once the CPU is the bottleneck the unzipping is pure extra cost.

saying these in an interview costs you the question

  • Assuming 4-bit always makes inference faster
  • Thinking quantization shrinks the KV cache too
  • Benchmarking only at batch size 1
  • Believing 4-bit weights mean 4-bit arithmetic
  • Treating W8A8 and 4-bit weight-only as interchangeable

context

open as a page

Which quantization scheme suits an A10G versus an H100 inference server?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Hardware decides. FP8 tensor cores exist only on Ada and Hopper generations and newer, so an H100 should serve FP8 weights and activations. An Ampere card like the A10G or A100 has no FP8 math and should serve 4-bit GPTQ or AWQ through a Marlin-class kernel, or INT8.

open as a page

Before switching a production endpoint to a quantized model, what do you validate?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Validate both halves of the trade. On quality, compare against the full-precision baseline on real task evals and the failure modes quantization hits hardest. On serving, measure TTFT, inter-token latency and throughput at production concurrency — then canary traffic with the old replica warm for rollback.

open as a page

Should a server load a prequantized checkpoint or quantize weights at startup?

level: middleimportance: should knowfreq 45%

basics

~20 s

Prefer a prequantized checkpoint for production: calibration is already done, the layout matches a fast kernel, and startup stays short. Startup quantization is convenient for experiments but adds load time and usually lands on slower kernels — the exception is FP8, which quantizes online cheaply.

open as a page

Do you standardize one quantization scheme across a serving fleet, or choose per model?

level: principalimportance: should knowfreq 35%

basics

~20 s

Standardize per GPU generation, not across the whole fleet — hardware decides which schemes have fast kernels. Allow per-model exceptions only with evidence, because every model-and-scheme pair carries its own validation, artifact and re-validation-on-upgrade cost.

open as a page