skip to content

In vLLM, when do you pass --quantization, and which values does it accept?

level: seniorimportance: should knowfreq 40%

answer

  1. method, not bit width
  2. the checkpoint usually already says
  3. a fixed literal value set
  4. one popular string is not a value
  5. two methods now refuse to load

basics

~20 s

Usually you pass nothing: vLLM reads the checkpoint's own quantization config and picks a kernel. You pass --quantization to force a specific kernel path, to quantize an unquantized checkpoint at load time, or to name a method the checkpoint does not declare. Values come from a fixed list.

solid answer

~40 s

The flag names a *method*, not a bit width, and it is drawn from a closed set — `awq`, `gptq`, `gptq_marlin`, `awq_marlin`, `compressed-tensors`, `fp8`, `modelopt`, `bitsandbytes`, `mxfp4`, `quark`, `torchao`, `experts_int8`, `moe_wna16` and online shorthands such as `fp8_per_tensor` and `nvfp4_per_token`. For a prequantized checkpoint you normally omit it: vLLM inspects the config and selects a kernel, often the faster Marlin path on supported hardware. You reach for it to pin a kernel, to quantize a 16-bit checkpoint online (`--quantization fp8`), or when a checkpoint's metadata is ambiguous. Two traps: `W8A8` is not a value — it is a compressed-tensors scheme served with `--quantization compressed-tensors` — and in 0.27 `fbgemm_fp8` and `fp_quant` are deprecated and raise unless you also pass `--allow-deprecated-quantization`.

code

bash · 11 lines
bash
# Prequantized checkpoint: let vLLM read the config and pick the kernel
vllm serve TheBloke/Llama-2-13B-chat-GPTQ

# Explicit method for a compressed-tensors checkpoint (this is how W8A8 is served)
vllm serve neuralmagic/Meta-Llama-3-8B-Instruct-quantized.w8a8 \
  --quantization compressed-tensors

# Quantize a 16-bit checkpoint at load time, and shrink the KV cache separately
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --quantization fp8 \
  --kv-cache-dtype fp8

go deeper

for a junior

Know that vLLM usually detects a prequantized checkpoint on its own, and that --quantization names a method from a fixed list rather than a number of bits.

for a middle

Explain the inference path — the checkpoint's config selects the method and often a faster kernel variant — and name the legitimate reasons to override: pinning a kernel, quantizing online, ambiguous metadata.

for a senior

Be precise about the traps: W8A8 is a compressed-tensors scheme, not a flag value, and deprecated methods now raise unless explicitly allowed. Pair any change with a task-level quality comparison, not a smoke test.

for a principal

Set the standard: which quantization formats your fleet supports, who produces and validates those checkpoints, and how a serving-engine upgrade that deprecates a method is caught before it reaches a production launch script.

## What the flag actually selects `--quantization` (short form `-q`) tells vLLM which quantization **method implementation** to use for the weights it is about to load. It does not convert a checkpoint from one format to another, and it does not pick a bit width by itself. It selects the code path that knows how to interpret the stored tensors and which kernels to run against them. The accepted values are a fixed literal set, and passing anything outside it fails at argument parsing. In vLLM 0.27 the set includes `awq`, `auto_awq`, `gptq`, `auto_gptq`, `gptq_marlin`, `awq_marlin`, `compressed-tensors`, `fp8`, `modelopt` and `modelopt_fp4`, `bitsandbytes`, `experts_int8`, `moe_wna16`, `mxfp4`, `quark`, `torchao`, `inc`, and a set of online shorthands such as `fp8_per_tensor`, `fp8_per_block`, `int8_per_channel_weight_only`, `nvfp4_per_token` and `mxfp8`. ## Why the usual answer is "don't pass it" A prequantized checkpoint carries its own quantization configuration. vLLM reads it and chooses the method, and where a faster kernel exists for the same format on the current hardware it will generally prefer it — GPTQ checkpoints served through the Marlin kernels being the canonical example. That inference is a feature: it means the checkpoint, not the launch script, is the source of truth about how the weights are stored, and it means you do not have to encode per-model knowledge in your deployment templates. The legitimate reasons to override are narrower than people assume: - **Pin a kernel path.** Forcing `gptq` instead of the Marlin variant, for instance, when debugging a numerical discrepancy or working around a kernel that misbehaves on your hardware. - **Quantize online.** Passing `fp8` (or one of the online shorthands) against a 16-bit checkpoint makes vLLM quantize the weights during load. This costs start-up time and gives up the quality control that a calibrated offline pipeline provides, but it needs no separate artifact. - **Disambiguate.** Some checkpoints are packaged without complete metadata, and the flag is how you tell vLLM what it is holding. ## The two traps worth naming **`W8A8` is not a value.** It is a *scheme* — 8-bit weights and 8-bit activations — expressed inside a compressed-tensors checkpoint. You serve such a checkpoint by letting vLLM read its config, or explicitly with `--quantization compressed-tensors`. Candidates who type `--quantization w8a8` are describing the numerics rather than the method, and vLLM rejects it. The strings `fp8_w8a8` and `int8_w8a8` do appear inside vLLM, but as names of kernel tuning configuration files, not as CLI values. **Deprecated methods now fail closed.** In 0.27, `fbgemm_fp8` and `fp_quant` are deprecated: launching with them raises unless you also pass `--allow-deprecated-quantization`. This is a good pattern to be able to discuss — a deprecation that stops the server rather than silently degrading it forces the operator to make a decision at a known moment instead of discovering a changed numerical path in production. It also means an upgrade can break a launch script that worked on the previous version, which is one concrete argument for pinning the image version and reading release notes for a serving component. ## The neighbouring flag `--kv-cache-dtype` is separate and orthogonal: it quantizes the KV cache rather than the weights, and accepts `auto` (match the model dtype), `fp8`, `fp8_e4m3` and `fp8_e5m2`. You can run 4-bit weights with a 16-bit cache, or 16-bit weights with an fp8 cache; they solve different problems — weights quantization shrinks the fixed cost, cache quantization buys concurrency and context. Being clear that these are two different budgets is a reliable senior signal. ## Verifying you got what you asked for The startup log states the quantization method in effect and the KV-cache size that resulted, and a failure to match method to checkpoint typically surfaces as a load-time error about unexpected tensor shapes or missing scale tensors rather than as degraded output. The dangerous outcome is not a crash but a silent difference in numerics after a version upgrade or a method override — which is why a quantized deployment needs a task-level check before and after any change to these flags, not just a smoke test that the server answers.

  • Why does vLLM raise on fbgemm_fp8 rather than warning and continuing?
    Because a silent fallback would change the numerics of a production endpoint at upgrade time with no signal. Failing closed puts the decision in front of the operator at a moment they control: either migrate the checkpoint to a supported method, or acknowledge the risk explicitly with --allow-deprecated-quantization. It costs one broken launch instead of an unexplained quality regression discovered weeks later.
  • How is --kv-cache-dtype different from --quantization?
    They quantize different things and buy different resources. --quantization applies to model weights and shrinks the fixed memory cost plus, on the right kernels, the compute cost. --kv-cache-dtype applies to the per-token cache and accepts auto, fp8, fp8_e4m3 and fp8_e5m2; halving cache bytes roughly doubles how much context and concurrency the same block pool holds. They are independent and are commonly used together.
  • After forcing a different quantization method, what do you check before shipping?
    Task-level quality, not just that the server starts. Run your own eval set through both configurations and compare on the metric your product cares about, plus a check on refusal and formatting behaviour, which quantization disturbs more visibly than aggregate scores. Also compare throughput and TTFT, since a method change is a kernel change and can move performance in either direction.

saying these in an interview costs you the question

  • Passes --quantization w8a8 as if it were a value
  • Thinks the flag converts a checkpoint's format
  • Always sets it explicitly instead of letting vLLM infer
  • Assumes deprecated methods fall back with a warning
  • Conflates weight quantization with KV-cache dtype

context