Which quantization scheme suits an A10G versus an H100 inference server?
answer
- kernels decide, not bit width
- FP8 needs Ada or Hopper
- INT8 portable back to Turing
- loading is not the same as accelerating
- 4-bit buys capacity, 8-bit buys math
basics
~20 sHardware decides. FP8 tensor cores exist only on Ada and Hopper generations and newer, so an H100 should serve FP8 weights and activations. An Ampere card like the A10G or A100 has no FP8 math and should serve 4-bit GPTQ or AWQ through a Marlin-class kernel, or INT8.
solid answer
~50 sPick the scheme your GPU has *fast kernels* for, not the one with the smallest bit width. FP8 arithmetic requires tensor cores that only appeared with Ada (L4, L40S) and Hopper (H100, H200); on those cards an FP8 checkpoint narrows weights and activations and roughly doubles the math rate with a very small accuracy cost, which makes it the default for a throughput-oriented deployment. On Ampere (A100, A10G) there is no FP8 tensor core: an engine can still *load* an FP8 checkpoint — vLLM will dequantize it through a Marlin-style kernel — so you get the memory saving without the speedup. On that generation the useful options are 4-bit GPTQ/AWQ in a Marlin-compatible layout, which buys VRAM and low-batch latency, or an INT8 W8A8 checkpoint, since INT8 tensor cores have been available since Turing. Always confirm the engine reports the fast kernel path at startup rather than a generic fallback.
code
bash · 2 lines# Hopper/Ada: quantize a bf16 checkpoint to FP8 online, no calibration needed
vllm serve my-org/llama-3.1-8b-instruct --quantization fp8go deeper
Know that quantization support depends on the GPU, and that the newest low-precision formats need newer cards. Say you would check the engine's supported values and the card's generation.
Explain that FP8 arithmetic needs FP8 tensor cores while 4-bit weight-only runs anywhere with a mixed-precision kernel, and name what each buys — math rate versus memory.
Demonstrate that you verify the resolved kernel path in the logs and benchmark against bf16 on the same SKU, rather than trusting that a checkpoint loading means it is accelerated.
Own the fleet consequence: a mixed GPU estate means a per-generation quantization default and two artifacts per model, and that maintenance cost belongs in the capacity plan.
## The rule: kernels, not bit widths A quantization scheme is only as good as the kernel that executes it on your specific GPU. Two questions decide everything: 1. Does this GPU have **tensor cores for that numeric type**? If yes, the multiply itself runs faster. 2. If not, does the engine have a **mixed-precision kernel** that unpacks the stored weights and multiplies in 16-bit? If yes, you keep the memory saving and lose the arithmetic saving. Candidates who answer "4-bit is smaller so it's better" miss both questions. ## What each generation gives you - **Turing and later** have INT8 tensor cores. INT8 W8A8 is therefore the broadly portable 8-bit option. - **Ampere** (A100, A10G, A30) adds bf16 and is the first generation with good support for the mixed-precision 4-bit kernels — Marlin-class GEMMs target Ampere and newer. No FP8 math. - **Ada Lovelace** (L4, L40S) and **Hopper** (H100, H200) add FP8 tensor cores. This is where FP8 becomes a *performance* feature rather than only a memory one. - **Blackwell** adds still narrower formats (NVFP4-style 4-bit with block scales); vLLM exposes these through `modelopt`-family and `mxfp4`/`nvfp4` methods. Treat support as newer and thinner than FP8's. ## The trap: loading is not accelerating An FP8 checkpoint on an A100 does not fail. vLLM will serve it by dequantizing the FP8 weights through a Marlin path and doing 16-bit math. Memory drops roughly by half; tokens per second do not improve, and can regress at high batch for exactly the reason weight-only 4-bit does. Teams have shipped this configuration believing they got "H100-class FP8 performance on Ampere" because nothing errored. The tell is in the startup logs, which report the resolved quantization method and kernel. The same trap exists at 4-bit: if the checkpoint's bit width, group size or symmetry falls outside what the fast kernel supports, the engine falls back to a generic GPTQ or AWQ kernel that is much slower at serving batch sizes. Re-quantizing into a supported layout is usually worth more than any other tuning you will do that week. ## Choosing per SKU **H100 / L40S, throughput-oriented service.** FP8. Either serve a checkpoint already quantized to FP8, or let vLLM quantize a bf16 checkpoint online with `--quantization fp8` — online FP8 is dynamic and cheap, needs no calibration pass, and is usually within noise of bf16 on quality. You narrow weights *and* activations, so both memory and math improve. **A100 80G, large model that must fit.** 4-bit AWQ or GPTQ in a Marlin-compatible layout. The goal here is capacity: shrinking a 70B model's weights from ~140 GB to ~35 GB is what makes it fit on two cards with room left for KV cache. Accept that the arithmetic is not faster. **A10G 24G, small model, cost-sensitive.** 4-bit again, mostly to leave enough VRAM for a usable KV cache. An 8B model at bf16 already eats ~16 GB of a 24 GB card, leaving very little for concurrency; at 4-bit you have room to batch. **Mixed fleet.** You will end up with a per-generation default rather than one fleet-wide scheme, and that means maintaining two quantized artifacts per model. ## Engines spell this differently vLLM uses `--quantization`, whose accepted values include `awq`, `gptq`, `gptq_marlin`, `awq_marlin`, `compressed-tensors`, `fp8`, `bitsandbytes`, `mxfp4` and `torchao`. TGI uses `--quantize`, with values including `awq`, `gptq`, `bitsandbytes` and `fp8`. TensorRT-LLM does it at build time instead: the quantization is baked into a compiled engine by `trtllm-build`, which is why those artifacts are pinned to a GPU architecture and precision. ## How to answer well Name the hardware constraint first (FP8 needs Ada/Hopper-or-newer tensor cores; INT8 is portable back to Turing; 4-bit mixed-precision kernels want Ampere-or-newer). Then say what you are optimising for — capacity or arithmetic — because that picks between 4-bit and 8-bit. Then close with verification: check the engine's reported kernel path, and benchmark at production concurrency before believing any of it.
- How would you catch a deployment that loaded FP8 on Ampere and quietly lost the speedup?Two ways. Read the startup logs — the engine reports the resolved quantization method and kernel, and a dequantizing fallback path is visible there. Then benchmark tokens/s at your real concurrency against the bf16 baseline on the same card: if memory dropped but throughput did not move (or regressed), you are on the fallback path and should either move the workload to an FP8-capable SKU or switch scheme.
- Why does TensorRT-LLM force this decision earlier than vLLM or TGI?Because TensorRT-LLM compiles an engine ahead of time with `trtllm-build`, and that artifact is pinned to a GPU architecture and precision. vLLM and TGI resolve kernels at load time, so the same checkpoint can move between SKUs and simply pick a different path. With a compiled engine, changing GPU generation or precision means a rebuild, which turns quantization choice into a build-pipeline concern.
- Does online FP8 quantization need a calibration dataset?Not for the common dynamic case: vLLM's `--quantization fp8` computes scales on the fly from a bf16 checkpoint, so there is no calibration pass and no extra artifact to store. Static per-tensor or per-channel schemes produced offline can use calibration to fix scales ahead of time, which removes the runtime scaling work at the cost of an offline step and a checkpoint you must version.
saying these in an interview costs you the question
- Choosing the smallest bit width available regardless of GPU
- Assuming FP8 runs fast on Ampere because it loads
- Ignoring which kernel the engine actually selected
- Treating a compiled engine as portable across GPU models
- Believing INT8 and FP8 are interchangeable on any card