skip to content

In TensorRT-LLM, how do FP8 or INT4-AWQ get baked into an engine?

level: seniorimportance: should knowfreq 38%

answer

  1. precision is chosen before the build
  2. a calibration pass writes the scales
  3. format lives in the checkpoint config
  4. no runtime switch, only a rebuild
  5. FP8 needs Ada/Hopper-class tensor cores

basics

~20 s

A calibration pass runs before the build: the quantization script reads the model with a --qformat such as fp8 or int4_awq, computes scales over a calibration set, and writes a quantized checkpoint. trtllm-build then compiles kernels for that format, so precision is frozen in the engine, not switchable at serve time.

solid answer

~40 s

Precision in TensorRT-LLM is a build-time property. Instead of a plain checkpoint conversion you run the quantization example script, passing `--qformat` (values include `fp8`, `int4_awq`, `int8_sq`, `w4a8_awq`, `full_prec`) and a `--calib_size`; it runs the model over calibration data, derives the scaling factors that format needs, and writes a quantized TensorRT-LLM checkpoint. `--kv_cache_dtype fp8` quantizes the KV cache as a separate decision from the weights. `trtllm-build` then compiles the kernels matching that scheme. The consequence people miss: there is no runtime flag to change precision. A BF16 engine and an FP8 engine are two artifacts, each needing its own build, its own storage and its own accuracy validation. Hardware constrains the choice too — FP8 tensor cores exist on Ada- and Hopper-class GPUs and newer, while INT4-AWQ weight-only quantization also runs on Ampere.

code

bash · 14 lines
bash
# calibrate + quantize -> quantized TensorRT-LLM checkpoint
python quantize.py \
  --model_dir ./Meta-Llama-3.1-8B-Instruct \
  --dtype bfloat16 \
  --qformat fp8 \
  --kv_cache_dtype fp8 \
  --calib_size 512 \
  --output_dir ./ckpt-llama3-8b-fp8

# build kernels for that scheme
trtllm-build \
  --checkpoint_dir ./ckpt-llama3-8b-fp8 \
  --output_dir ./engines/llama3-8b-fp8 \
  --gemm_plugin auto

go deeper

for a junior

Know that a quantized TensorRT-LLM deployment starts from a quantization step before the build, and that the resulting engine is a different file from the full-precision one.

for a middle

Explain the pipeline — calibration pass with a chosen qformat, quantized checkpoint, then a build that compiles matching kernels — and say why no runtime flag can change it afterwards.

for a senior

Show the production discipline: pick the format your GPUs have fast kernels for, validate the built engine against the baseline on task metrics before it ships, and keep the previous artifact as a rollback.

for a principal

Own the cost of the matrix. Decide how many precisions the fleet standardizes on, given that each one multiplies builds, storage and validation runs, and whether uniform hardware is cheaper than supporting several schemes.

## Where precision is decided In a Python-level engine you often point a server at a prequantized checkpoint and it picks kernels at load. In TensorRT-LLM's ahead-of-time flow the decision happens two steps earlier, and it happens twice: once when the checkpoint is produced with quantization scales, and once when the builder compiles kernels for that scheme. By the time the engine exists, precision is a property of the binary. ## The calibration pass The quantization example script takes the Hugging Face model directory and a format: - `--qformat fp8` — 8-bit floating point for weights and activations, using per-tensor scales. - `--qformat int4_awq` — 4-bit weight-only, activation-aware, using per-group scales. - `--qformat int8_sq` — 8-bit with smoothing applied to make activations easier to quantize. - `--qformat w4a8_awq` — 4-bit weights with 8-bit activations. - `--qformat full_prec` — no weight quantization, used when you only want, say, an FP8 KV cache. `--calib_size` sets how many samples the pass runs. The output is a checkpoint in the normal TensorRT-LLM layout, with the quantization scheme recorded in its config and scaling factors stored beside the weights. `--kv_cache_dtype fp8` is orthogonal: it changes how cached keys and values are stored, which buys sequence capacity rather than GEMM speed, and it can be combined with an otherwise full-precision model. ## What the builder does with it `trtllm-build` reads the scheme from the checkpoint config and selects the kernel implementations that consume those scales. This is why there is no `--quantization` flag on the build command doing the work: the checkpoint already carries the answer. It is also why mixing steps across library versions is risky — the scale layout and the kernels that read it evolve together. ## Hardware pins the menu FP8 tensor cores appear on Ada- and Hopper-class GPUs and their successors; on older cards an FP8 engine simply is not the right build target. INT4 weight-only schemes such as AWQ run on Ampere as well, because the weights are dequantized into a supported compute type inside the kernel. This is why the same organisation can end up with an INT4-AWQ engine for its A100 pool and an FP8 engine for its H100 pool — the same model, two artifacts, because the hardware differs. The two formats also behave differently under load. Weight-only 4-bit shrinks the bytes moved per token, which is exactly what a memory-bandwidth-bound decode step wants at small batch. As batch grows and the work becomes compute-bound, the dequantize step stops paying for itself, while FP8 — which runs the math in 8 bits rather than only storing weights that way — keeps its advantage. Choosing between them is a question about your batch regime, not a question about which number is smaller. ## The validation obligation Because the artifact is compiled, "we quantized it" is not a claim you can check by reading a config in production — you check it by evaluating the engine you are about to ship. A workable gate: run your task evals against the quantized engine and against the BF16 baseline on the same prompts, compare on the metrics the product actually cares about, and eyeball a set of long-context and structured-output prompts, which is where quantization damage tends to surface first. Store the result alongside the artifact. When someone asks in three months whether the FP8 engine regressed, the answer should be a recorded number, not a memory. ## Operational consequences - **The artifact matrix multiplies.** Precision joins model, parallel degree, GPU SKU and library version as an artifact key. Two precisions across two SKUs is four builds to produce and validate for one model. - **You cannot A/B by flag.** Comparing FP8 against BF16 in production means running two deployments and splitting traffic, not toggling a setting. - **Rollback is a redeploy.** If the quantized engine regresses, you fall back to a previously built artifact — which only exists if you kept it. - **KV-cache dtype is a separate lever.** If the problem is that you cannot fit enough concurrent sequences, an FP8 KV cache may buy more than shrinking the weights, and it can be adopted independently. ## The interview shape The reflex answer to a cost question is "quantize it". What distinguishes a senior answer here is knowing that in a compiled engine the decision is upstream and irreversible without a rebuild, that the hardware constrains which formats are even on the menu, and that shipping it responsibly means a recorded accuracy comparison against the baseline you replaced.

  • You want to compare FP8 against BF16 on live traffic. How do you set that up?
    Two deployments. Each precision is a separate compiled engine, so there is no flag to flip — you build both artifacts, run them behind a splitter, and compare on product metrics plus latency and cost per token. Keep the BF16 artifact around afterwards; it is your rollback path if the quantized engine regresses later.
  • When is quantizing only the KV cache the better move?
    When the constraint is concurrency rather than compute. An FP8 KV cache halves the bytes per cached token against 16-bit, so more sequences or longer contexts fit in the same VRAM, and it can be applied with an otherwise full-precision model. If instead your decode step is bandwidth-bound on weights at small batch, weight quantization is the lever that helps.
  • Why can a 4-bit weight-only engine lose its advantage at large batch?
    Weight-only quantization shrinks the weight bytes read per forward pass, which dominates at small batch where decode is memory-bandwidth-bound. As batch grows, the same weights are reused across many sequences and the step becomes compute-bound, so the dequantization work in the kernel starts to cost more than the bandwidth it saves.

saying these in an interview costs you the question

  • Thinks precision can be switched with a serving flag
  • Assumes FP8 engines run on any NVIDIA GPU
  • Ships a quantized engine with no accuracy comparison
  • Confuses logits_datatype with compute precision
  • Treats KV-cache dtype and weight format as one decision

context