Do you standardize one quantization scheme across a serving fleet, or choose per model?
answer
- hardware forces a per-generation default
- the cost is validation, not compute
- every scheme re-validated on engine upgrade
- deprecation now fails closed
- keep a full-precision reference replica
basics
~20 sStandardize per GPU generation, not across the whole fleet — hardware decides which schemes have fast kernels. Allow per-model exceptions only with evidence, because every model-and-scheme pair carries its own validation, artifact and re-validation-on-upgrade cost.
solid answer
~50 sOne fleet-wide scheme is usually impossible: a mixed estate of Ampere and Hopper cards has no scheme that is fast on both, so the realistic default is one per GPU generation — FP8 where FP8 tensor cores exist, 4-bit or INT8 elsewhere. Above that, the argument for standardizing is cost of ownership: every model-and-scheme combination needs its own accuracy evaluation, its own performance benchmark, its own artifact in the registry, and re-validation whenever the engine changes kernels or deprecates a method — vLLM 0.27 now refuses `fbgemm_fp8` and `fp_quant` outright unless you pass `--allow-deprecated-quantization`. Compiled-engine stacks tighten this further, since a TensorRT-LLM artifact is pinned to a GPU architecture and precision and must be rebuilt to move. So: a documented default per SKU class, a written exception process requiring measured evidence, a bf16 reference replica retained for comparison, and quantization scheme treated as a versioned property of the deployment rather than a per-team preference.
go deeper
Know that different GPUs support different quantization schemes, so a single company-wide choice is rarely possible, and that each choice has to be tested before it serves traffic.
Explain the per-generation default — FP8 where the hardware has FP8 tensor cores, 4-bit or INT8 elsewhere — and why the checkpoint layout has to match a fast kernel.
Argue the maintenance side concretely: validation per model-and-scheme pair, re-validation on engine upgrade, and a deprecation that now fails closed rather than warning.
Own the policy itself — documented defaults per SKU class, an evidence-gated exception process, pinned artifacts and engine versions, a retained full-precision reference, and a compatibility gate on engine upgrades.
## Why the naive answer fails in both directions "Standardize on one scheme" fails because the fleet is heterogeneous. FP8 arithmetic exists only on Ada/Hopper-generation cards and newer; mandating FP8 fleet-wide means every Ampere node runs a dequantizing fallback that saves memory and buys no speed. "Choose per model" fails because the cost is not the choosing — it is everything that follows each choice. The workable position is a **default per hardware class, with an evidence-gated exception process**. ## The real cost of a long tail of schemes For each (model, scheme, GPU generation) triple you are on the hook for: - **An accuracy validation run** against the full-precision baseline, on task evals that stress structured output and long context. This is human time, not just compute. - **A performance benchmark** at production concurrency, because the same scheme behaves differently per model shape and per card. - **An artifact** in the model registry, with its own storage, provenance and revision pin. - **Re-validation on engine upgrade**, because the engine may route the same checkpoint to a different kernel, or stop accepting the method entirely. Ten models times three schemes is thirty of those, and the marginal one always looks cheap. Standardization is a bet that the aggregate maintenance bill exceeds the per-model tuning gains — which it usually does, except for the two or three highest-volume models where a few percent of GPU spend is real money. ## Deprecation is an operational, not academic, risk Quantization methods churn faster than almost anything else in a serving stack, because they are close to fast-moving kernels. The concrete current example: vLLM 0.27 raises on `fbgemm_fp8` and `fp_quant` unless the operator explicitly passes `--allow-deprecated-quantization`. Failing closed is the right design — silently substituting a different scheme would change model numerics without notice — but it means an engine bump can crash-loop a deployment that nobody touched. The fewer distinct schemes in production, the smaller that blast radius, and the shorter the release-note review before an upgrade. ## Compiled engines raise the stakes vLLM and TGI resolve kernels at load time, so a checkpoint can move between SKUs and simply take a different path — degraded, perhaps, but running. TensorRT-LLM compiles ahead of time: `trtllm-build` produces an engine pinned to a GPU architecture and precision, served through Triton's TensorRT-LLM backend. Moving that workload to a different card generation is a rebuild in a build pipeline, not a scheduling decision. If your fleet uses compiled engines, quantization scheme becomes a build-matrix dimension, and the argument for keeping that matrix small gets much stronger. ## A defensible policy 1. **Default per SKU class.** FP8 on FP8-capable cards; 4-bit AWQ/GPTQ in a fast-kernel-compatible layout, or INT8 W8A8, on Ampere. Write it down, including the acceptance thresholds. 2. **Keep a full-precision reference.** At least one bf16 replica or an on-demand path for the highest-value models, so quality regressions can always be A/B'd against ground truth and a rollback exists that is not itself quantized. 3. **Exception process.** A team may deviate with measured evidence: quality delta against the baseline and throughput at production concurrency, both on the target SKU. "It felt faster" is not evidence. 4. **Pin and version.** Checkpoint revision, scheme, engine version and eval results recorded together. Quantized artifacts get republished upstream; an unpinned dependency is an unreviewed model change. 5. **Gate engine upgrades** on a scheme-compatibility check across the deployed set — a cheap script, and the difference between a routine bump and a fleet incident. ## Where per-model choice genuinely wins Do not over-rotate to uniformity. Real reasons to deviate: a model whose only credible open checkpoint exists in one scheme; an MoE model where expert-specific methods matter; a very high-volume endpoint where a measured few percent of throughput pays for the extra validation many times over; or a model that demonstrably degrades at the fleet default and passes at a wider bit width. The policy exists to make those decisions explicit and evidenced, not to forbid them. ## The framing to give an interviewer Say that hardware sets the floor (per-generation defaults are forced), that the marginal cost of a scheme is validation and re-validation rather than compute, that deprecation now fails closed and therefore couples quantization choice to engine upgrades, and that you would run a documented default plus an evidence-gated exception path with a full-precision reference retained.
- What makes engine upgrades riskier when many quantization schemes are in production?Two mechanisms. A method can be removed or gated, as vLLM 0.27 does by raising on fbgemm_fp8 and fp_quant unless --allow-deprecated-quantization is passed, which turns an untouched deployment into a startup failure. Less visibly, the engine can route the same checkpoint to a different kernel, changing throughput or numerics without any error. Each distinct scheme in production is one more surface to check before a bump.
- Where is a per-model exception clearly worth the maintenance cost?On the highest-volume endpoint, where a measured few percent of throughput is a large absolute GPU bill; on a model whose only credible open checkpoint exists in one scheme; and on any model that fails your quality bar at the fleet default and passes at a wider bit width. In each case the exception is justified by a number, not a preference, and it inherits the same pinning and re-validation obligations.
- Why keep a full-precision replica once quantized serving is proven?Because it is the ground truth for every later question. Regression investigations need a non-quantized comparison, new evals need a reference to score against, and an incident rollback must land on something you trust. It costs GPU-hours, but weight-load cold starts run into minutes, so a scaled-to-zero reference is not available during the incident where you need it.
saying these in an interview costs you the question
- Mandating one scheme across mixed GPU generations
- Treating a new scheme's cost as just the quantization run
- Letting each team pick a scheme without evidence
- Assuming compiled engines move between GPU models freely
- Deleting the full-precision baseline once quantized serving ships