Before switching a production endpoint to a quantized model, what do you validate?
answer
- compare against the full-precision baseline
- perplexity flat, capabilities broken
- greedy decoding isolates the weights
- measure throughput at production concurrency
- canary with the old replica warm
basics
~20 sValidate both halves of the trade. On quality, compare against the full-precision baseline on real task evals and the failure modes quantization hits hardest. On serving, measure TTFT, inter-token latency and throughput at production concurrency — then canary traffic with the old replica warm for rollback.
solid answer
~50 sTreat it as a model rollout, not a config change. First fix the comparison: same engine, same server settings, same prompts, greedy decoding, so differences come from the weights and not from sampling. Then evaluate on your own task suite rather than a generic perplexity number — quantization damage shows up first in the brittle behaviours: structured output and tool-call formatting, long-context recall, non-English text, and multi-step arithmetic or code. Second, measure the serving side at your real concurrency: a checkpoint that halves VRAM but costs 30% of your tokens/s may be a net loss once you price it per million tokens. Third, roll out behind a canary — a small traffic share, side-by-side dashboards for TTFT, inter-token latency, error rate and any quality signal you can compute online — and keep the full-precision replica able to take traffic back. Record the exact checkpoint revision you validated.
go deeper
Know that a quantized model must be tested against the original before it serves users, and that the test should use your own real prompts rather than only a public benchmark score.
Explain why perplexity is a weak gate and name the fragile surfaces — structured output, long context, non-English, math — plus the need to pin sampling so the weights are the only variable.
Show you validate both halves: quality delta against the baseline and serving metrics at production concurrency, converted into cost per token at the SLO, then rolled out behind a canary with a warm rollback.
Own the policy: what quality delta is acceptable, who signs it off, how quantized artifacts and engine versions are pinned, and what the standing GPU cost of a warm rollback path is worth.
## Why this needs a rollout, not a flag flip Quantization changes the model's numerics. Everything downstream that was tuned against the old numerics — prompts, few-shot examples, JSON schemas, regexes over the output, thresholds on a classifier's phrasing — was validated against weights you are about to replace. The change is invisible in the deployment manifest and very visible in the product, which is the worst combination for an incident. ## Step 1: make the comparison honest Before any numbers mean anything, control the variables: - **Same engine and version** on both sides. Comparing a quantized model on a new engine build against a bf16 model on an old one confounds two changes. - **Same server settings**: context length, batching limits, and especially anything that changes numerics. - **Greedy decoding** (temperature 0, no top-p truncation) for the quality run, so output differences come from the weights rather than from the sampler. Note that even greedy output is not bit-identical run to run on a batching server, because kernel reduction order varies with batch composition — so compare distributions of scores, not string equality. - **The same prompt set**, ideally sampled from real production traffic rather than a public benchmark your model may have memorised. ## Step 2: evaluate what quantization actually breaks Aggregate perplexity is a poor gate. It averages over an enormous number of easy tokens and can move by a fraction of a percent while a capability you depend on falls over. Target the fragile surfaces: - **Structured output and tool calls.** Schema adherence and argument formatting degrade before prose does. If your product depends on parseable JSON, measure parse rate and schema-validity rate explicitly. - **Long context.** Retrieval-style recall from deep in a long prompt is disproportionately sensitive. Evaluate at the context lengths you actually serve, not at 2K. - **Languages other than English**, and any domain vocabulary your calibration data did not contain. - **Multi-step reasoning, math and code**, where a single wrong token cascades. - **Refusal and safety behaviour**, which is a post-training artifact and can shift. Score against the full-precision model as the reference, and decide the acceptable delta *before* you look at the results. ## Step 3: prove the serving win is real The entire reason for the change is cost or capacity, so measure it under production conditions: - **Throughput at your real concurrency**, not batch size 1. A weight-only 4-bit checkpoint can look brilliant single-stream and lose to bf16 at batch 128. - **TTFT and inter-token latency** distributions, not means. Quantization changes where the bottleneck sits, and a scheme that helps decode can leave prefill unchanged. - **Which kernel the engine chose.** A fallback to a generic kernel silently erases most of the expected gain; the startup logs report the resolved method. - **Realised VRAM headroom**, and what you did with it — usually more KV cache, which means higher achievable concurrency, which is where the throughput win actually comes from. Then convert to the number the business cares about: cost per million tokens at the latency SLO. A configuration that is 15% cheaper and 25% slower may fail the SLO and be worth nothing. ## Step 4: canary, watch, and keep the exit open Route a small slice of live traffic to the quantized replica while the full-precision replica keeps serving the rest. Watch, side by side: error and timeout rate, TTFT/ITL percentiles, output length distribution (a suddenly shorter or longer mean is a strong early signal something shifted), tool-call parse failures, and whatever human or model-graded quality signal you can compute online. Give it enough traffic and enough time to cover the long tail of prompt shapes. Rollback must be cheap. Keeping the old replica warm costs GPU-hours, but weight-loading cold starts are minutes long, so a scale-from-zero rollback is not a rollback during an incident. ## Step 5: record what you validated Pin the exact checkpoint revision, the quantization scheme and its parameters, the engine version and the eval results together. Quantized artifacts get re-published; "we tested the 4-bit one" is not a reproducible statement six months later, and an engine upgrade can change the kernel path underneath an artifact you never touched. ## The short answer Quality against the full-precision baseline on task evals that stress structured output and long context; serving metrics at production concurrency converted to cost per token at the SLO; a canary with a warm rollback path; and a recorded, pinned artifact.
- Why can't you just check that the quantized model produces identical output to bf16?Because it never will. Quantization changes numerics by design, and even two bf16 runs on a batching server can diverge because kernel reduction order depends on batch composition. The right test is statistical: score both models on a task suite and compare aggregate quality, with a delta threshold agreed in advance. Exact-match comparison produces false alarms and tells you nothing about whether the product still works.
- What single online signal is the best early warning that a quantized rollout has gone wrong?A shift in the output-length distribution, paired with downstream parse or validation failure rate. Degraded models tend to ramble, truncate, or drift out of the expected format long before an offline eval refresh would catch it, and both signals are cheap to compute on live traffic without any grader. Error rate and timeout rate cover the serving side; the length and parse signals cover quality.
- How does calibration data choice show up in this validation?As a domain-shaped hole. A checkpoint calibrated on generic English web text can look fine on generic evals and lose meaningfully on your domain — legal, medical, code, or a non-English language. That is why the eval set should be sampled from real production prompts: it is the only set that surfaces a mismatch between the calibration distribution and your traffic.
saying these in an interview costs you the question
- Gating a rollout on perplexity alone
- Benchmarking latency only at batch size one
- Comparing quantized and baseline with sampling enabled
- Assuming smaller weights automatically means cheaper serving
- Scaling the full-precision replica to zero before the canary finishes