After quantizing Llama to 4-bit, how do you verify the accuracy loss is acceptable?
answer
- Compare against an unquantized control
- Perplexity rejects, it does not accept
- Score your own task, not benchmarks
- Structured output degrades first
- Small models lose more than large ones
basics
~20 sCompare the quantized build against the same unquantized weights on your own task, not just on perplexity. Perplexity averages over tokens and hides narrow failures, so add task evals for structured output, tool-call validity, code and long context, plus a side-by-side sample review.
solid answer
~50 sTreat it as a regression test against a reference. Keep the FP16 or Q8_0 build as the control, then measure three layers: 1. **Perplexity** on a held-out corpus is a cheap smoke test — a large delta means something is badly wrong. A small delta proves very little, because it is a token-averaged likelihood and your failures are usually concentrated. 2. **Task evals on your own data.** Score the things that break first under quantization: exact structured output, JSON that validates against your schema, tool-call argument correctness, code that compiles, multi-step arithmetic, and behaviour at the long end of your context distribution. Score with greedy decoding so sampling noise does not swamp the comparison. 3. **Human spot-checks** of paired outputs on real prompts, which catch tone and instruction-following drift that no metric flags. Run the ladder — Q8_0, 4-bit, 3-bit — and pick the lowest level whose deltas you can defend. Also expect smaller models to lose more than large ones at the same bit width.
go deeper
Know that you must compare the quantized model against the unquantized one on the same prompts, and that lower bit widths trade accuracy for memory rather than being free.
Explain why perplexity is only a smoke test and name what degrades first — exact formatting, schema-valid JSON, tool arguments, multi-step arithmetic, long-context recall.
Demonstrate a real harness: a fixed control build, deterministic scoring on production-shaped data, greedy decoding, enough examples for the delta to be meaningful, and paired human review for what metrics miss.
Own the acceptance policy — who signs off on a quantized build, what regression budget is acceptable against the memory saved, and how the harness is re-run automatically whenever weights or the quantizer change.
## Why this needs a process Quantization does not fail loudly. A 4-bit Llama still speaks fluently, still answers general questions, and still scores respectably on public benchmarks. What it loses is *precision*: exact formatting, reliable schema adherence, correct arithmetic chains, tool arguments that match the declared JSON Schema, and stability at the far end of the context window. Those are exactly the behaviours production pipelines depend on and public benchmarks under-weight. So the verification job is to build a measurement that looks where the damage actually lands. ## Establish a control You cannot say "the 4-bit build is fine" in isolation; you can only say it is or is not different from a reference. Use the FP16/BF16 weights as the control where you can afford to run them, or Q8_0 / int8 as a practical proxy — 8-bit is close enough to lossless that a 4-bit-versus-8-bit comparison is honest. Fix everything else: same prompts, same chat template, same seeds, same decoding parameters, same stop conditions. Any of those drifting between runs will produce differences you will wrongly attribute to quantization. ## Layer 1 — perplexity as a smoke test Perplexity on a held-out corpus (llama.cpp ships a perplexity tool for exactly this) is cheap and catches gross breakage: a build with a mangled tensor mix or a bad calibration run shows a large delta immediately. Its limits matter just as much: - It averages over every token, so a model that is right 99 % of the time and catastrophically wrong on closing braces looks fine. - It is measured on generic text, which may look nothing like your traffic. - Absolute values are not comparable across corpora or tokenizers, only deltas on an identical setup are. So use it to *reject* builds, never to *accept* them. ## Layer 2 — task evals on your own data This is the load-bearing layer. Build a few hundred graded examples from real traffic, and score deterministically wherever possible: - **Schema validity.** Does the structured output parse and validate? Percentage-valid is a brutally clear metric and is one of the first things to sag at low bit widths. - **Tool-call correctness.** Right tool chosen, arguments well-formed and semantically right. - **Executable checks.** For code, does it compile and pass tests? For extraction, does the field match ground truth exactly? - **Multi-step reasoning.** Chained arithmetic or multi-hop questions, where small per-step degradation compounds. - **Long context.** Sample at the long end of your real length distribution, not just short prompts. Retrieval-from-context accuracy is quantization-sensitive. - **Non-English and domain jargon**, if you serve them — a build calibrated on English web text can regress specifically there. Use greedy decoding (temperature 0) for the comparison so you are measuring the weights, not the sampler. Then, separately, sanity-check with your production sampling settings, because greedy hides some instabilities. Be realistic about statistics: a two-point difference on 100 examples is noise. Size the eval set so the difference you care about is detectable, and report the delta with an interval rather than a bare number. ## Layer 3 — paired human review Run the same fifty to a hundred real prompts through both builds and read the pairs side by side, blind if you can. Metrics miss register shifts, verbosity creep, instruction-following slippage and repetition loops. This is also the step that catches the pathological failure mode of aggressive quantization: degenerate repetition at long generations. ## What the ladder usually looks like Run 8-bit, 4-bit and 3-bit through the same harness. Typically 8-bit is indistinguishable from the reference, 4-bit shows small but measurable regressions concentrated in the precise tasks, and 3-bit and below fall off sharply — with the fall being far steeper for small models. A 70B at 4-bit is a mild compromise; a 1B–3B model at 3-bit often is not usable for structured work at all. Model size buys redundancy, and redundancy is what quantization spends. ## Deciding and keeping it honest Pick the lowest bit width whose regressions you can state and accept, with the memory saving written next to it — the decision is a tradeoff, so present it as one. Then keep the harness: requantize when weights are updated and re-run it, because a new build of the "same" quant level is a new artifact, and if the calibration corpus or the quantizer version changed, so did the model. Finally, watch for confounds that masquerade as quantization damage: a wrong or subtly different chat template, a tokenizer mismatch, changed stop tokens, or a different sampling default in the serving stack. Confirm the control and the candidate share all of those before blaming the bits.
- Why is greedy decoding the right setting for this comparison?Because sampling noise is often larger than the quantization effect you are trying to measure. At temperature 0 the same prompt gives the same continuation, so any difference between control and candidate comes from the weights. Once you have the greedy delta, re-run a smaller pass with production sampling settings — some instabilities, notably repetition loops at long generations, only appear under sampling.
- How would you catch a regression that only appears at long context lengths?Sample your eval prompts from the long tail of your real length distribution rather than a uniform short set, and include retrieval-style items where a fact is planted early and must be used late. Quantization error compounds with sequence length and interacts with position handling, so a build that is clean at 2k tokens can degrade noticeably at 64k. If you never test there, you ship the regression.
- A quantized build scores fine on public benchmarks but users report worse results. Where do you look first?At the confounds before the bits: chat template, tokenizer version, stop tokens, and the serving stack's default sampling parameters — a template mismatch produces exactly this symptom. If those match, build an eval from the failing traffic itself, since public benchmarks under-weight structured output, tool calls, long context and non-English text, which are where 4-bit quantization actually costs you.
saying these in an interview costs you the question
- Accepting a build on perplexity delta alone
- Comparing against a published benchmark instead of your own control
- Testing only short prompts and general chat
- Sampling at high temperature while comparing builds
- Assuming a small model tolerates 3-bit like a 70B does