Why can two 4-bit LLM builds differ sharply in quality at the same bit width?
answer
- the number of bits is not the spec
- how many weights share one scale
- trained at low precision, or converted after
- weight-only is not the same as everything
- ask what else got quantized
basics
~20 sBit width is a container, not a quality level. How finely the format shares scaling factors, whether the checkpoint was trained at low precision or converted afterwards, and what else got quantized alongside the weights all matter more than the number four.
solid answer
~50 sThree things separate two 4-bit builds. First, **scale granularity**: formats differ in how many weights share one scaling factor, and finer blocks tolerate outliers better — NVFP4's 16-element blocks give it an accuracy edge over MXFP4's 32-element blocks at the same nominal width. Second, **how the checkpoint reached 4 bits**: a vendor quantization-aware-trained checkpoint, where the model learned under low precision, generally beats a post-hoc conversion of the same weights, and published QAT releases such as Gemma's int4 checkpoints demonstrated exactly that gap. Third, **scope**: weight-only 4-bit is a much gentler change than also quantizing activations or the KV cache, so two builds labelled 4-bit may not be quantizing the same things. As of mid-2026 4-bit is a shipping default rather than a compromise — GPT-OSS shipped in MXFP4, and NVIDIA reports pretraining in NVFP4 with no measurable loss — but the label alone still tells you almost nothing.
go deeper
Know that the bit width alone does not tell you how good a quantized model is, and that two 4-bit releases of the same model can behave quite differently. Ask which format and where the checkpoint came from.
Explain the axes: how finely scales are shared, whether the checkpoint was trained quantization-aware or converted afterwards, and whether activations and the KV cache were quantized too. Be able to say why finer blocks tolerate outliers better.
Show that you check scope and provenance before comparing numbers, prefer a vendor quantization-aware checkpoint when one exists at your width, and validate on your own task rather than trusting a published comparison.
Own the format standardization decision across a fleet — which formats your hardware executes natively, whether to depend on vendor low-precision releases or maintain your own conversion pipeline, and what that commits you to as hardware turns over.
## The label is a container, not a guarantee "4-bit" states how many bits encode each weight value. It says nothing about how the values were chosen, how they are scaled, what else in the pipeline runs at low precision, or whether the model ever saw low precision during training. Two builds carrying the same label routinely land several points apart on the same task suite. Treating the width as the quality specification is the most common mistake in this area. ## Axis one: scale granularity Every low-precision format pairs a small integer or float code with a shared scaling factor. The question is how many weights share one factor. A scale shared across a whole tensor must accommodate the largest value anywhere in it; a scale shared across a small block only has to accommodate that block's range. Finer sharing means outliers damage fewer neighbours. That difference is the main reason NVFP4, which scales in 16-element blocks, tends to preserve more accuracy than MXFP4, which scales in 32-element blocks, at identical nominal width. Open-weight quantization tooling exposes the same axis under different names — per-tensor versus per-channel versus per-group scales — with the same directional effect: finer is more accurate and slightly larger, because the scales themselves take space. The internals of these formats are a topic of their own; what matters for the quality question is that granularity, not width, is doing much of the work. ## Axis two: how the checkpoint reached 4 bits There are three routes, in increasing order of typical quality: 1. **Post-training conversion.** Take released high-precision weights and map them into the low-precision format, possibly using a small calibration corpus to fit the scales. Cheapest and most common. 2. **Quantization-aware training or fine-tuning.** The model is trained or fine-tuned with the low-precision behaviour simulated in the loop, so it learns weights that survive the mapping. Vendors increasingly publish these directly — Google's published int4 QAT Gemma checkpoints are the canonical public example, and they beat post-hoc conversions of the same models at equal width. 3. **Native low-precision training.** The pretraining itself runs in the low-precision format. NVIDIA reports training in NVFP4 with no measurable quality loss versus higher-precision baselines, which is what turned 4-bit from a serving compromise into a shipping format. If a vendor QAT checkpoint exists at the width you want, it is almost always the better starting point than converting the high-precision release yourself. ## Axis three: what else is quantized "4-bit model" in casual usage often means weight-only 4-bit with activations still in 16-bit and the KV cache untouched. A build that also quantizes activations, or stores the KV cache at low precision, is a materially different system with a different quality profile — activation quantization exposes the outlier problem far more sharply than weight quantization does. Low-precision KV cache is now common because it halves that footprint and lets you hold far more concurrent context, but it is a separate quality decision from the weights. When comparing two builds, establish scope before comparing numbers. ## Axis four: hardware-native execution A format that the hardware executes natively behaves differently in practice from one emulated by dequantizing back to a wider type before each matmul. Blackwell-class hardware runs MXFP4 and NVFP4 natively; elsewhere a 4-bit checkpoint may be stored small and computed wide. That does not change the stored-quality story much, but it changes the throughput half of the trade entirely, and it changes which formats are worth choosing on a given fleet. ## The mid-2026 picture Four-bit is a mainstream production format, not an emergency measure. MXFP4 is an open standard with frontier models shipped in it; NVFP4 pairs finer blocks with hardware support and is used for training as well as serving; GPTQ, AWQ, SmoothQuant and GGUF k-quants remain the workhorses of the open-weight ecosystem. The live interview question is no longer "is 4-bit acceptable" but "which 4-bit, from which route, quantizing what". ## What to do with all this Never accept a width as a specification. Ask which format and granularity, whether the checkpoint is QAT or converted, whether activations and KV cache are included, and whether the target hardware runs it natively. Then measure on your own task — because the interaction between all four axes and your particular workload is not predictable from the labels.
- Why does a quantization-aware-trained checkpoint beat a post-hoc conversion at the same width?Because the model gets to adapt. In quantization-aware training the low-precision mapping is simulated during the forward pass, so gradients push the weights toward values that survive rounding and toward representations that do not depend on precision the format cannot carry. Post-hoc conversion has to take whatever the high-precision training produced and approximate it, with only a calibration corpus to fit the scales.
- If finer scale blocks are more accurate, why not make every block tiny?Because the scales themselves cost memory and compute. Shrink the block far enough and the metadata approaches the size of the data it describes, erasing the compression, and per-block work starts to dominate the kernel. The chosen block sizes — 32 elements in MXFP4, 16 in NVFP4 — are settlements between accuracy, storage overhead and what hardware can execute efficiently.
- Two teams report different quality for 'the same 4-bit model'. Where would you look first?Scope and route. Confirm whether both quantized weights only or also activations and the KV cache, and whether one started from a vendor quantization-aware checkpoint while the other converted the high-precision release. Then confirm both evaluated the same task with the same decoding settings, since sampling configuration alone can account for a visible gap.
saying these in an interview costs you the question
- Four bits is four bits, so quality is fixed
- Quantization always means weight-only quantization
- Post-hoc conversion matches quantization-aware training at equal width
- Block size is an implementation detail with no quality effect
- Any 4-bit format runs natively on any modern GPU