skip to content

Why do 4-bit serving recipes keep certain layers at higher precision?

level: seniorimportance: should knowfreq 42%

answer

  1. some layers cost far more than others
  2. a shared scale must cover the extremes
  3. first and last blocks matter most
  4. tiny tensors, outsized influence
  5. the MLP stacks carry the memory mass

basics

~20 s

Quantization error is not spread evenly. A few tensors carry extreme activation outliers or sit where error cannot be absorbed downstream, so squeezing them costs far more accuracy than the memory it saves. Recipes leave those in higher precision and push the bulk to 4-bit.

solid answer

~50 s

Sensitivity varies enormously across a transformer. Some projections see activations with a handful of channels whose magnitudes are orders larger than the rest; a shared scale must stretch to cover those outliers, which crushes every ordinary value into a few representable levels and destroys the signal. Position matters too — the first block operates on raw embeddings and the last block plus the output head shape logits directly, with no downstream layer left to absorb the error. Small decisive tensors like embeddings, the output head, normalization parameters and the routing weights in a mixture-of-experts model are usually left alone as well. The trade is cheap because those tensors are a tiny fraction of the parameters: the MLP stacks hold most of the mass, so quantizing them to 4-bit captures nearly all the memory win. You find the sensitive layers empirically — quantize one block at a time, measure the task delta, rank, and keep the worst offenders high.

go deeper

for a junior

Know that quantization is usually not applied uniformly, and that some parts of a model — the embedding and output layers especially — are commonly left at higher precision because they are small and important.

for a middle

Explain the mechanism: a shared scale must cover the largest magnitude in its group, so outlier channels crush everything else onto a few levels. Be able to say why early and late blocks are more error-sensitive than middle ones.

for a senior

Demonstrate the empirical loop — a per-layer sensitivity sweep with a task-level readout, a ranked exclusion list, and awareness that mixed layouts can cost throughput even while they buy accuracy.

for a principal

Own the direction of travel: whether to invest in hand-tuned mixed precision per checkpoint or standardize on a uniform hardware-native format with quantization-aware training, given your hardware fleet and how often you re-quantize.

## The premise: error is not uniform It is tempting to model quantization as a single knob applied to a homogeneous pile of weights. Measured layer by layer, a transformer behaves nothing like that. Quantize one block at a time and measure the resulting drop on a task metric and you get a spiky profile: most blocks cost almost nothing, and a handful cost more than all the others combined. Every practical low-precision recipe exploits that shape. ## What makes a tensor sensitive: outliers Quantization represents a group of values with a small set of levels plus a shared scaling factor. The scale must reach the largest magnitude in the group. In several places inside a transformer — most notoriously the projections that consume attention output and the inputs to certain feed-forward projections — a few activation channels carry magnitudes one or two orders larger than the rest. The scale stretches to cover them, and everything else in the group collapses onto a couple of representable levels. The result is not gentle rounding noise; it is the near-total loss of the small-magnitude signal that most of the computation actually depends on. This is why "outlier channels" is the vocabulary of the field. Whole families of algorithms exist to move that difficulty around — that machinery belongs to the algorithms discussion — but even with them, the outlier-heavy tensors remain the least tolerant of aggressive width reduction. ## Position in the stack Two positional effects show up repeatedly. **Early layers.** The first block or two operate on the raw embedded representation. Error injected there propagates through every subsequent layer, and the network has had no opportunity to build redundancy that could route around it. **Late layers and the head.** The final block and the output projection map hidden state to logits directly. There is no downstream computation left to absorb a perturbation — a small logit shift is immediately a changed token distribution. That is precisely the regime where decisive-token errors appear. ## Small but decisive tensors A separate category is tensors that are tiny in parameter count yet control routing of information: - **Embedding and output-head matrices** — large in raw element count for big vocabularies, but degrading them hits every token directly. - **Normalization parameters and biases** — negligible memory, and quantizing them is essentially free damage. - **Router or gate weights in a mixture-of-experts model** — with fine-grained expert layouts now standard in most frontier open models, a tiny router matrix decides which experts see each token. A rounding error there does not blur an output; it sends the token to a different expert entirely. These are routinely excluded from quantization. ## Why keeping them high is cheap The memory math is the reason mixed precision works at all. The bulk of parameters in a dense transformer sit in the feed-forward stacks, and in a mixture-of-experts model in the expert weights. Sensitive tensors — a few attention projections, the first and last blocks, routers and norms — are a small share of the total. Keeping them at 8 or 16 bits while pushing the mass to 4-bit gives up only a few percent of the memory saving and recovers most of the accuracy that uniform 4-bit gave away. A representative recipe: attention output projections plus the first and last transformer blocks stay high, the MLP stacks go to 4-bit. ## Finding them Do not guess from architecture diagrams alone; the profile is checkpoint-specific. A sensitivity sweep is the standard approach: hold everything else at high precision, quantize one block or one tensor type at a time, and measure either the end-task delta or the relative error of that layer's output against the high-precision reference. Rank the layers, keep the top offenders high, and stop where the accuracy curve flattens. Aggregate perplexity is a poor readout for this sweep; use a task metric with enough resolution to see the differences. ## The cost of mixed precision, and the 2026 direction Mixed precision is not free on the serving side. Heterogeneous layouts complicate kernels, prevent some fusion, and can leave throughput below what a uniform, hardware-native format achieves — a build that is more accurate per bit but slower per token. That is one reason the field has been moving toward uniform low-precision formats executed natively on hardware, combined with quantization-aware training to recover the accuracy that hand-tuned mixed precision used to buy. Hand-selected exclusion lists remain common in open-weight tooling and remain the right move when a uniform recipe measurably breaks your task.

  • How would you actually identify the sensitive layers in an unfamiliar checkpoint?
    A sweep. Hold the model at high precision, quantize one block or tensor type at a time, and measure either the end-task delta or the relative error of that layer's output against the reference. Rank the results and keep the worst offenders high, stopping where the curve flattens. Use a task metric rather than perplexity as the readout — the sweep needs resolution the aggregate does not provide.
  • Can a mixed-precision build ever be slower than a uniform 4-bit one?
    Yes, and it often is. Heterogeneous layouts break kernel fusion and force extra conversion steps, so the accuracy you buy per bit can cost you tokens per second. That is a real part of the trade: a uniform format executed natively on hardware may deliver more throughput at slightly worse quality, and the right answer depends on whether your bottleneck is memory or latency.
  • Why are mixture-of-experts routers usually excluded from quantization?
    Because a router is a tiny matrix with discrete consequences. Rounding error in a normal weight blurs an output slightly; rounding error in a gate flips which experts process the token, which changes the computation entirely rather than perturbing it. The parameter count is negligible, so keeping it in higher precision costs essentially no memory and removes a sharp failure mode.

saying these in an interview costs you the question

  • All layers degrade equally, so uniform width is fine
  • Sensitivity is about parameter count, not activation distribution
  • Keeping some layers high precision wastes most of the memory saving
  • Late layers matter least because they are near the output
  • Mixed precision is always faster than uniform low precision

context