skip to content

Why are activations harder to quantize than weights in a transformer?

level: middleimportance: should knowfreq 56%

answer

  1. one operand is static, one is not
  2. ranges must be estimated, not measured
  3. a few channels dominate every token
  4. the spiky axis is the summed axis
  5. difficulty can be moved into the weights

basics

~20 s

Weights are fixed after training, so their ranges can be measured once, offline. Activations change with every input and contain a few hidden channels whose magnitudes are far above the rest, so a single scale either clips those channels or wastes almost all the range on them.

solid answer

~50 s

Weight-only quantization (say 4-bit weights, 16-bit activations) is the easy half: the tensor is static, so you can inspect it, group it and pick scales once. Weight-and-activation quantization — for example 8-bit weights with 8-bit activations — is what actually lets the matrix multiply run on integer or FP8 hardware paths, because both operands have to be low precision for that to work. The difficulty is that activations are data-dependent and, in large transformers, structurally spiky: a small number of hidden channels carry persistently large values across essentially all tokens. A per-tensor scale sized for those outliers leaves the ordinary channels compressed into a handful of levels. Two mitigations dominate. Finer scale granularity — per-token, or per-channel where the layout allows — and **outlier migration**, the SmoothQuant idea of dividing the outlier channels by a factor and multiplying the matching weight rows by the same factor, which is mathematically identity and can be folded into the preceding layer offline.

code

python · 11 lines
python
import torch

x = torch.randn(8, 512)
x[:, 17] *= 60                      # one persistent outlier channel
w = torch.randn(512, 512) * 0.05

s = x.abs().amax(0).sqrt() / w.abs().amax(1).sqrt()
x_s, w_s = x / s, w * s[:, None]

print(x.abs().max().item(), x_s.abs().max().item())
print(torch.allclose(x @ w, x_s @ w_s, atol=1e-3))

go deeper

for a junior

Know the two shapes: weight-only quantization compresses parameters only, while quantizing activations as well is what lets the multiply itself run in low precision. Say plainly that activations depend on the input.

for a middle

Explain the mechanics — static calibrated ranges versus per-token dynamic scaling, and the persistent per-channel outliers that make a single scale useless. Be able to state the smoothing identity that moves outliers into the weights.

for a senior

Show you have debugged this: which layers saturate, how you detect clipping in production, why the outlier axis is the accumulation axis so per-channel activation scales do not factor out, and what you would try when W8A8 still regresses.

for a principal

Own the decision of whether activation quantization is worth pursuing at all for a given fleet — it buys a compute path that only some accelerator generations expose, and it adds calibration, monitoring and revalidation obligations that weight-only compression does not.

## Two different jobs "Quantizing a model" can mean two quite different things. **Weight-only** quantization compresses the parameters and leaves the arithmetic in floating point: the low-bit weights are dequantized on the way into the matrix multiply. **Weight-and-activation** quantization compresses both operands, so the multiply itself can execute on an integer or low-precision floating-point path. The first is primarily about the size of the thing you have to move; the second is about what instruction the accelerator gets to issue. That difference in purpose is why the two are attempted so differently, and why the second is much harder. ## Weights are a static, inspectable tensor After training, a weight matrix is a fixed array. You can compute its exact distribution, split it into groups of 64 or 128 along whichever axis suits the kernel, and choose a scale per group with full knowledge of what it must cover. If a group has an awkward spread you can shrink the group. Nothing at runtime can surprise you. Rounding error is then a one-time, offline optimization problem — which is exactly what algorithms in the weight-only family attack. ## Activations are a distribution, not a tensor An activation tensor exists only while a particular batch of tokens is flowing through. Its range depends on the input, and you are choosing scales for a distribution you can only sample. There are two strategies: - **Static scaling**: estimate ranges during calibration and freeze them. Cheapest at runtime, but any production input outside the calibrated range gets clipped. - **Dynamic scaling**: compute the scale on the fly, typically per token, from the values actually present. Robust to distribution shift, but it costs a reduction over the tensor before every quantized matmul and complicates kernel fusion. Even with dynamic scaling, the shape of the distribution is the real problem. ## The outlier structure Large transformers develop **persistent per-channel outliers**: specific hidden dimensions whose activation magnitudes sit far above the rest, consistently, across nearly all tokens and layers. They are not noise and you cannot clip them away — they carry information the model depends on, and truncating them damages quality out of proportion to their count. The phenomenon becomes pronounced as models scale past a few billion parameters, and it is why naive 8-bit activation quantization that worked fine on small vision models fails on an LLM. The geometry is unforgiving. A symmetric per-tensor scale must cover the largest absolute value present. If one channel is dozens of times larger than the typical channel, then the typical channels are represented by only a few distinct levels out of the available range, and their relative error explodes. ## Why per-channel does not simply solve it For weights, per-output-channel scales are easy: each output channel's scale multiplies a whole row of the result, so it factors cleanly out of the accumulation. For activations, the outlier structure runs along the *input*-channel axis, which is the axis being summed over inside the matmul. A per-input-channel scale therefore cannot be pulled out of the accumulator — it would have to be applied inside the inner loop. That structural mismatch, not a lack of imagination, is why per-token scaling (which does factor out) is the granularity that hardware kernels actually offer. ## Outlier migration SmoothQuant's contribution is to notice that the difficulty can be moved rather than solved where it sits. For a linear layer computing `y = x @ W`, pick a per-input-channel factor `s` and compute `(x / s) @ (diag(s) @ W)` instead. The product is unchanged — it is a diagonal rescaling that cancels — but the activation's outlier channels have been shrunk and the corresponding weight rows enlarged. Since weights tolerate a much wider spread (they are static and finely grouped), you have traded an intractable activation problem for a tractable weight one. The factor is usually chosen with an exponent that balances how much difficulty each side absorbs. Crucially the division of `x` is not a runtime cost: it can be folded into the weights and biases of whatever produced `x` — the preceding normalization or linear layer — during the offline quantization step. At serving time the model just has different weights. ## Other levers Keeping a handful of known-outlier channels in higher precision and running them as a separate small matmul is the mixed-precision alternative; it works but fragments the kernel and is hardware-unfriendly. Rotation-based methods multiply activations and weights by paired orthogonal matrices to spread outlier energy across channels before quantizing, achieving something similar to migration without a hand-picked factor. ## What an interviewer is listening for The weak answer is "activations are dynamic". The strong answer adds the structural fact — persistent per-channel outliers — explains why the outlier axis is the one the matmul sums over, and names migration and per-token scaling as the two ways out.

  • Why does the smoothing rescale cost nothing at inference time?
    Because it is a diagonal transform that cancels: dividing the activation's input channels by s and multiplying the matching weight rows by s leaves the product identical. The division of the activation can be folded offline into the weights and biases of whatever produced it — the preceding normalization or linear layer — so the served model simply has different constants and issues no extra operation.
  • When would you use static activation ranges instead of per-token dynamic ones?
    When the runtime reduction is worth avoiding and the input distribution is genuinely narrow and known — a fixed-format classification or extraction workload, for instance, on hardware where fusing the scale computation into the kernel is awkward. The risk you accept is clipping on any input outside the calibrated range, so static scaling wants monitoring for saturation and a wider safety margin than the calibration maxima suggest.
  • Why not just clip the outlier channels and quantize what is left?
    Because those channels are not noise. They are consistent across tokens and layers and the model's downstream computation depends on their magnitude; truncating them degrades quality far more than their small count suggests. Everything useful in this area either preserves them at higher precision, rescales them into a manageable range, or rotates their energy across channels — never discards them.

saying these in an interview costs you the question

  • Says activation ranges can be baked in offline like weights
  • Treats outlier channels as random noise to clip
  • Claims weight-only 4-bit runs the matmul in 4-bit arithmetic
  • Thinks per-input-channel activation scales are free in the kernel
  • Believes outlier migration changes the layer's mathematical output

context