How do GPTQ and AWQ differ when quantizing the same weights to 4-bit?
answer
- both weight-only, both calibrated
- one compensates, one rescales
- importance measured from activations, not weights
- inverse-Hessian column updates
- avoids mixed precision on purpose
basics
~20 sGPTQ quantizes a layer's weights in sequence and uses second-order information from calibration activations to nudge the not-yet-quantized weights so they absorb each rounding error. AWQ instead searches for per-channel scale factors that protect the small fraction of channels the activations show to matter most.
solid answer
~60 sBoth are calibration-based, weight-only post-training methods, both usually run with group-wise scales, and both need only a few hundred samples — so the difference is in what they do with those samples. **GPTQ** is error compensation. Working one layer at a time, it quantizes weight columns in order and, after each one, updates the remaining unquantized weights using inverse-Hessian information estimated from the calibration activations, so the layer's output error introduced by that rounding is partly cancelled by weights not yet fixed. It descends from optimal brain surgeon-style pruning. **AWQ** is salience-aware scaling. Its observation is that a small fraction of weight channels matter disproportionately, and that you identify them by the magnitude of the *activations* flowing through them, not by the weights' own magnitude. Rather than keeping those channels in higher precision — which fragments the kernel — it searches a per-input-channel scale that reduces their relative rounding error, then quantizes everything at one uniform width. GPTQ does more work per layer and is more sensitive to calibration overfitting; AWQ is faster and simpler to reason about.
code
python · 14 linesimport torch
w = torch.randn(4096, 4096)
w[:, 0] *= 30 # one dominant input channel
def q_err(t, bits=4, group=None):
x = t.reshape(1, -1) if group is None else t.reshape(-1, group)
qmax = 2 ** (bits - 1) - 1
s = x.abs().amax(-1, keepdim=True) / qmax
deq = torch.clamp(torch.round(x / s), -qmax - 1, qmax) * s
return (deq - x).abs().mean().item()
print(q_err(w)) # one scale for the whole tensor
print(q_err(w, group=128)) # a scale per 128 weightsgo deeper
Recognize both as post-training, weight-only methods that need a small calibration set, and know that they differ in how they use it rather than in whether they retrain the model.
Explain each mechanism concretely: sequential column quantization with inverse-Hessian corrections on the remaining weights, versus a searched per-channel scale that protects channels the activations mark as salient. Know what group-wise scales are.
Show judgment about calibration sensitivity, the systems argument against mixed precision inside a tensor, and the fact that the ranking between them is empirical — describe how you would produce both and score them on your own traffic before choosing.
Own the standardization decision: one quantization recipe across a model portfolio has real value in tooling, kernels and revalidation cost, and that value usually outweighs chasing a small per-model advantage between two comparable algorithms.
## Common ground first GPTQ and AWQ are frequently presented as rivals, but they agree on almost everything structurally. Both are **post-training**: no gradients on the model weights, no training pipeline. Both are **weight-only**: the activations stay in a floating-point format and the low-bit weights are dequantized into the matmul. Both are **calibration-based**: a few hundred sequences are pushed through the model so the algorithm can see what the layer actually receives. Both normally quantize with **group-wise scales** — a separate scale for every group of, say, 128 weights along the input dimension — because a single scale per tensor is far too coarse at 4 bits. Both operate layer by layer, so a model much larger than one GPU's memory can be quantized on one GPU. The divergence is in the objective they optimize with the calibration statistics. ## GPTQ: compensate the error you just made GPTQ frames each linear layer as a reconstruction problem: choose quantized weights so that the layer's output on the calibration inputs stays as close as possible to the original layer's output. That objective is quadratic in the weights, and its curvature is captured by a Hessian built from the calibration activations. The algorithm proceeds through the weight columns in order. Each column is rounded to the grid, which introduces a known output error. GPTQ then applies a correction to the columns not yet quantized, derived from the inverse Hessian, so that those still-free weights move to values that partially cancel the error just committed. By the time the last column is quantized, most of the earlier damage has been absorbed by weights that were free to move at the time. This is a direct descendant of optimal brain surgeon / optimal brain quantization, made tractable for billion-parameter layers by processing columns in blocks and reusing a Cholesky factorization of the Hessian. Implementations expose an activation-ordering option that quantizes the most important columns first, which usually helps and slightly complicates the kernel's indexing. The costs are the ones you would expect from a second-order method: the Hessian estimate comes from a finite calibration sample, so the corrections can overfit it. Give GPTQ a calibration set that does not resemble production traffic and it will diligently optimize for the wrong distribution. ## AWQ: protect what the activations say is salient AWQ starts from a different observation: weight importance is not visible in the weights. Two channels with identical weight magnitudes contribute very differently to the output if one of them multiplies large activation values and the other multiplies near-zero ones. Ranking channels by average activation magnitude identifies a small salient fraction — on the order of one percent — whose rounding error dominates the layer's output error. The naive response is mixed precision: keep those channels in floating point and quantize the rest. AWQ argues against this on systems grounds — a tensor with a ragged precision layout defeats the packed kernels that make low-bit serving fast in the first place. Instead it keeps a single uniform bit width and applies a per-input-channel scaling: multiply the salient channels' weights up before rounding (and divide the corresponding activation channels down, folded away offline, exactly as in outlier migration), so those weights land on a finer effective grid relative to their own magnitude. The scaling exponent is found by a small grid search that minimizes the layer's output error on the calibration set. Because the search space is one scalar per layer rather than a full second-order correction, AWQ runs faster than GPTQ and depends on the calibration data mainly through a channel-magnitude statistic — a much lower-variance quantity than a Hessian — which is why it is generally considered the less calibration-sensitive of the two. ## SmoothQuant is a third thing, not a competitor It is worth keeping the families straight. SmoothQuant uses the same diagonal-rescaling identity AWQ uses, but for a different purpose: making *activations* quantizable so that weight-and-activation arithmetic becomes viable. GPTQ and AWQ are both about weight-only compression. They can be composed with activation-side techniques rather than chosen instead of them. ## Choosing in practice Run both, on a calibration set drawn from your own traffic, and evaluate on your own task suite — they are cheap enough that this is an afternoon, and the ranking between them is model- and width-dependent rather than universal. Beyond that, the practical differentiators are ecosystem ones: which kernel your serving stack has optimized, whether the group size and activation-ordering variant you want is supported end to end, and whether a vendor QAT checkpoint at the same width exists and makes the question moot. ## What a strong answer contains Name the mechanism of each in one clause — error compensation via inverse-Hessian updates versus activation-informed per-channel scaling — say that both are weight-only PTQ on group scales, and finish with the honest conclusion that the choice is empirical and stack-dependent rather than settled.
- What does a group size of 128 mean, and what does shrinking it buy you?It means one scale (and zero-point, if asymmetric) is shared by every 128 consecutive weights along the input dimension, rather than by the whole tensor or a whole row. Shrinking the group gives each scale a narrower range to cover, so rounding error drops — at the price of storing more scale metadata and, below some size, kernels that unpack less efficiently. 128 and 64 are the common compromises.
- Why does AWQ deliberately avoid keeping salient channels in floating point?Because mixed precision inside a single tensor breaks the packed layout that low-bit kernels rely on. You end up issuing a second, small, irregular matmul and reassembling the result, which can erase the speed advantage that motivated quantizing at all. Applying a scale instead keeps every weight at one uniform width and one memory layout, so the fast kernel still applies.
- Is one of them simply better, or does the answer depend?It depends, and honest practitioners say so. The ranking shifts with model family, bit width, group size and the calibration corpus, and the two are close enough that ecosystem support — which kernel your serving stack has tuned, which variant your toolchain packages — often decides. The right move is to produce both from your own calibration data and score them on your own task suite.
saying these in an interview costs you the question
- Says AWQ keeps one percent of weights in floating point
- Describes GPTQ as a fine-tuning or training method
- Claims either method also quantizes the activations
- Assumes a bigger calibration set always improves GPTQ
- Treats the two names as labels for the same algorithm