skip to content

In ExecuTorch int8 quantization, when do you need calibration data?

level: middleimportance: must knowfreq 52%

answer

  1. the question is only about activations
  2. weights need no data either way
  3. one variant decides the scale at runtime
  4. fixed scales enable cross-op integer fusion
  5. matmul-heavy versus conv-heavy models differ

basics

~20 s

Only for static quantization, where activation scales are frozen ahead of time and must come from observed data. Dynamic quantization quantizes weights offline and computes activation scales at runtime per input, so no calibration set is required.

solid answer

~40 s

The split is about **activations**, not weights — weights are known offline either way. In **static** PTQ you fix an activation scale ahead of time, so you must run representative inputs through the prepared module and let observers record the ranges; without that pass the scales are meaningless. In **dynamic** PTQ, configured with `get_symmetric_quantization_config(is_dynamic=True)`, the kernel measures each incoming activation tensor at runtime and derives its scale on the spot, so calibration data is unnecessary. The trade: dynamic is nearly free to set up and robust to input distribution shift, but pays a per-inference cost and only helps ops where the weights dominate — linear/matmul-heavy models. Static is what you want for convnets, because a fixed scale lets the backend chain int8 convs without dequantizing between them.

code

python · 14 lines
python
from executorch.backends.xnnpack.quantizer.xnnpack_quantizer import (
    XNNPACKQuantizer,
    get_symmetric_quantization_config,
)

# Static: activation scales frozen from a calibration pass
static_q = XNNPACKQuantizer().set_global(
    get_symmetric_quantization_config(is_per_channel=True)
)

# Dynamic: activation scales derived per call, no calibration data
dynamic_q = XNNPACKQuantizer().set_global(
    get_symmetric_quantization_config(is_per_channel=True, is_dynamic=True)
)

go deeper

for a junior

Know the one-line distinction: static freezes activation scales from a calibration pass, dynamic computes them at runtime and needs no data. Weights are quantized offline in both.

for a middle

Explain why fixed scales let a backend chain int8 ops without round-tripping through float, and why dynamic pays a per-inference scan that only amortises over large matmuls.

for a senior

Show judgment about the calibration set as production artifact: coverage, drift, and the fact that a badly calibrated static model fails silently rather than loudly.

for a principal

Decide the policy across a model fleet — which model families default to which variant, who owns the calibration corpus, and when drift triggers requantization and revalidation.

## Weights are easy, activations are the problem After training, weights are constants. You can look at them, compute a scale, and store int8 values — no data needed, ever. Activations are different: their range depends on the input, and you do not have the input until inference. Every quantization variant is really an answer to "what do we do about activation scales?", and the answer decides whether you need calibration data. ## Static PTQ: freeze the scale ahead of time Static quantization commits to one scale per activation tensor, chosen offline. To choose it you must observe realistic activations, which is what the calibration pass does: run the prepared module forward on representative inputs so its observers accumulate min/max or a histogram, then let `convert_pt2e` turn those into constants. What you get for that effort is the fastest option on device. Because the scale of a conv's output is already known, the backend can hand int8 straight to the next int8 op without a dequantize/requantize round trip. Whole chains — conv, batchnorm, relu, conv — collapse into fused integer kernels. This is why static is the default choice for CNNs and for anything where convolutions dominate the runtime. The cost is a dependency on the calibration set. It needs to look like production traffic: same preprocessing, same resolution, same lighting or speaker or language mix. A few dozen to a few hundred batches is usually plenty; the number matters far less than the representativeness. Ranges that are too narrow clip real inputs; ranges stretched by one freak outlier waste resolution on values that never occur, which is the argument for a histogram-based observer over plain min/max. ## Dynamic PTQ: compute the scale at runtime Dynamic quantization stores weights as int8 and leaves the activation scale to inference time: the kernel scans the incoming tensor, computes its range, quantizes, multiplies in integer, and rescales the output. Enable it in the ExecuTorch flow with `get_symmetric_quantization_config(is_per_channel=True, is_dynamic=True)`. Because the scale adapts per call, there is nothing to calibrate — you skip the data pass entirely, and the model cannot be embarrassed by an input distribution it never saw. That robustness is real: dynamic quantization degrades gracefully where a badly calibrated static model falls off a cliff. The cost has two parts. First, the per-tensor scan and rescale are real work on every inference; for a small op, that overhead can exceed the integer speedup. Second, the model cannot stay in the integer domain across ops, since each op re-derives its own scale, so you lose the fusion that makes static fast. Dynamic wins when the arithmetic is dominated by large weight matrices — transformer and RNN-style stacks, embedding-plus-linear models. There the memory traffic of the weights is the bottleneck, 4x smaller weights is the whole win, and the runtime scan is amortised over a large matmul. It is also the pragmatic first move when you have no calibration data available at all. ## Weight-only, the third point on the line Some flows quantize weights alone and dequantize them into float for the multiply. This buys the download and memory-footprint win with essentially no accuracy risk and no calibration, but little or no compute speedup, because the arithmetic is still float. It is a size play, not a latency play. ## Choosing, in practice A workable decision order: 1. **Convolutional vision model with a latency budget?** Static, per-channel weights, with a calibration set drawn from production data. 2. **Linear/matmul-heavy model, or no calibration data at hand?** Dynamic — cheap to try and hard to get badly wrong. 3. **Size is the constraint, latency is fine?** Weight-only. 4. **Static chosen but accuracy unacceptable?** Fix calibration coverage and per-channel weights first; escalate to QAT only after that. Always confirm the choice on device rather than from the table above: which variants a backend actually accelerates differs by backend and by operator, and a config the delegate cannot execute becomes explicit quantize/dequantize work in float — slower than the model you started with. ## Interview traps The most common wrong answer is "dynamic quantization needs a smaller calibration set". It needs none; that is the definition. The second is treating calibration as training — no labels, no loss, no backward pass; it is a forward-only instrumentation run. The third is assuming dynamic is always safer, which ignores that its per-inference scan can wipe out the speedup on small ops and that it forecloses cross-op integer fusion.

  • How much calibration data is enough for static quantization?
    Representativeness matters far more than volume — a few dozen to a few hundred batches drawn from real traffic usually saturates the observers. What breaks a calibration set is coverage: missing a lighting condition, an input resolution or a rare-but-real class leaves ranges too narrow, and those inputs clip at inference.
  • Why can dynamic quantization end up slower than the float model on a small operator?
    Each call must scan the activation tensor to find its range, quantize it, then rescale the output. On a small matmul that bookkeeping can cost more than the integer arithmetic saves, and because every op re-derives its own scale, the backend cannot keep values in int8 between ops.
  • Does calibration require labels or a backward pass?
    Neither. It is a forward-only instrumentation run: you push representative inputs through the prepared module so observers accumulate statistics. No loss is computed and no parameter changes. That is what distinguishes it from QAT, where the same-looking loop really is fine-tuning.
  • What happens if the deployed input distribution drifts away from your calibration set?
    Static activation scales start clipping or wasting range, and accuracy degrades quietly — no error is raised. That is a monitoring obligation: watch output distributions on device, and requantize with a refreshed calibration set when the input pipeline or the capture hardware changes.

saying these in an interview costs you the question

  • Says dynamic quantization needs a small calibration set
  • Treats calibration as a training run with labels
  • Thinks weights need calibration data too
  • Claims static is always faster regardless of operator mix
  • Uses whatever calibration images are lying around

context