skip to content

Quantization for Mobile

Mobile inference is usually memory- and latency-bound, so weights get squeezed from float32 to int8. The practical distinctions are dynamic vs static post-training quantization, when calibration data is required, and when accuracy loss forces quantization-aware training — now expressed through the PT2 export flow rather than the deprecated eager quantization API.

on this pageshow

questions

6

Why does PT2 export quantization need a backend quantizer like XNNPACKQuantizer?

level: middleimportance: must knowfreq 48%

answer

  1. policy versus mechanism
  2. the framework has no opinion here
  3. only the target knows its fusable patterns
  4. annotating unsupported ops costs latency
  5. scope it per module when needed

basics

~20 s

Because only the backend knows which operator patterns it can actually execute in int8. The quantizer encodes that contract, annotating exactly the nodes the delegate can fuse; quantizing anything else adds quantize/dequantize work around an op that still runs in float.

solid answer

~40 s

`prepare_pt2e`/`convert_pt2e` are mechanism, not policy — they insert observers and rewrite the graph, but they have no opinion about *what* should be quantized. That opinion is backend-specific: XNNPACK can run a fused int8 conv-bn-relu, a different accelerator supports a different set, and each has its own constraints on symmetric versus affine, per-channel weights, and supported dtypes. `XNNPACKQuantizer` is that knowledge as an object: it walks the graph, matches the patterns XNNPACK implements, and attaches quantization specs to them. You shape it with `set_global(get_symmetric_quantization_config(...))` and can override a submodule with `set_module_name(...)`. Use the wrong quantizer, or annotate ops the delegate cannot claim, and lowering leaves literal quantize/dequantize nodes around a float op — measurably slower than the float model you started with.

code

python · 11 lines
python
from executorch.backends.xnnpack.quantizer.xnnpack_quantizer import (
    XNNPACKQuantizer,
    get_symmetric_quantization_config,
)

global_cfg = get_symmetric_quantization_config(is_per_channel=True)

quantizer = XNNPACKQuantizer()
quantizer.set_global(global_cfg)
# keep one numerically sensitive submodule out of the int8 region
quantizer.set_module_name("classifier", None)

go deeper

for a junior

Know that a quantizer object is a required argument to prepare_pt2e, and that you pick the one matching your target backend rather than a generic one.

for a middle

Explain the policy/mechanism split: the quantizer annotates which patterns to quantize based on what the backend can fuse, while prepare/convert only carry out the rewrite.

for a senior

Be ready to diagnose a mismatch in production: unclaimed subgraphs, leftover quantize/dequantize pairs around float ops, and a quantized model that benchmarks slower than float.

for a principal

Own backend portability as a strategy question — how many quantization configurations the team maintains across device tiers, and what it costs to requalify accuracy for each target.

## Separating policy from mechanism The PT2 export quantization design splits the job in two. `prepare_pt2e` and `convert_pt2e` own the **mechanism**: insert observers, read them back, compute scales, rewrite the graph with quantize/dequantize ops. A **quantizer** owns the **policy**: which nodes get quantized, to what dtype, symmetric or affine, per-tensor or per-channel, with which observer, and which nodes are deliberately left in float. That split exists because policy is not a property of PyTorch — it is a property of the hardware and kernel library you are targeting. A new accelerator can ship its own quantizer without changing anything in the framework. ## What the quantizer actually knows A backend quantizer encodes a kernel library's real capability surface: - **Which patterns are fusable.** XNNPACK implements conv→batchnorm→relu as one int8 kernel; the quantizer annotates the whole pattern so the pieces stay in the integer domain instead of dequantizing between them. - **What numerics the kernels accept.** Symmetric int8 weights, a particular qmin/qmax, per-channel weight scales but per-tensor activation scales — these are kernel constraints, not preferences. - **What is unsupported.** Ops the library has no int8 kernel for are left un-annotated on purpose, so they stay float and the graph does not pay for a pointless round trip. That is why `get_symmetric_quantization_config()` lives beside `XNNPACKQuantizer` in the same package: the config's knobs (`is_per_channel`, `is_dynamic`, `is_qat`) select among the shapes that backend can actually run. ## The failure this prevents Suppose you annotate an operator the delegate cannot execute in int8. `convert_pt2e` will happily wrap it in a quantize/dequantize pair — it does what it is told. During lowering, the partitioner inspects the graph, cannot claim that subgraph, and leaves it for the portable kernels. What runs on device is: quantize the tensor to int8, immediately dequantize it back to float, run the original float op, quantize the result again for the next boundary. Every one of those steps is real arithmetic and real memory traffic that the unquantized model never performed. The output is numerically *worse* (you took the rounding hit) and the latency is *higher*. This is the single most common reason a candidate reports "we quantized and it got slower", and it is precisely what a backend-matched quantizer is designed to make impossible. ## Scoping the policy The quantizer is configurable, not all-or-nothing: - `set_global(config)` applies one config to everything the quantizer recognises — the usual starting point. - `set_module_name(name, config)` overrides a specific submodule by its qualified name. This is the lever for excluding a numerically sensitive block (a final classifier head, an attention softmax path) while quantizing everything else. - `set_module_type(type, config)` does the same for all instances of a module class. These scoping calls are how a real deployment is tuned: quantize globally, find the two layers that cost accuracy, exempt them, re-measure. ## Ordering constraints The quantizer runs against the exported graph **before** lowering, and that order is not negotiable. The partitioner's claim decisions depend on the quantize/dequantize pattern already being present; if you lower first and then try to quantize, the graph the backend consumed is already fixed. Similarly, you cannot reuse annotations across backends: a graph annotated for one kernel library is not automatically valid for another, because the fusable pattern sets differ. Retargeting means re-running the flow with that backend's quantizer. ## How to verify you got it right Do not trust the config — check the outcome. Inspect the lowered artifact and confirm the quantized region was actually claimed by the delegate rather than left to portable kernels. Then benchmark on a real device: a correctly matched quantizer produces a clear latency drop on CPU-bound models, while a mismatch shows up as latency at or above the float baseline. Comparing float and quantized intermediate outputs tells you where the numerics diverged; comparing per-op timings tells you where the fusion did not happen. ## The one-sentence version The quantizer is the backend's contract, expressed as graph annotations: quantize what this hardware can execute in integer, leave everything else alone — because quantization that the target cannot execute is pure cost.

  • What exactly goes wrong if you quantize an operator the delegate cannot execute in int8?
    The converted graph carries a quantize/dequantize pair around it, but no int8 kernel claims the pattern during lowering. On device that becomes quantize, dequantize, run the op in float — extra arithmetic and memory traffic on top of the rounding error. You get worse numerics and worse latency simultaneously.
  • Can you take a graph annotated for one backend and lower it for a different one?
    Not safely. The annotations encode one kernel library's fusable patterns and numeric constraints; another backend supports a different set. The right move is to re-run prepare/convert with that backend's own quantizer, and to re-measure accuracy, since a different scheme can shift results.
  • Why is the quantizer applied before lowering rather than after?
    Because the partitioner decides which subgraphs it can claim by pattern-matching the quantized graph. The quantize/dequantize structure has to exist for an int8 kernel to be selected. Lowering first fixes the backend's decisions against a float graph, and there is nothing left for quantization to influence.
  • How would you exclude a single sensitive layer from quantization?
    Scope the policy: set_global for the bulk of the model, then set_module_name (or set_module_type) to give that submodule a different config or none at all. Then re-measure — exempting a layer costs some of the speedup, so it should be justified by a real accuracy recovery.

saying these in an interview costs you the question

  • Thinks prepare_pt2e decides what to quantize by itself
  • Believes any quantizer works with any backend
  • Assumes unsupported ops silently stay float with no cost
  • Applies quantization after lowering to the target
  • Quantizes everything and never checks what the delegate claimed

context

open as a page

In ExecuTorch int8 quantization, when do you need calibration data?

level: middleimportance: must knowfreq 52%

basics

~20 s

Only for static quantization, where activation scales are frozen ahead of time and must come from observed data. Dynamic quantization quantizes weights offline and computes activation scales at runtime per input, so no calibration set is required.

open as a page

What do prepare_pt2e and convert_pt2e do to an exported PyTorch graph?

level: middleimportance: must knowfreq 60%

basics

~20 s

prepare_pt2e inserts observer modules at the tensors a quantizer annotated, so a calibration or training run records their value ranges. convert_pt2e then replaces each observer with quantize/dequantize op pairs carrying the scale and zero-point it computed.

open as a page

Which PyTorch API quantizes a model for on-device ExecuTorch use today?

level: juniorimportance: should knowfreq 45%

basics

~10 s

PT2 export quantization: capture the model with torch.export, then run prepare_pt2e and convert_pt2e from torchao, driven by a backend quantizer such as XNNPACKQuantizer. The older eager torch.ao.quantization path is deprecated.

open as a page

Int8 PTQ cost you 4 points of accuracy on device — how do you recover it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Work cheapest-first: verify calibration coverage, switch weights to per-channel, then find the sensitive layers by comparing float and quantized intermediates and exempt them. Only if that falls short move to QAT with prepare_qat_pt2e and a short fine-tune.

open as a page

On a phone latency and app-size budget, do you quantize or ship a smaller model?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Usually both, in order: fix the budgets and accuracy floor first, take int8 quantization because it is the cheapest lever, and change architecture when quantization alone misses the target. They compose — a smaller model quantizes too.

open as a page