Which PyTorch API quantizes a model for on-device ExecuTorch use today?
answer
- two generations of the API
- the runtime brand changed too
- the graph comes from torch.export
- prepare_pt2e and convert_pt2e, via torchao
- quantize_dynamic and optimize_for_mobile are legacy
basics
~10 sPT2 export quantization: capture the model with torch.export, then run prepare_pt2e and convert_pt2e from torchao, driven by a backend quantizer such as XNNPACKQuantizer. The older eager torch.ao.quantization path is deprecated.
solid answer
~40 sOn-device PyTorch means ExecuTorch, and ExecuTorch consumes an **exported graph**, so quantization happens on that graph rather than on a live `nn.Module`. The current flow is: `torch.export.export(model, example_inputs).module()` to capture the graph, build a backend quantizer (`XNNPACKQuantizer` with `get_symmetric_quantization_config(...)`), then `prepare_pt2e` → run calibration or QAT fine-tuning → `convert_pt2e`. The quantized graph is exported again and lowered to a `.pte`. The predecessor was eager-mode quantization in `torch.ao.quantization` — `quantize_dynamic`, `qconfig`/`QuantStub`, plus `torch.utils.mobile_optimizer.optimize_for_mobile` for the TorchScript lite interpreter. That module is deprecated in favour of `torchao`, and "PyTorch Mobile" is a retired brand name; recognise the old calls in legacy code, but do not start there.
code
python · 19 linesimport torch
from torch import nn
from executorch.backends.xnnpack.quantizer.xnnpack_quantizer import (
XNNPACKQuantizer,
get_symmetric_quantization_config,
)
from torchao.quantization.pt2e.quantize_pt2e import prepare_pt2e, convert_pt2e
model = nn.Sequential(nn.Conv2d(3, 8, 3), nn.ReLU(), nn.AdaptiveAvgPool2d(1)).eval()
example_inputs = (torch.randn(1, 3, 32, 32),)
graph = torch.export.export(model, example_inputs).module()
quantizer = XNNPACKQuantizer().set_global(
get_symmetric_quantization_config(is_per_channel=True)
)
prepared = prepare_pt2e(graph, quantizer)
for _ in range(16):
prepared(torch.randn(1, 3, 32, 32))
quantized = convert_pt2e(prepared)go deeper
Recall the shape of the flow: export the model, prepare, calibrate, convert, then lower. Name prepare_pt2e and convert_pt2e, and know that quantize_dynamic belongs to the older, deprecated API.
Explain why the API moved: eager quantization only saw modules and needed QuantStub boilerplate, while the exported graph exposes every op so a quantizer can annotate patterns directly.
Show you track deprecations in production. Be ready to say what a port from the eager path costs, what breaks, and how you would validate that the ported model matches the old one numerically.
Own the migration decision across a model portfolio: when to freeze legacy models on the old path versus re-qualifying them, and how to keep one quantization pipeline rather than two.
## The question behind the question An interviewer asking this is usually checking whether your on-device knowledge is current. PyTorch's mobile story has been rewritten twice: the runtime brand changed (PyTorch Mobile → ExecuTorch) and the quantization API changed (eager-mode `torch.ao.quantization` → PT2 export quantization in `torchao`). A candidate who answers "call `quantize_dynamic` and then `optimize_for_mobile`" is describing a path the vendor has deprecated. ## The current path, step by step ExecuTorch is ahead-of-time: nothing about your Python model class survives to the phone. What ships is a `.pte` file produced from an **exported graph** — a flat, traced representation of the operations. Quantization therefore has to be a graph transformation, applied before lowering: 1. **Capture.** `torch.export.export(model, example_inputs).module()` gives you a graph module you can still run in Python. 2. **Choose a quantizer.** A quantizer is a backend-specific object that decides *which* nodes get quantized and *how*. For CPU inference that is `XNNPACKQuantizer`, configured with `get_symmetric_quantization_config(is_per_channel=True)` and attached with `set_global(...)`. 3. **Prepare.** `prepare_pt2e(graph, quantizer)` inserts observers at the annotated tensors. 4. **Calibrate or fine-tune.** Run representative inputs through the prepared module (static PTQ), or fine-tune it if you used `prepare_qat_pt2e`. 5. **Convert.** `convert_pt2e(prepared)` replaces observers with quantize/dequantize op pairs carrying concrete scales and zero-points. 6. **Lower.** Export again and lower to `.pte` for the target backend. The imports that matter: `torchao.quantization.pt2e.quantize_pt2e` for `prepare_pt2e` / `convert_pt2e` / `prepare_qat_pt2e`, and `executorch.backends.xnnpack.quantizer.xnnpack_quantizer` for `XNNPACKQuantizer` and `get_symmetric_quantization_config`. ## What the old path looked like Eager-mode quantization worked on the module tree, not a graph. You either called `torch.ao.quantization.quantize_dynamic(model, {nn.Linear}, dtype=torch.qint8)` for a one-liner dynamic conversion, or you inserted `QuantStub`/`DeQuantStub` into your `forward`, set a `qconfig`, called `prepare`/`convert`, and — for the old lite interpreter — passed the TorchScript module through `torch.utils.mobile_optimizer.optimize_for_mobile` before saving. Every one of those symbols still imports today; presence is not currency. `torch.ao.quantization` carries a deprecation notice pointing at `torchao`. ## Why the migration happened at all Eager quantization had two structural problems. First, it could only see modules, so anything expressed as a functional call in `forward` (a `+`, a `torch.matmul`, a `F.relu`) was invisible to it — hence the `QuantStub` boilerplate and the `FloatFunctional` shims. Second, its module swaps produced objects that were awkward to export; the quantized model and the deployable model were built by two different mechanisms. PT2 export quantization fixes both: the graph shows every operation, so the quantizer can annotate patterns like conv→batchnorm→relu directly, and the artifact it produces is the same exported graph the lowering pipeline already consumes. It also decouples "what to quantize" (the quantizer, owned by the backend) from "how to do it" (`prepare_pt2e`/`convert_pt2e`, owned by the framework), which is why a new accelerator can ship a quantizer without patching the core flow. ## What to say, and what not to Say: quantization for on-device PyTorch is a graph pass over an exported model, done with `prepare_pt2e`/`convert_pt2e` and a backend quantizer, before lowering to `.pte`. Mention that the eager API is the deprecated predecessor and that you would recognise it in an old repo. Do not say that quantization is a flag on the runtime or the interpreter — the `.pte` you ship is already quantized; the phone just executes int8 kernels. Do not say you quantize after lowering; by then the backend has already decided which subgraphs it can take, and that decision depends on the quantization annotations being present. And do not describe the on-device runtime as "PyTorch Mobile" — the TorchScript lite-interpreter path is the legacy alternative, not the default. ## Version caveat The exact import path for `prepare_pt2e` has moved once already (from `torch.ao.quantization.quantize_pt2e` to `torchao.quantization.pt2e.quantize_pt2e`), and `XNNPACKQuantizer` is imported from the ExecuTorch backend package. When you name these in an interview, say which release you are describing; the flow shape is stable even where the module path is not.
- Where in the pipeline does quantization sit relative to lowering the model for a backend?Before it. You quantize the exported graph, then export the converted graph again and lower it. The backend partitioner looks for quantized patterns it can claim; if the quantize/dequantize nodes are not there yet, it simply claims float subgraphs and you get a float model.
- If you inherit a repo that calls quantize_dynamic and optimize_for_mobile, what does that tell you?That it targets the legacy TorchScript lite-interpreter path, not ExecuTorch. Both symbols still import, so it will keep working, but it is the deprecated predecessor: eager module-swapping quantization plus a TorchScript-level optimizer. A port means re-expressing it as export plus prepare_pt2e/convert_pt2e with a backend quantizer.
- Does the phone need any quantization library at runtime?No. All the numeric decisions — scales, zero-points, which ops run in int8 — are baked into the .pte ahead of time. The device-side runtime just executes the kernels the delegate provides. That is the whole point of an ahead-of-time flow: no calibration, no observers and no Python on device.
saying these in an interview costs you the question
- Says optimize_for_mobile is the current mobile export step
- Describes the on-device runtime as PyTorch Mobile
- Thinks quantization is a runtime flag on the interpreter
- Claims you quantize after lowering to .pte
- Assumes eager qconfig/QuantStub is still the recommended API