How does a PyTorch nn.Module become an ExecuTorch .pte file?
answer
- three ahead-of-time stages, no Python on device
- capture, lower, emit
- ExportedProgram, then Edge dialect
- one-step lowering call, not two
- to_executorch().buffer written as .pte
basics
~20 sThree ahead-of-time steps: torch.export.export captures the model as an ExportedProgram graph, to_edge_transform_and_lower converts it to the Edge dialect and hands supported subgraphs to a backend, and to_executorch emits the flatbuffer you write out as .pte.
solid answer
~40 sOn-device PyTorch is an ahead-of-time pipeline, because the phone has no Python. First `torch.export.export(model.eval(), example_args)` captures the whole model into an `ExportedProgram` — a full-graph ATen-dialect FX graph with weights lifted out and shape guards recorded. Then `to_edge_transform_and_lower(ep, partitioner=[XnnpackPartitioner()])` lowers to the Edge dialect and lets the backend claim the subgraphs it can compile, returning an `EdgeProgramManager`. Then `.to_executorch()` runs memory planning and emits an `ExecutorchProgramManager`, whose `.buffer` you write to `model.pte`. That single file carries the instruction stream, constant weights, any delegate blobs and the memory plan, and the C++ runtime loads it directly. Note that the older two-step `to_edge()` + `to_backend()` still works but emits a deprecation warning; the one-step call is the recommended path.
code
python · 22 linesimport torch
from torch.export import export
from executorch.exir import to_edge_transform_and_lower
from executorch.backends.xnnpack.partition.xnnpack_partitioner import XnnpackPartitioner
class Net(torch.nn.Module):
def __init__(self):
super().__init__()
self.fc = torch.nn.Linear(8, 2)
def forward(self, x):
return torch.relu(self.fc(x))
model = Net().eval()
example = (torch.randn(1, 8),)
exported = export(model, example)
edge = to_edge_transform_and_lower(exported, partitioner=[XnnpackPartitioner()])
program = edge.to_executorch()
with open("model.pte", "wb") as f:
f.write(program.buffer)go deeper
Know that a mobile app cannot run Python, so the model is converted ahead of time into a single .pte file, and that the conversion happens in your build environment, not on the phone.
Be ready to name the three calls in order and say what each produces: torch.export.export gives an ExportedProgram, to_edge_transform_and_lower gives an EdgeProgramManager with backend subgraphs claimed, to_executorch gives the buffer you write as .pte.
Show that you treat export as a build step that must be gated: eval mode enforced, output parity checked against eager in CI, and the artifact regenerated whenever the model or the torch version moves. Explain what memory planning buys you at runtime.
Own where the export step lives in the release pipeline — who regenerates artifacts, how model versions are pinned against app versions, and what the rollback story is when a re-export changes numerics for users already on the old binary.
## Why there is a pipeline at all Eager PyTorch is Python: `forward()` runs Python bytecode, dispatches through the ATen operator library, and allocates as it goes. A phone app has no Python interpreter and a hard binary-size budget, so on-device PyTorch splits the work in two. Everything dynamic and expensive happens **ahead of time** on your laptop or CI machine and produces one self-contained artifact; on the device, a small C++ runtime just walks that artifact. ExecuTorch is that runtime, and `.pte` (PyTorch ExecuTorch program) is that artifact. ## Stage 1 — capture: `torch.export.export` ``` ep = torch.export.export(model.eval(), (example_input,)) ``` This produces an `ExportedProgram`. Three things are worth knowing about it: - It is a **full-graph** capture. Unlike `torch.compile`, which is allowed to "graph break" and fall back to Python for anything it cannot trace, export must capture the entire model or fail loudly. There is no Python left to fall back to on device. - Parameters and buffers are **lifted** out of the module and become explicit graph inputs, recorded in a graph signature. The graph is a pure function; the weights travel alongside it. - It is traced with your **example inputs**, so concrete shapes and dtypes are baked in as guards unless you explicitly declare dimensions dynamic. Call `model.eval()` first. Exporting in training mode captures dropout and the batch-norm training branch, and the exported model then produces different numbers on device than it did in your validation loop — a silent-wrong-answer bug, not a crash. ## Stage 2 — lower: `to_edge_transform_and_lower` ``` edge = to_edge_transform_and_lower(ep, partitioner=[XnnpackPartitioner()]) ``` This does two jobs at once. It converts the ATen graph to the **Edge dialect** — the same operators, but constrained to what an edge runtime can serialize and execute (explicit dtypes, no Python objects, no training-only constructs) — and it runs the **partitioner**, which tags the subgraphs the chosen backend can accept. Each claimed subgraph is compiled by that backend into an opaque blob and replaced in the graph by a call-delegate node. Whatever the backend does not claim stays as ordinary Edge operators. The result is an `EdgeProgramManager`, a container around one or more exported methods. ## Stage 3 — emit: `to_executorch` ``` prog = edge.to_executorch() with open("model.pte", "wb") as f: f.write(prog.buffer) ``` `to_executorch()` runs **memory planning**: because the graph is static, the toolchain can compute up front how large the intermediate-tensor arena must be and where each intermediate lives inside it, so the runtime does almost no dynamic allocation. It returns an `ExecutorchProgramManager`; `.buffer` is the serialized program bytes. A `.pte` therefore contains far more than weights — the instruction stream, the constant tensors, any delegate blobs, the memory plan, and the signature of each exported method. ## The deprecated two-step Older tutorials show `to_edge(ep)` followed by `edge.to_backend(XnnpackPartitioner())`. That path still runs, but ExecuTorch prints a deprecation warning naming the `to_edge() + to_backend()` workflow and pointing at `to_edge_transform_and_lower`. The one-step call interleaves the edge transformation passes with partitioning so the backend sees a graph in the shape it expects; the split version could hand the partitioner a graph that had already been rewritten in ways it did not recognise. If you inherit code using the two-step form, migrating is usually a one-line change. ## Verifying before you ship You do not need a phone to smoke-test the artifact. The Python runtime bindings load the same `.pte`: ``` from executorch.runtime import Runtime method = Runtime.get().load_program("model.pte").load_method("forward") out = method.execute([example_input]) ``` Compare that output against the eager model's on the same input, with a tolerance — numerical drift from a delegate is normal, a completely different answer means you exported the wrong thing (training mode, a frozen branch, wrong preprocessing). ## Where it commonly goes wrong Export is the stage that fails, and it fails for a small set of reasons: data-dependent Python control flow that cannot be captured, shapes that got specialized to your example input, and operators with no Edge-dialect or backend implementation. None of these are runtime problems — they surface on your build machine, which is the whole point of doing the work ahead of time.
- What is actually inside a .pte file besides the weights?A serialized flatbuffer program: the instruction stream for each exported method, the constant tensors, any compiled backend delegate blobs, the ahead-of-time memory plan describing the intermediate-tensor arena, and the input/output signature of each method. That is why the device runtime can execute it with almost no dynamic allocation and no compiler present.
- Why does ExecuTorch warn when you use to_edge() followed by to_backend()?Because that two-step workflow is deprecated in favour of to_edge_transform_and_lower. The one-step call interleaves the edge dialect transformation passes with partitioning, so the backend partitioner inspects the graph in the form it expects instead of one that edge passes have already rewritten. Migrating is normally a single-line change and removes the warning.
- How would you check a .pte is correct before wiring it into a mobile app?Load it with the Python runtime bindings — Runtime.get().load_program(path).load_method("forward") — run the same inputs through it and through the eager model, and compare with a numeric tolerance. Small drift from a delegate is expected; a large divergence usually means the model was exported in training mode or a branch was frozen at export time.
- Can one .pte hold more than one entry point?Yes. The export flow accepts a dictionary of named methods, and the resulting program carries each one with its own signature and memory plan. On device you load the program once and then load the method you want by name, which is how models with separate encode and decode entry points, or a reset method alongside forward, are shipped as a single file.
saying these in an interview costs you the question
- Thinks the device runs Python or a TorchScript interpreter
- Believes a .pte is just a state_dict of weights
- Names torch.jit.trace as the ExecuTorch entry point
- Uses the deprecated to_edge() plus to_backend() two-step
- Exports without model.eval(), keeping dropout and BN training behaviour