skip to content

What does XnnpackPartitioner do to the graph when lowering an ExecuTorch model?

level: seniorimportance: should knowfreq 45%

answer

  1. the backend's admissions policy over the graph
  2. claimed subgraphs become opaque blobs
  3. one call-delegate node per claimed region
  4. the rest falls back to runtime kernels
  5. partial delegation is normal; measure coverage

basics

~20 s

It walks the Edge-dialect graph and claims the subgraphs XNNPACK can execute. Each claimed subgraph is compiled into an opaque delegate blob and replaced by a call-delegate node; everything it declines stays as ordinary operators run by the runtime's own kernels.

solid answer

~50 s

A partitioner is the backend's admissions policy. During `to_edge_transform_and_lower`, `XnnpackPartitioner` inspects the Edge graph and tags the operator groups XNNPACK supports for the given dtypes and shapes. Each tagged group is handed to the backend's ahead-of-time compilation step, which turns it into an opaque binary blob stored in the `.pte`; in the graph it becomes an `executorch_call_delegate` node. Anything the partitioner declines — an unsupported operator, an unsupported dtype, sometimes a dynamic dimension — remains as regular Edge operators executed by the runtime's portable or optimized kernel library. So delegation is normally **partial**, and the interesting production question is how partial: many small delegated islands separated by fallback operators cost you boundary overhead and can be slower than you expect. Core ML and Vulkan work the same way through their own partitioners, and you can pass several so different subgraphs land on different backends.

code

python · 19 lines
python
import torch
from torch.export import export
from executorch.exir import to_edge_transform_and_lower
from executorch.backends.xnnpack.partition.xnnpack_partitioner import XnnpackPartitioner

class Block(torch.nn.Module):
    def __init__(self):
        super().__init__()
        self.conv = torch.nn.Conv2d(3, 8, 3, padding=1)

    def forward(self, x):
        return torch.relu(self.conv(x))

ep = export(Block().eval(), (torch.randn(1, 3, 32, 32),))
edge = to_edge_transform_and_lower(ep, partitioner=[XnnpackPartitioner()])

graph = edge.exported_program().graph_module.graph
for node in graph.nodes:
    print(node.op, node.target)

go deeper

for a junior

Know that lowering hands parts of the model to a hardware backend such as XNNPACK, and that whatever the backend cannot take still runs, just on the runtime's own CPU kernels.

for a middle

Explain the mechanism: the partitioner tags supported subgraphs, each is compiled ahead of time into an opaque blob stored in the .pte, and the graph keeps a call-delegate node in its place.

for a senior

Show that you measure delegated coverage and fragmentation rather than trusting that lowering succeeded, and that you keep the set of backends in the app build in sync with the backends the artifact was lowered to.

for a principal

Own the backend strategy across the device fleet — which accelerators you target per platform, whether you ship per-platform artifacts, and how partitioner choice and ordering are validated against measured latency rather than assumed.

## The role of a partitioner ExecuTorch does not assume any single backend can run a whole model. Instead each backend ships a **partitioner** whose job is to look at the Edge-dialect graph and answer, node by node, "can I take this?". `to_edge_transform_and_lower(ep, partitioner=[XnnpackPartitioner()])` runs that pass, and the outcome is a graph split between two kinds of execution. ## What happens to claimed subgraphs A claimed group of nodes is removed from the main graph and handed to the backend's ahead-of-time `preprocess` step. XNNPACK, for example, serializes those operators into its own representation; a Core ML lowering produces a Core ML model; a Vulkan lowering produces shader-oriented artifacts. The output is an **opaque blob** — the ExecuTorch toolchain does not interpret it — stored inside the `.pte` alongside the program. In the graph, the whole subgraph collapses into a single `executorch_call_delegate` node that references the blob by id. At run time, the runtime hits that node, looks up the registered backend by name, and hands it the blob plus the input tensors. If the backend is not compiled into the app binary, this fails at load time — which is the classic "works in Python, fails in the app" report: the artifact was lowered to XNNPACK or Core ML, but the mobile build did not include that backend. ## What happens to everything else Declined nodes are not an error. They stay in the graph as ordinary Edge operators, and the runtime executes them with its own kernel library — the portable kernels (small, reference implementations covering broad operator surface) or the optimized ones where available. That fallback is why lowering usually *succeeds* even for models full of unusual operators, and it is also why success tells you nothing about performance. ## Partial delegation is the normal case The important production metric is not "did it lower" but **how much of the compute landed on the backend, and in how many pieces**. Two failure shapes come up: - **A rejected hot operator.** One unsupported activation or normalization in the middle of a residual block can split what should have been one delegated region into two, and the tensors crossing each boundary may need conversion between the backend's preferred layout and the runtime's. The model is correct and disappointingly slow. - **Fragmentation.** Dozens of tiny delegated islands each pay call overhead. Sometimes fewer, larger delegated regions beat more total delegated operators. You diagnose this by printing the lowered graph and counting `executorch_call_delegate` nodes against the operators left behind, and by profiling on the device rather than reasoning about it. ## Multiple partitioners `to_edge_transform_and_lower` accepts a **list** of partitioners. They are consulted in order, so an earlier partitioner claims what it can and a later one sees the remainder. That is how you express "prefer the NPU-backed path, fall back to XNNPACK on CPU". It also means partitioner order is a tuning knob: a greedy first partitioner that takes a subgraph it runs poorly denies it to a backend that would have run it well. The practical constraint on iOS is that Core ML delegation and XNNPACK delegation have different device coverage — the same `.pte` behaves differently across hardware generations — so teams often build per-platform artifacts rather than one universal file. ## Why the old two-step made this worse The deprecated `to_edge()` then `to_backend(partitioner)` flow ran the edge transformation passes first and partitioning afterwards, so a partitioner could be shown a graph already rewritten in ways its pattern matching did not recognise — silently reducing what it claimed. `to_edge_transform_and_lower` interleaves the two, which is a correctness-and-coverage reason to migrate, not just a warning to silence. ## What a strong answer includes Name the mechanism (tag, compile to blob, replace with a call-delegate node), state that the remainder falls back to runtime kernels rather than failing, and then move to the operational consequence: measure delegated coverage and fragmentation, expect the app build to contain every backend you lowered to, and treat partitioner choice and order as a tuning decision made against on-device measurements.

  • What happens on device if the app binary lacks the backend a .pte was lowered to?
    Loading fails, because the call-delegate node names a backend that is not registered with the runtime. This is a frequent cause of a model that works through the Python bindings on a laptop but refuses to load in the app: the lowering step and the mobile build configuration have drifted apart, and both must name the same backend.
  • Why can adding more delegated operators make a model slower?
    Because delegation is measured in regions, not operators. Many small delegated islands separated by fallback operators each pay a call boundary and possibly a layout or memory-format conversion. Fewer, larger contiguous regions often beat higher total coverage, which is why you profile on the device instead of maximizing the count of claimed nodes.
  • What does passing a list of partitioners to to_edge_transform_and_lower express?
    Priority. They are consulted in order, so the first claims what it can and later ones see only the remainder — the way you say "prefer the accelerator, fall back to CPU". Order becomes a tuning knob, since a greedy first partitioner can take a subgraph it executes poorly and deny it to a backend that would have been faster.
  • How do you tell how much of a model actually got delegated?
    Inspect the lowered exported program's graph and count call-delegate nodes against the operators still present, which shows both coverage and fragmentation. Then confirm with an on-device profile, because node counts do not weight by cost — one undelegated convolution can outweigh dozens of delegated elementwise operators.

saying these in an interview costs you the question

  • Assumes lowering delegates the entire model or fails
  • Thinks an unsupported operator aborts the export
  • Treats successful lowering as evidence of good performance
  • Forgets the backend must be compiled into the app binary
  • Ignores boundary overhead between delegated and fallback regions

context