skip to content

Why does an exported ExecuTorch model reject a different input size on device?

level: seniorimportance: should knowfreq 48%

answer

  1. example inputs become recorded guards
  2. static shapes are what memory planning needs
  3. nothing recompiles on the phone
  4. dynamic_shapes plus a named Dim with a range
  5. delegates may decline symbolic dimensions

basics

~20 s

torch.export traces with example inputs and specializes every dimension you did not declare dynamic, recording the concrete sizes as guards. Anything else violates a guard at run time. Declare flexible dimensions with torch.export.Dim through the dynamic_shapes argument.

solid answer

~50 s

Export is a tracing capture: it records the shapes of the example inputs as constraints, and by default every dimension is frozen to the value it had at trace time. Memory planning in `to_executorch()` then sizes the intermediate-tensor arena against those exact shapes, which is part of why the runtime allocates so little. A different input size therefore fails a guard rather than reshaping. To keep a dimension flexible you pass `dynamic_shapes` to `torch.export.export`, marking that axis with a `torch.export.Dim` and giving it a bounded range — bounds matter, because the plan needs a worst case to size against. Two caveats: a dynamic dimension can force the model down paths that reason about symbolic sizes, and backend delegates vary in dynamic-shape support, so a partitioner may decline subgraphs it would otherwise have taken, leaving them on the runtime's own kernels.

code

python · 20 lines
python
import torch
from torch.export import Dim, export

class Encoder(torch.nn.Module):
    def __init__(self):
        super().__init__()
        self.proj = torch.nn.Linear(16, 16)

    def forward(self, tokens):
        return self.proj(tokens).sum(dim=1)

model = Encoder().eval()
seq = Dim("seq", min=1, max=512)

ep = export(
    model,
    (torch.randn(1, 128, 16),),
    dynamic_shapes={"tokens": {1: seq}},
)
print(ep.range_constraints)

go deeper

for a junior

Know that the example input you export with fixes the model's accepted input size, so a phone camera frame usually has to be resized to that exact geometry before it is passed in.

for a middle

Explain that export records shapes as guards and that torch.export.Dim through dynamic_shapes is how a dimension stays flexible, including why the Dim needs a range.

for a senior

Show the diagnosis and the trade: reproduce with the .pte in Python at the device's real shape, then choose between host-side normalization, a bounded dynamic dimension, or shape buckets, and verify delegated coverage did not collapse afterwards.

for a principal

Own the contract between the app and the artifact — which shapes are supported, where preprocessing lives, and how a bounded worst case protects the memory budget across the whole fleet of devices you support.

## Specialization is the default, not an accident `torch.export.export(model, (torch.randn(1, 3, 224, 224),))` does not merely note that the input is a 4-D float tensor. It records `1, 3, 224, 224` as concrete facts and lets every downstream shape computation constant-fold against them. Reshapes become literal target shapes, broadcasts become fixed, and the graph is valid only for that geometry. The exporter attaches guards expressing those assumptions, and the artifact enforces them. This is deliberate. Static shapes are what allow `to_executorch()` to perform **memory planning** — computing ahead of time exactly how large the intermediate-tensor arena must be and where each temporary lives inside it. That is the reason the device runtime can execute a model with almost no dynamic allocation, which is exactly what you want in an app with a tight memory budget and a hostile OS-level memory killer. ## What the failure looks like On device you get a load-or-execute failure when the input tensor you hand `forward()` does not match the recorded signature. It is not a helpful "expected 224 got 256" from the model author's perspective; it is a contract violation at the runtime boundary. Candidates who have not shipped this often assume the runtime will resize or that the delegate will recompile — nothing recompiles on the phone, because there is no compiler there. ## Declaring a dimension dynamic ``` from torch.export import Dim, export batch = Dim("batch", min=1, max=8) ep = export(model.eval(), (x,), dynamic_shapes={"x": {0: batch}}) ``` The `dynamic_shapes` argument maps each input to the axes that are allowed to vary, and each such axis carries a `Dim` with a name and a range. Naming matters: reusing the same `Dim` object across two inputs asserts that those axes are *equal*, which is how you express "sequence length of the ids matches sequence length of the mask". Bounds matter for a practical reason. The memory plan has to size the arena for the worst case, so an unbounded dimension is either rejected or forces a conservative plan. On a device you usually *want* a tight upper bound — it converts a potential out-of-memory crash in the field into an export-time decision. More recent torch versions also let you write `Dim.AUTO` to have the exporter infer which dimensions must be dynamic rather than asserting them yourself; it is convenient for exploration, but for a shipped artifact an explicit named `Dim` with an explicit range documents the contract. ## Where dynamic shapes get expensive Three things go wrong once a dimension is symbolic: - **Guard failures during export.** A model that internally does `x.view(x.shape[0] * 4, -1)` or indexes by a computed size may need extra constraints, and export will tell you that a guard was added on a dynamic dimension. Each such message is a place the model reasons about size in a way you must either bound or rewrite. - **Reduced delegation.** Backends differ in dynamic-shape support. A partitioner that cannot handle a symbolic dimension in a given operator simply declines that subgraph, which falls back to the runtime's portable kernels. The model still runs; it just runs slower than the static version, and the regression is invisible unless you measure delegated coverage. - **Looser memory planning.** Planning against the maximum means holding the worst-case arena even for small inputs. ## The alternative: normalize on the host For a lot of mobile work, the better answer is not to make the model dynamic at all. Vision pipelines resize and letterbox the camera frame to a fixed network input anyway; doing that in the app with the platform's image APIs keeps the exported graph fully static and fully delegated. Reserve dynamic dimensions for cases where padding genuinely costs you — variable-length sequence models being the obvious one, where padding every request to the maximum length wastes real compute. A middle option is **bucketing**: export a small set of shapes (say 128, 256 and 512 tokens) as separate methods or files and dispatch to the nearest bucket at run time. You keep static-shape performance and bound the padding waste, at the cost of artifact size — which is a size-versus-latency trade you should be explicit about. ## How to answer the diagnosis question "The model works in Python but rejects our camera frames" is nearly always this: exported at the notebook's example resolution, fed the device's native resolution. Confirm by inspecting the exported program's input signature, reproduce in Python by loading the `.pte` through the runtime bindings with the device's actual shape, then choose deliberately — normalize on the host, declare a bounded dynamic dimension, or ship buckets — and re-check how much of the graph still lands on the delegate afterwards.

  • Why does a dynamic dimension usually need an upper bound?
    Because to_executorch() plans memory ahead of time and must size the intermediate arena for the worst case. An unbounded dimension leaves no worst case to plan against, and on a phone that is exactly the situation you want to avoid: an unbounded input becomes an out-of-memory kill in the field instead of a decision you made at export time.
  • What happens to backend delegation when you mark a dimension dynamic?
    Delegate support for symbolic shapes is uneven, so a partitioner may decline subgraphs it would have claimed under static shapes. Those operators then execute on the runtime's own kernels. The model still produces correct results, but latency can regress sharply, so measure how much of the graph is still delegated after making a dimension dynamic.
  • When is shape bucketing better than a single dynamic export?
    When padding waste is real but delegation matters more than artifact size. Exporting a handful of fixed sizes and dispatching to the nearest bucket keeps every graph static and fully delegated, bounds the padding overhead, and costs you extra megabytes plus the dispatch logic in the app. It suits sequence models with a known length distribution.
  • Does reusing the same Dim object across two inputs mean anything?
    Yes — it asserts those axes are equal. That is how you express constraints like token ids and attention mask sharing one sequence length. Using two separately named Dim objects would let them vary independently, and export will then require any code that assumes they match to be reconciled or will add a guard you must satisfy.

saying these in an interview costs you the question

  • Assumes the runtime resizes or recompiles for a new shape
  • Thinks any tensor of the right rank and dtype is accepted
  • Marks every dimension dynamic without measuring the delegation loss
  • Declares an unbounded dynamic dimension on a memory-constrained device
  • Debugs on device instead of reproducing with the .pte in Python

context