How do you load and run an ExecuTorch .pte model in a mobile app?
answer
- one small Module object per platform
- copy the asset to a real file path
- tensor plus shape, wrapped in a value type
- first call pays load and arena allocation
- preprocessing is the app's job, not the graph's
basics
~20 sShip the .pte to a readable file path, open it with the ExecuTorch Module API for your platform, wrap input tensors in the runtime's value type, call forward, and unwrap the returned tensor. No Python and no PyTorch training stack are involved.
solid answer
~40 sEvery platform exposes the same abstraction, a `Module` wrapping one loaded program. On Android you use the `org.pytorch.executorch` package: `Module.load(path)`, build inputs with `Tensor.fromBlob(data, shape)`, wrap each in an `EValue`, call `module.forward(...)`, and read the result back through the returned `EValue` array. On iOS the ExecuTorch Swift and Objective-C packages expose an equivalent `Module`, and the shared C++ layer is `executorch::extension::Module` with `load()` and `forward()`. Two practical details bite people: the `.pte` usually ships as an app asset and must be copied out to a real filesystem path before loading, because the loader wants a file it can memory-map; and the first call on a method is where the program is loaded and its planned memory arena allocated, so it is measurably slower than steady state. Warm the model off the UI thread.
code
java · 10 linesimport org.pytorch.executorch.EValue;
import org.pytorch.executorch.Module;
import org.pytorch.executorch.Tensor;
public float[] classify(String ptePath, float[] pixels) {
Module module = Module.load(ptePath);
Tensor input = Tensor.fromBlob(pixels, new long[] {1, 3, 224, 224});
EValue[] outputs = module.forward(EValue.from(input));
return outputs[0].toTensor().getDataAsFloatArray();
}go deeper
Be able to describe the four steps — get the file onto a real path, load it into a Module, wrap your input tensor with its shape, call forward and unwrap the output — and say that no Python is involved.
Explain what the value wrapper and explicit shape are for, why the first call is slow, and why preprocessing such as resizing and normalization must match exactly what the model was exported with.
Show you can split a device-side bug quickly: reproduce with the same .pte through the Python runtime bindings, then decide whether the artifact, the backend availability in the app build, or the app's tensor construction is at fault.
Own how models reach devices at all — bundled versus downloaded artifacts, versioning the .pte against app releases, warm-up and threading policy, and the rollback path when a new artifact regresses on a subset of hardware.
## One abstraction, three language surfaces ExecuTorch's runtime API is deliberately small. The central object is a **Module**: it owns a loaded program, resolves methods by name (usually `forward`), and executes them. The C++ type is `executorch::extension::Module`; Android exposes it as `org.pytorch.executorch.Module`; iOS exposes it through the ExecuTorch Swift/Objective-C package. Java and Swift call the same underlying implementation, so mental models transfer between platforms. ## Getting the file onto the device A `.pte` is a plain file. On Android it is normally bundled in `assets/` and copied once at first launch into the app's private files directory; on iOS it goes into the app bundle and you resolve its path from the bundle. The reason for the copy on Android is mundane but catches everyone: assets inside an APK are not regular files with a filesystem path, and the loader wants a path it can open and memory-map. If your artifact is large, this also matters for download-on-demand strategies — you keep the app small and fetch the `.pte` after install, which works precisely because the runtime just needs a path. ## The Android shape of the call ``` Module module = Module.load(filePath); Tensor input = Tensor.fromBlob(floatData, new long[]{1, 3, 224, 224}); EValue[] out = module.forward(EValue.from(input)); float[] scores = out[0].toTensor().getDataAsFloatArray(); ``` `EValue` is the runtime's tagged value type — the boxed thing a method's inputs and outputs are made of, capable of holding a tensor, a scalar, or a list. `Tensor.fromBlob` wraps a primitive array plus an explicit shape; note that you supply the shape, because a flat array has none. Get that shape wrong and you will not get an exception from the array, you will get wrong numbers or a shape-guard failure at execute time. The dtype must match what the model was exported with — a float model fed integer pixel data straight from a bitmap produces garbage. Normalization and layout conversion (bitmap to planar float, mean/std scaling) are your responsibility on the app side; the exported graph starts where your `forward()` started. ## First-call cost and threading Method loading is lazy. Constructing the `Module` is cheap; the first `forward()` resolves the method, materializes the planned memory arena and initializes any backend delegate. That first invocation can be an order of magnitude slower than the ones after it. Two consequences for real apps: warm the model during a screen transition or on a background thread at startup, and never measure latency from a single call. Treat a `Module` as not thread-safe. Use one per worker thread, or serialize calls behind a single-threaded executor. Sharing one instance across concurrent requests is a source of intermittent corruption that is miserable to debug on a device. ## Multiple methods A program can export more than `forward`. The Module API can execute a method by name, which is how models with separate entry points — an encoder and a decoder, a reset alongside a step — ship as one file. If you need this, name your methods at export time and load them explicitly rather than overloading `forward` with a mode flag argument (a mode flag is also exactly the kind of data-dependent branch that complicates export). ## Failure modes worth recognizing - **Load fails immediately.** Usually a path problem (asset not copied) or a missing backend: the `.pte` contains a delegate call for a backend that was not compiled into the app binary. - **Execute fails.** Almost always a shape or dtype mismatch against the exported signature. - **Runs but returns nonsense.** Preprocessing mismatch — wrong channel order, unnormalized inputs, or a model exported in training mode. ## Verifying without a device in the loop Before blaming the app, load the same `.pte` in Python through the runtime bindings and run the exact tensor the app would send. If Python agrees with the app's wrong answer, the artifact is wrong; if Python is right, the bug is in your preprocessing or tensor construction on the mobile side. That split is the fastest way to cut a device-side investigation in half.
- Why is the first forward() call so much slower than later ones?Constructing the Module is cheap, but method loading is lazy: the first call resolves the method, allocates the ahead-of-time planned memory arena and initializes any backend delegate. Warm the model on a background thread before the user needs it, and never benchmark from a single invocation — report a distribution over many warm calls instead.
- The .pte loads in Python but fails to load in the Android app. What do you check first?Two things. Whether the file was actually copied out of assets to a real filesystem path the loader can open, and whether the app binary includes the backend the model was lowered to. A delegate call naming XNNPACK or Core ML fails at load if that backend was not compiled into the mobile build, even though the desktop bindings had it.
- How do you ship a model with more than one entry point?Export the methods under distinct names so the program carries each with its own signature and memory plan, then load and execute them by name from the Module API. That is cleaner than adding a mode flag to forward(), which also introduces exactly the kind of data-dependent branching that makes export harder.
- Is a Module instance safe to share across threads?Treat it as not thread-safe. Use one Module per worker thread or serialize calls through a single-threaded executor. Sharing one instance across concurrent inference requests produces intermittent, hard-to-reproduce corruption, and on a phone that surfaces as rare wrong answers or crashes in the field rather than a clean failure in testing.
saying these in an interview costs you the question
- Expects to load the .pte straight from the APK asset stream
- Assumes the runtime normalizes or resizes inputs for you
- Passes raw integer pixel data to a float model
- Benchmarks using the first, cold invocation
- Shares one Module across concurrent threads