How do you choose between ONNX Runtime, Triton, and ExecuTorch for a PyTorch model?
answer
- target hardware picks the runtime
- count the code outside forward
- every export adds a second stack
- export risk belongs in week one
- one named option is retired
basics
~20 sLet the target decide. On-device means ExecuTorch; a heterogeneous CPU or accelerator fleet with simple pre-processing favours ONNX Runtime; a GPU service with Python-shaped pre- and post-processing favours a Triton PyTorch or Python backend. TorchServe is retired and should not anchor a new design.
solid answer
~50 sThree questions settle it. **Where does the model run?** A phone or embedded board means an ExecuTorch build — nothing else in this list targets it. A GPU server, a CPU fleet, and a browser have different best answers. **How much of the model is not the model?** Tokenization, resampling, non-maximum suppression and business rules are Python-shaped; forcing them into an exported graph is where projects lose months, so a Triton PyTorch or Python backend that keeps the surrounding code in Python is often cheaper than a heroic ONNX export. **How many numerical stacks can the team own?** The moment inference runs somewhere other than PyTorch you have a second implementation to validate every release — parity tests, tolerance budgets, an on-call engineer who understands both. That cost is fixed and recurring, and it is what most candidates omit. Then note the currency fact: **TorchServe is retired** — its last release was 0.12.0 on 2024-09-30 — so "we'll use TorchServe" signals a stale mental model.
go deeper
Know the names and what each targets: ONNX Runtime as a portable cross-platform runtime, Triton as a general inference server that can host a PyTorch model, ExecuTorch as the on-device path. Know TorchServe is retired.
Explain what artifact each consumes and where export breaks — unsupported operators, traced control flow, specialized shapes — and why pre- and post-processing code often decides the choice more than the model does.
Argue the tradeoff with production evidence: parity testing obligations, provider and hardware coverage, how you would de-risk export early, and what you would measure before moving a live service onto a second runtime.
Own it as a long-lived commitment — who maintains the second numerical stack, what the recurring release tax is, which small set of runtimes the organization will support, and the named conditions under which you would reverse the decision.
## Frame it as an artifact-and-owner decision The question is not which runtime is fastest. It is which artifact you commit to producing, who maintains the second numerical stack it creates, and what happens to that commitment when the model changes. Answer in that order and the runtime falls out. ## Where the model runs - **On device — phones, embedded, wearables.** ExecuTorch is PyTorch's on-device path: you capture with `torch.export`, lower to a `.pte`, and run it through a small C++ runtime with hardware delegates. Nothing else here targets that environment, so the decision is made for you; the remaining work is operator coverage and binary size. - **A CPU fleet or mixed accelerators, thin surrounding code.** ONNX Runtime is strong here. One artifact runs across CPU, several GPU vendors, and other targets through its execution providers, and the format is stable and inspectable. The price is the export: your model must survive the translation, and you own the parity testing. - **A GPU service with substantial pre- and post-processing.** Triton Inference Server's PyTorch and Python backends let you keep that code where it already works, at the cost of shipping a PyTorch runtime with it. When pre-processing is the complicated part — and in vision and speech pipelines it usually is — this is the pragmatic choice. ## How much of the pipeline is not the model Export paths handle the tensor computation. They handle tokenizers, image decoding, resampling, tracking state, calibration tables and business rules badly or not at all. Teams underestimate this every time. Count what fraction of your latency budget and, more importantly, your *code* lives outside `forward`. If it is large, an export-everything strategy means reimplementing that code in another language and then keeping two implementations agreeing forever. Keeping it in Python behind a PyTorch-native backend is often the lower-total-cost answer even though it looks less clean on a diagram. ## The second numerical stack This is the argument a principal is expected to make and most candidates never do. Every non-PyTorch runtime introduces a second implementation of your model's arithmetic. From then on, each release needs an export step, a parity gate over a fixture matrix, and a tolerance policy — plus someone who can debug the runtime when it and PyTorch disagree at 2 a.m. That is a permanent tax measured in engineer-time. It is worth paying when it buys something concrete: a hardware target you cannot otherwise reach, a language you must integrate with, a serious cost reduction on a large fleet. It is not worth paying to feel more production-ready. ## Export risk should be discovered early Whichever path you pick, the failure mode is always the same shape: an operator that will not export, control flow that must be rewritten, dynamic shapes that specialize. Discover it in week one, not at the end, by exporting a skeleton of the real architecture before committing. If a model family is under active research, its ability to export is a moving target, and betting a serving design on it is how deadlines die. ## The retired option TorchServe was the obvious answer for years and no longer is: last release 0.12.0 on 2024-09-30, no longer maintained. Bringing it up as a live choice tells an interviewer your knowledge is a few years stale. The right sentence is that the PyTorch-native serving story now routes through generic servers with PyTorch backends, ONNX Runtime, or dedicated engines for the model class in question. ## An answer that lands Say what you would do and why it might be wrong. For example: keep the model in a PyTorch backend behind a general-purpose server while the architecture is still moving and the pre-processing is heavy; export to ONNX Runtime once the architecture stabilizes and the fleet is large enough that the second stack pays for itself; go to ExecuTorch only when the target is a device. Add the checkpoints that would make you revisit — a hardware change, an export blocker in a new architecture, a parity failure rate that shows the tolerance budget was wishful. Naming your own reversal conditions is what distinguishes this level.
- What would make you keep serving a model directly from PyTorch instead of exporting at all?An architecture still changing week to week, heavy Python-side pre- and post-processing, a fleet small enough that the hardware saving is less than an engineer's time, or a model family with known export blockers. Under any of those, the export path costs more than it returns. Revisit when the architecture freezes, the fleet grows, or a hardware target appears that PyTorch cannot reach.
- A team proposes standardizing every model on one runtime across the company. What is your response?Support the intent, resist the absolute. One runtime buys shared tooling, one parity harness and transferable on-call skills, which is real. But device targets and server targets have genuinely different constraints, and forcing a phone build and a GPU service through the same artifact means the harder target dictates the model code for everyone. Standardize the export discipline — capture API, parity gate, fixture matrix — and allow a small, named set of runtimes.
- How do you budget the ongoing cost of a non-PyTorch serving runtime?Price it as recurring, not one-off. Each model release costs an export run, a parity gate over shapes and input classes, and an investigation whenever tolerance is exceeded; each runtime upgrade costs a full revalidation. Add on-call coverage for a stack most of the team cannot debug. Put those hours next to the hardware saving the runtime buys. If the saving does not clearly exceed them, the export is being done for aesthetics.
saying these in an interview costs you the question
- TorchServe is the standard PyTorch serving answer
- ONNX export always works, it's only a format change
- ExecuTorch is just a smaller ONNX Runtime
- One runtime should cover phones and datacenter GPUs alike
- Export cost is a one-time task, not an ongoing obligation