skip to content

How does Triton decide which backend executes a model in its repository?

level: middleimportance: should knowfreq 54%

answer

  1. The config picks it, the image supplies it
  2. backend: versus the older platform: spelling
  3. Auto-complete guesses from the artifact filename
  4. libtriton_<name>.so under /opt/tritonserver/backends
  5. vLLM and TensorRT-LLM ship in separate image variants

basics

~20 s

Triton picks the backend from the model's configuration: either an explicit backend field such as "python" or "vllm", or a platform field such as "onnxruntime_onnx" or "tensorrt_plan". If neither is set, auto-complete infers it from the artifact filename in the version directory.

solid answer

~40 s

Every model resolves to exactly one backend — a shared library the server loads. You select it in `config.pbtxt` with `backend: "python"`, `backend: "vllm"`, `backend: "onnxruntime"`, `backend: "tensorrt"`, or with the older `platform:` spelling for framework models (`onnxruntime_onnx`, `tensorrt_plan`, `pytorch_libtorch`, `tensorflow_savedmodel`). When both are omitted, Triton's auto-complete infers the backend from the artifact it finds — `model.onnx`, `model.plan`, `model.pt`, `model.py` — which is why many framework models need no configuration file at all. The library itself lives at `/opt/tritonserver/backends/<name>/libtriton_<name>.so`, and this is the crucial operational point: **backends ship in the image, not with the model**. The base `nvcr.io/nvidia/tritonserver:26.07-py3` carries TensorRT, ONNX Runtime, PyTorch and Python; the vLLM and TensorRT-LLM backends ship in the separate `26.07-vllm-python-py3` and `26.07-trtllm-python-py3` variants. Asking for a backend the image does not contain fails at load with a missing-backend error.

go deeper

for a junior

Know that Triton runs several frameworks through pluggable backends, and that config.pbtxt names the one a model uses with either a backend or a platform line.

for a middle

Explain the backend versus platform spellings, what auto-complete infers from the artifact filename, and that the backend library lives in the container image rather than in the model repository.

for a senior

Read a load failure correctly — a missing backend library is an image problem, not a config problem — and choose image variants deliberately when co-locating LLM and classic-framework models.

for a principal

Own the image strategy: which backends the platform supports, how many image variants you are willing to operate, and whether an LLM backend's upgrade cadence should be coupled to the rest of the fleet.

## One model, one backend Triton is a thin request-scheduling shell around a set of pluggable **backends**. A backend is a shared library implementing Triton's backend C API: it knows how to load one kind of artifact and how to run inference on it. ONNX Runtime, TensorRT, PyTorch, OpenVINO, Python, vLLM and TensorRT-LLM are all backends. This plug-in design is exactly what lets one server host a TensorRT detector, an ONNX ranker and a Python pre-processor side by side. ## The two fields that choose it In `config.pbtxt` there are two spellings: - `backend: "<name>"` — the general form, and the only form for non-framework backends: `"python"`, `"vllm"`, `"tensorrtllm"`, `"onnxruntime"`, `"pytorch"`, `"tensorrt"`, `"openvino"`. - `platform: "<name>"` — the older framework-oriented form, which encodes both the framework and the artifact format: `"onnxruntime_onnx"`, `"tensorrt_plan"`, `"pytorch_libtorch"`, `"tensorflow_savedmodel"`. `platform: "ensemble"` is a special value that means no backend at all — the model is a graph of other models. For framework models the two are interchangeable in practice; for Python, vLLM and TensorRT-LLM you use `backend`. ## What auto-complete infers If you write neither field, Triton looks at what is inside the version directory and maps the filename to a backend: `model.plan` to TensorRT, `model.onnx` to ONNX Runtime, `model.pt` to PyTorch, `model.savedmodel/` to TensorFlow, `model.py` to Python, `model.json` to vLLM. For most framework artifacts it then reads the artifact's own metadata for input and output names, types and shapes, so the model loads with no `config.pbtxt` at all. `--disable-auto-complete-config` (the current spelling; `--strict-model-config` is deprecated) turns this inference off and makes an explicit configuration mandatory — a reasonable production setting, because it converts a silent wrong guess into a startup error. If your file is not named the default, `default_model_filename` in `config.pbtxt` tells Triton what to look for. ## Where backends physically live At startup Triton scans `/opt/tritonserver/backends/`. Each subdirectory is a backend name and must contain `libtriton_<name>.so`. `--backend-directory` moves that root, and a backend can also be shipped inside a model repository for custom builds. The practical consequence is that **your image determines your backend menu**: - `nvcr.io/nvidia/tritonserver:26.07-py3` — the general image, with TensorRT, ONNX Runtime, PyTorch, Python and OpenVINO. - `nvcr.io/nvidia/tritonserver:26.07-vllm-python-py3` — adds the vLLM backend. - `nvcr.io/nvidia/tritonserver:26.07-trtllm-python-py3` — adds the TensorRT-LLM backend. A `config.pbtxt` that names `vllm` on the base image fails to load with a message about not finding the backend library. There is no download-on-demand. ## The Python backend `backend: "python"` runs your `model.py`, which defines a class `TritonPythonModel` with `initialize`, `execute` and optionally `finalize` and `auto_complete_config`. Inside, `triton_python_backend_utils` (conventionally imported as `pb_utils`) provides `get_input_tensor_by_name`, `Tensor` and `InferenceResponse`. It runs in a separate process from the server core and is how you serve anything the compiled backends do not cover — arbitrary Python libraries, tokenizers, business logic wrapping other models. The cost is Python: a process per model instance, and the overhead of moving tensors across the process boundary. ## The vLLM backend `backend: "vllm"` reads a `model.json` in the version directory containing vLLM engine arguments — `model`, `tensor_parallel_size`, `gpu_memory_utilization` and the rest — and stands up a vLLM engine behind Triton's request interface. It is itself implemented on top of the Python backend. Note the trap when copying arguments from old examples: engine arguments are validated against the vLLM version inside the image, and vLLM 0.27 accepts `enable_log_requests`, not the older inverted `disable_log_requests`. ## Diagnosing it When a model does not load, the startup table names the model and the reason. "Unable to find backend library" means the image, not the config, is wrong. An unknown platform value means a typo in `platform:`. A model that loads but produces wrong outputs, having had no `config.pbtxt`, is usually auto-complete having inferred something you did not intend — write the file and disable auto-complete.

  • What does the Python backend cost you compared with a compiled backend?
    It runs your model in a separate Python process per instance, so every request crosses a process boundary with the tensors, and you inherit Python's execution speed for anything outside vectorised library calls. It buys arbitrary Python — tokenizers, custom logic, calling other models — which no compiled backend gives you. Use it for glue and for models with no compiled path, not as the default.
  • When would you keep config.pbtxt even though auto-complete could generate it?
    Whenever the configuration expresses intent the artifact cannot: batching caps, scheduling, how many instances run, warmup, or an explicit contract you want reviewed in a pull request. Many teams go further and run with --disable-auto-complete-config so a missing or malformed configuration is a startup failure instead of a silent inference from file metadata.
  • Can one Triton server run models on several different backends at once?
    Yes — that is the point of the design. A repository can hold a TensorRT plan, an ONNX model and a Python pre-processor, and the server loads a separate backend library for each. The constraint is that all of those backends must be present in the running image, which is why mixing a vLLM model with classic framework models forces the vllm-python image variant.

saying these in an interview costs you the question

  • Assuming Triton downloads a missing backend at load time
  • Thinking the backend is chosen by the client at request time
  • Writing platform: "vllm" instead of backend: "vllm"
  • Believing the Python backend runs inside the server process
  • Treating the base -py3 image as containing every backend

context