skip to content

Model Repository and Backends

You will learn how Triton discovers models — a versioned repository plus a config.pbtxt per model — and how one server hosts ONNX, PyTorch, TensorRT, Python, and vLLM backends side by side. Interviewers ask because this layout is what lets a platform team serve many models without a service per model.

on this pageshow

questions

6

How must a Triton model repository be laid out on disk for a model to load?

level: juniorimportance: must knowfreq 78%

answer

  1. It is a directory convention, not an API
  2. Three levels: root, model, version
  3. Version directory names are integers
  4. The artifact lives inside the version directory
  5. config.pbtxt sits one level above 1/

basics

~20 s

A Triton model repository is a directory of model directories. Each model directory holds an optional config.pbtxt plus one or more numerically named version subdirectories, and the model file itself lives inside a version directory.

solid answer

~50 s

You point the server at a root directory with `tritonserver --model-repository=/models`. Inside it, every subdirectory is a model, named exactly as clients will address it. A model directory contains `config.pbtxt` (the model configuration) and one or more **version directories whose names are positive integers** — `1/`, `2/`. The actual artifact goes inside the version directory, with a name the backend expects: `model.onnx` for ONNX Runtime, `model.plan` for TensorRT, `model.pt` for TorchScript, `model.savedmodel/` for TensorFlow, `model.py` for the Python backend, `model.json` for the vLLM backend. Directories that are not numeric are ignored, so the classic beginner failure — putting `model.onnx` next to `config.pbtxt` with no `1/` — makes Triton report that the model has no valid versions. The repository path may also be an `s3://`, `gs://` or `as://` URI, and `--model-repository` can be passed more than once.

code

bash · 24 lines
bash
mkdir -p models/densenet_onnx/1
cp densenet.onnx models/densenet_onnx/1/model.onnx
cat > models/densenet_onnx/config.pbtxt <<'EOF'
name: "densenet_onnx"
platform: "onnxruntime_onnx"
max_batch_size: 8
input [
  {
    name: "data_0"
    data_type: TYPE_FP32
    dims: [ 3, 224, 224 ]
  }
]
output [
  {
    name: "fc6_1"
    data_type: TYPE_FP32
    dims: [ 1000 ]
  }
]
EOF
docker run --gpus all --rm -p8000:8000 -p8001:8001 -p8002:8002 \
  -v "$PWD/models:/models" nvcr.io/nvidia/tritonserver:26.07-py3 \
  tritonserver --model-repository=/models

go deeper

for a junior

Be able to sketch the three levels from memory — repository root, model directory, numbered version directory — and say that the model file goes in the version directory while config.pbtxt goes above it.

for a middle

Explain which artifact filename each backend expects and why non-numeric version directories are ignored, and name --model-repository plus its s3, gs and as URI support.

for a senior

Show how this layout becomes your deployment mechanism: artifacts built by CI, copied into an object-store repository, version directories as the unit of rollout, and the startup model table as your first diagnostic.

for a principal

Own the convention across teams: who writes config.pbtxt, whether repositories are per-team or shared, how artifacts get promoted into them, and what pulling multi-gigabyte artifacts from object storage does to server start time.

## What the repository is Triton Inference Server does not have a registration API you call to "install" a model. It has a **convention over a filesystem**: you hand the server one or more root directories with `--model-repository=<path>`, and everything it will serve is discovered by walking that tree at startup (and, in poll or explicit mode, later as well). The layout is therefore part of the contract, not a suggestion, and most first-day Triton failures are layout failures rather than model failures. ## The three levels ``` <repository root>/ <model-name>/ config.pbtxt <version>/ <model file> ``` **Level one — the root.** Passed to `--model-repository`. It can be a local path (usually a volume mounted into the container at `/models`) or a cloud URI: `s3://bucket/prefix`, `gs://bucket/prefix`, `as://account/container/prefix`. The flag may be repeated to serve several repositories from one server; with `--model-namespacing=true` two repositories may even contain a model of the same name. **Level two — the model directory.** Its directory name *is* the model name that clients use in `/v2/models/<name>/infer`. If `config.pbtxt` also sets a `name:` field, it must match the directory. Anything else the model needs but does not version — label files listed as `label_filename`, for example — sits here too. **Level three — the version directory.** Its name must be a **positive integer with no leading zeros**: `1`, `2`, `17`. Triton ignores any subdirectory that is not numeric and any that begins with `.`, which is why `v1/`, `latest/` or `20260820/` silently do nothing. A model with no valid version directory fails to load. ## Which file name each backend wants The default artifact name is fixed per backend: - TensorRT (`platform: "tensorrt_plan"`) → `model.plan` - ONNX Runtime (`platform: "onnxruntime_onnx"`) → `model.onnx` - TorchScript (`platform: "pytorch_libtorch"`) → `model.pt` - TensorFlow SavedModel (`platform: "tensorflow_savedmodel"`) → `model.savedmodel/` (a directory) - Python backend (`backend: "python"`) → `model.py` - vLLM backend (`backend: "vllm"`) → `model.json` If your artifact is named something else, set `default_model_filename` in `config.pbtxt` rather than renaming it by hand in a build step. ## Where config.pbtxt goes `config.pbtxt` sits in the **model** directory, one level above the version directories — one configuration for all versions of the model. Putting it inside `1/` is a common mistake and leaves Triton with a model it cannot configure. For several backends (TensorRT, ONNX, TensorFlow SavedModel, OpenVINO, and Python models implementing `auto_complete_config`) Triton can infer the configuration from the artifact's own metadata, so `config.pbtxt` may be omitted entirely; `--disable-auto-complete-config` turns that inference off and requires an explicit file. (`--strict-model-config` is the deprecated spelling of the same idea.) ## Why versioning is in the path Because versions are directories, deployment is a file copy. Adding `2/` next to `1/` is the whole rollout mechanism; which of them the server actually loads is decided by `version_policy` in `config.pbtxt`, and by default only the highest-numbered one is loaded. Clients can pin a version with `/v2/models/<name>/versions/2/infer` or omit it and get the highest loaded version. ## How it fails, and how to check When a model is missing at runtime, read the server's startup table: Triton prints every model it found with `READY`, `UNAVAILABLE` or an error string. Typical causes are a non-numeric version directory, the artifact placed beside `config.pbtxt` instead of inside a version directory, a wrong artifact name for the declared platform, or file permissions inside the container (the mounted volume must be readable by the server user). Programmatically, `GET /v2/models/<name>/ready` returns 200 only when that model is loaded and ready, and `GET /v2/health/ready` covers the server. Starting with `--exit-on-error=false` keeps the server up when one model fails, which is useful while you fix a repository. ## A mental model Think of the repository as a docroot: the server has no database of models, only a directory it re-reads. Everything you can express about a model — its backend, its shapes, which versions live — is either a directory name or a line in `config.pbtxt` inside that tree. That is exactly what makes it easy to ship from object storage and to diff in Git.

  • Your model directory has config.pbtxt and model.onnx but the server says the model has no versions. What is wrong?
    The artifact is at the wrong level. Triton only looks for model files inside numerically named version subdirectories, so `model.onnx` must move to `1/model.onnx` while `config.pbtxt` stays in the model directory. A directory named `v1` or `latest` will not help — non-numeric directories are ignored.
  • Can Triton serve a model repository straight from object storage?
    Yes. `--model-repository` accepts `s3://`, `gs://` and `as://` URIs as well as local paths, and the flag can be repeated to serve several repositories at once. Triton copies the contents into a local temporary directory when it loads a model, so startup cost scales with artifact size — which matters a lot for multi-gigabyte LLM weights.
  • What if the artifact cannot be named model.onnx because your build system emits another name?
    Set `default_model_filename` in `config.pbtxt` to the actual filename. That is preferable to a rename step in your image build, because it keeps the repository a faithful copy of what the training pipeline produced and keeps the mapping visible in the configuration a reviewer reads.

The repository is a docroot rather than a registry: the server serves whatever the directory tree says exists, so deploying a model is copying files into the right place.

saying these in an interview costs you the question

  • Placing the model file next to config.pbtxt with no version directory
  • Naming version directories v1 or latest instead of 1
  • Putting config.pbtxt inside the version directory
  • Assuming a REST call registers a model rather than the directory layout
  • Thinking the model name comes from config.pbtxt rather than the directory

context

open as a page

In a Triton config.pbtxt, what do max_batch_size and dims declare together?

level: middleimportance: must knowfreq 64%

basics

~20 s

When max_batch_size is greater than 0, Triton assumes a variable leading batch dimension the model accepts, and dims lists the shape of a single instance without it. When max_batch_size is 0 the model cannot batch, and dims must be the complete tensor shape.

open as a page

How do you load a new model into a running Triton server without restarting it?

level: middleimportance: must knowfreq 58%

basics

~10 s

Start the server with --model-control-mode=explicit and then call the repository API: POST /v2/repository/models/<name>/load to load and /unload to remove it. The alternative is poll mode, where Triton rescans the repository on a timer.

open as a page

How does Triton decide which backend executes a model in its repository?

level: middleimportance: should knowfreq 54%

basics

~20 s

Triton picks the backend from the model's configuration: either an explicit backend field such as "python" or "vllm", or a platform field such as "onnxruntime_onnx" or "tensorrt_plan". If neither is set, auto-complete infers it from the artifact filename in the version directory.

open as a page

In Triton, how do you roll out model version 2 and roll back without downtime?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Versions are numbered directories under the same model directory, so rollout is adding a 2/ directory and rollback is pointing version_policy back at 1. Clients keep calling the same model name; only the version Triton loads changes.

open as a page

Should a platform team host 40 models in one Triton server or one per model?

level: principalimportance: should knowfreq 36%

basics

~20 s

Co-locate models that are small, share a GPU comfortably and tolerate a shared failure domain; isolate anything that wants a whole GPU or a different container image. Triton's repository and load API already give deploy independence inside one server, so the reason to split is resource and blast radius.

open as a page