skip to content

How must a Triton model repository be laid out on disk for a model to load?

level: juniorimportance: must knowfreq 78%

answer

  1. It is a directory convention, not an API
  2. Three levels: root, model, version
  3. Version directory names are integers
  4. The artifact lives inside the version directory
  5. config.pbtxt sits one level above 1/

basics

~20 s

A Triton model repository is a directory of model directories. Each model directory holds an optional config.pbtxt plus one or more numerically named version subdirectories, and the model file itself lives inside a version directory.

solid answer

~50 s

You point the server at a root directory with `tritonserver --model-repository=/models`. Inside it, every subdirectory is a model, named exactly as clients will address it. A model directory contains `config.pbtxt` (the model configuration) and one or more **version directories whose names are positive integers** — `1/`, `2/`. The actual artifact goes inside the version directory, with a name the backend expects: `model.onnx` for ONNX Runtime, `model.plan` for TensorRT, `model.pt` for TorchScript, `model.savedmodel/` for TensorFlow, `model.py` for the Python backend, `model.json` for the vLLM backend. Directories that are not numeric are ignored, so the classic beginner failure — putting `model.onnx` next to `config.pbtxt` with no `1/` — makes Triton report that the model has no valid versions. The repository path may also be an `s3://`, `gs://` or `as://` URI, and `--model-repository` can be passed more than once.

code

bash · 24 lines
bash
mkdir -p models/densenet_onnx/1
cp densenet.onnx models/densenet_onnx/1/model.onnx
cat > models/densenet_onnx/config.pbtxt <<'EOF'
name: "densenet_onnx"
platform: "onnxruntime_onnx"
max_batch_size: 8
input [
  {
    name: "data_0"
    data_type: TYPE_FP32
    dims: [ 3, 224, 224 ]
  }
]
output [
  {
    name: "fc6_1"
    data_type: TYPE_FP32
    dims: [ 1000 ]
  }
]
EOF
docker run --gpus all --rm -p8000:8000 -p8001:8001 -p8002:8002 \
  -v "$PWD/models:/models" nvcr.io/nvidia/tritonserver:26.07-py3 \
  tritonserver --model-repository=/models

go deeper

for a junior

Be able to sketch the three levels from memory — repository root, model directory, numbered version directory — and say that the model file goes in the version directory while config.pbtxt goes above it.

for a middle

Explain which artifact filename each backend expects and why non-numeric version directories are ignored, and name --model-repository plus its s3, gs and as URI support.

for a senior

Show how this layout becomes your deployment mechanism: artifacts built by CI, copied into an object-store repository, version directories as the unit of rollout, and the startup model table as your first diagnostic.

for a principal

Own the convention across teams: who writes config.pbtxt, whether repositories are per-team or shared, how artifacts get promoted into them, and what pulling multi-gigabyte artifacts from object storage does to server start time.

## What the repository is Triton Inference Server does not have a registration API you call to "install" a model. It has a **convention over a filesystem**: you hand the server one or more root directories with `--model-repository=<path>`, and everything it will serve is discovered by walking that tree at startup (and, in poll or explicit mode, later as well). The layout is therefore part of the contract, not a suggestion, and most first-day Triton failures are layout failures rather than model failures. ## The three levels ``` <repository root>/ <model-name>/ config.pbtxt <version>/ <model file> ``` **Level one — the root.** Passed to `--model-repository`. It can be a local path (usually a volume mounted into the container at `/models`) or a cloud URI: `s3://bucket/prefix`, `gs://bucket/prefix`, `as://account/container/prefix`. The flag may be repeated to serve several repositories from one server; with `--model-namespacing=true` two repositories may even contain a model of the same name. **Level two — the model directory.** Its directory name *is* the model name that clients use in `/v2/models/<name>/infer`. If `config.pbtxt` also sets a `name:` field, it must match the directory. Anything else the model needs but does not version — label files listed as `label_filename`, for example — sits here too. **Level three — the version directory.** Its name must be a **positive integer with no leading zeros**: `1`, `2`, `17`. Triton ignores any subdirectory that is not numeric and any that begins with `.`, which is why `v1/`, `latest/` or `20260820/` silently do nothing. A model with no valid version directory fails to load. ## Which file name each backend wants The default artifact name is fixed per backend: - TensorRT (`platform: "tensorrt_plan"`) → `model.plan` - ONNX Runtime (`platform: "onnxruntime_onnx"`) → `model.onnx` - TorchScript (`platform: "pytorch_libtorch"`) → `model.pt` - TensorFlow SavedModel (`platform: "tensorflow_savedmodel"`) → `model.savedmodel/` (a directory) - Python backend (`backend: "python"`) → `model.py` - vLLM backend (`backend: "vllm"`) → `model.json` If your artifact is named something else, set `default_model_filename` in `config.pbtxt` rather than renaming it by hand in a build step. ## Where config.pbtxt goes `config.pbtxt` sits in the **model** directory, one level above the version directories — one configuration for all versions of the model. Putting it inside `1/` is a common mistake and leaves Triton with a model it cannot configure. For several backends (TensorRT, ONNX, TensorFlow SavedModel, OpenVINO, and Python models implementing `auto_complete_config`) Triton can infer the configuration from the artifact's own metadata, so `config.pbtxt` may be omitted entirely; `--disable-auto-complete-config` turns that inference off and requires an explicit file. (`--strict-model-config` is the deprecated spelling of the same idea.) ## Why versioning is in the path Because versions are directories, deployment is a file copy. Adding `2/` next to `1/` is the whole rollout mechanism; which of them the server actually loads is decided by `version_policy` in `config.pbtxt`, and by default only the highest-numbered one is loaded. Clients can pin a version with `/v2/models/<name>/versions/2/infer` or omit it and get the highest loaded version. ## How it fails, and how to check When a model is missing at runtime, read the server's startup table: Triton prints every model it found with `READY`, `UNAVAILABLE` or an error string. Typical causes are a non-numeric version directory, the artifact placed beside `config.pbtxt` instead of inside a version directory, a wrong artifact name for the declared platform, or file permissions inside the container (the mounted volume must be readable by the server user). Programmatically, `GET /v2/models/<name>/ready` returns 200 only when that model is loaded and ready, and `GET /v2/health/ready` covers the server. Starting with `--exit-on-error=false` keeps the server up when one model fails, which is useful while you fix a repository. ## A mental model Think of the repository as a docroot: the server has no database of models, only a directory it re-reads. Everything you can express about a model — its backend, its shapes, which versions live — is either a directory name or a line in `config.pbtxt` inside that tree. That is exactly what makes it easy to ship from object storage and to diff in Git.

  • Your model directory has config.pbtxt and model.onnx but the server says the model has no versions. What is wrong?
    The artifact is at the wrong level. Triton only looks for model files inside numerically named version subdirectories, so `model.onnx` must move to `1/model.onnx` while `config.pbtxt` stays in the model directory. A directory named `v1` or `latest` will not help — non-numeric directories are ignored.
  • Can Triton serve a model repository straight from object storage?
    Yes. `--model-repository` accepts `s3://`, `gs://` and `as://` URIs as well as local paths, and the flag can be repeated to serve several repositories at once. Triton copies the contents into a local temporary directory when it loads a model, so startup cost scales with artifact size — which matters a lot for multi-gigabyte LLM weights.
  • What if the artifact cannot be named model.onnx because your build system emits another name?
    Set `default_model_filename` in `config.pbtxt` to the actual filename. That is preferable to a rename step in your image build, because it keeps the repository a faithful copy of what the training pipeline produced and keeps the mapping visible in the configuration a reviewer reads.

The repository is a docroot rather than a registry: the server serves whatever the directory tree says exists, so deploying a model is copying files into the right place.

saying these in an interview costs you the question

  • Placing the model file next to config.pbtxt with no version directory
  • Naming version directories v1 or latest instead of 1
  • Putting config.pbtxt inside the version directory
  • Assuming a REST call registers a model rather than the directory layout
  • Thinking the model name comes from config.pbtxt rather than the directory

context