How do you load a new model into a running Triton server without restarting it?
answer
- Three modes: none, poll, explicit
- Default refuses to change the loaded set
- Explicit hands control to an HTTP call
- Repository endpoints are POST under /v2/repository
- A load call returns whether it worked
basics
~10 sStart the server with --model-control-mode=explicit and then call the repository API: POST /v2/repository/models/<name>/load to load and /unload to remove it. The alternative is poll mode, where Triton rescans the repository on a timer.
solid answer
~50 sTriton's `--model-control-mode` flag has three values. `none` is the default: everything in the repository loads at startup and the set never changes. `poll` makes Triton rescan the repository every `--repository-poll-secs` seconds and pick up added, changed or removed models on its own. `explicit` puts you in charge — the server starts with only the models named by `--load-model` (or none at all), and models are loaded and unloaded on demand through the model-control protocol: `POST /v2/repository/models/<name>/load`, `POST /v2/repository/models/<name>/unload`, and `POST /v2/repository/index` to list what the server knows and its state. From Python, `tritonclient` wraps these as `load_model()`, `unload_model()` and `get_model_repository_index()`. Explicit mode is what you want in a delivery pipeline: your deploy step copies the artifact into the repository and then issues an explicit load, so a failed load is an error you can act on rather than a timer quietly doing something.
code
bash · 3 linestritonserver --model-repository=/models \
--model-control-mode=explicit \
--load-model=densenet_onnxgo deeper
Remember that Triton can load models without a restart, and that this is enabled by a server flag plus HTTP calls under /v2/repository rather than by dropping files in and hoping.
Name the three control modes and what each one does, and describe the load, unload and index endpoints together with the tritonclient methods that wrap them.
Argue for explicit mode in a delivery pipeline because a load call returns success or failure, and account for what a load actually costs — artifact read, GPU allocation, warmup — and for the drain on unload.
Decide the control-plane model for the fleet: who may call load and unload, how that port is protected given the API is unauthenticated, and whether models are preloaded or brought in on demand to fit a GPU budget.
## The problem A Triton server hosting many models cannot be restarted every time one of them changes — a restart reloads every model, and for large artifacts that is minutes of downtime for models that did not change. The model-control machinery exists to make loading a per-model operation. ## The three modes **`--model-control-mode=none`** (the default). Triton loads every model in the repository at startup and then freezes that set. Changes on disk are ignored, and the repository API's load and unload calls are rejected. This is the right mode for an immutable deployment where a new model means a new pod. **`--model-control-mode=poll`**. Triton rescans the repository every `--repository-poll-secs` seconds. A new model directory appears and it is loaded. A new version directory appears and it is picked up according to the model's `version_policy`. A model directory disappears and the model is unloaded. Convenient, and dangerous in exactly the way you would expect: a partially copied artifact can be picked up mid-write, and there is no acknowledgement anywhere that a load succeeded. If you use poll, write into the repository atomically — stage elsewhere, then rename. **`--model-control-mode=explicit`**. Nothing loads at startup except models named with `--load-model=<name>` (repeatable; `--load-model=*` loads everything once). After that, the model-control protocol governs the set. ## The repository API The endpoints are extensions to the KServe v2 protocol and are all POST: - `POST /v2/repository/models/<name>/load` — load, or reload if already loaded - `POST /v2/repository/models/<name>/unload` — unload, freeing its GPU memory - `POST /v2/repository/index` — return every model in the repository with its state (`READY`, `UNAVAILABLE`, `LOADING`, `UNLOADING`) and version A load call blocks until the load finishes and returns an error if it failed, which is what makes it usable as a deployment step. Note that per-model metadata lives at `GET /v2/models/<name>` and readiness at `GET /v2/models/<name>/ready`; the repository index is the endpoint that enumerates the repository. The load request body can also carry an override: a replacement model configuration, and even the model files themselves base64-encoded as `file:1/model.onnx` parameters. That lets a control plane push a model without shared storage, at the cost of pushing the artifact through an HTTP request. ## From a client `tritonclient` 2.71 exposes the same operations over HTTP and gRPC: `InferenceServerClient.load_model(name)`, `unload_model(name)`, `get_model_repository_index()`, plus `is_model_ready(name)` and `is_server_ready()`. Wrapping the deploy in these calls gives you a load with a return value, which a file copy into a polled directory never does. ## Reloading in place Calling load on an already-loaded model reloads it — Triton brings up the new instance and swaps. That is how you apply a changed `config.pbtxt` or a new artifact under the same version without an unload gap. During a reload the model's memory footprint can transiently be that of both copies, which matters when the model is a multi-gigabyte LLM on a nearly full GPU. ## Why explicit mode is the production default for platforms Three reasons. First, **acknowledgement**: your pipeline learns whether the load worked. Second, **startup time**: a server hosting fifty models does not need to load all fifty to become ready; load them on demand or preload a hot subset with `--load-model`. Third, **capacity control**: unloading is how you reclaim GPU memory from a model nobody is calling, which is the whole basis of hosting more models than fit at once. ## The cautions The repository endpoints are unauthenticated like the rest of Triton's API — anyone who can reach the port can unload your production model, so keep the HTTP port on an internal network behind your own gateway. Loading is not free either: it reads the artifact (possibly from object storage), allocates GPU memory and may run warmup, so a load triggered by the first request for a cold model is a real latency event, not a millisecond of overhead. And an unload only completes when in-flight requests for that model drain.
- Why is poll mode risky for a production repository backed by shared storage?Because the timer, not you, decides when a directory is complete. A copy still in progress can be picked up half-written, a failed load is reported only in the server log, and you have no acknowledgement to gate a pipeline on. If you must poll, stage the model elsewhere and move it in atomically, and treat the poll interval as added deploy latency.
- What happens if you call the load endpoint while the server runs in the default control mode?The request is rejected — in none mode the loaded model set is fixed at startup and the model-control protocol is disabled. This is the usual first surprise, and the fix is a server flag rather than a client change, so it requires restarting the server with --model-control-mode=explicit.
- How do you make one hot model available immediately in explicit mode without loading the other forty?Pass --load-model=<name> at startup for the models you want warm; it is repeatable, and --load-model=* loads everything. Everything else stays unloaded until a control-plane load call brings it in, which keeps server startup fast and GPU memory free for the models actually receiving traffic.
saying these in an interview costs you the question
- Expecting load and unload calls to work in the default mode
- Thinking a file copy alone loads a model in explicit mode
- Assuming the repository endpoints are authenticated
- Believing unload is instant regardless of in-flight requests
- Treating a model load as a millisecond-scale operation