In Triton, how do you roll out model version 2 and roll back without downtime?
answer
- Versions are numbered directories, not registry entries
- One config.pbtxt covers every version
- Default loads only the highest version
- specific { versions: [...] } pins what serves
- No weighted traffic split between versions
basics
~20 sVersions are numbered directories under the same model directory, so rollout is adding a 2/ directory and rollback is pointing version_policy back at 1. Clients keep calling the same model name; only the version Triton loads changes.
solid answer
~50 sA model directory holds `1/`, `2/` and so on, and `config.pbtxt` decides which of them are actually loaded through `version_policy`. The default is `latest { num_versions: 1 }`, so only the highest-numbered version is served — copy in `2/`, trigger a reload, and Triton loads 2 and unloads 1 while the model name and the client URL stay identical. To run both during a validation window, set `version_policy: { latest: { num_versions: 2 } }` or `version_policy: { specific: { versions: [1, 2] } }`; clients then choose with `/v2/models/<name>/versions/2/infer` while requests with no version get the highest loaded one. Rollback is the same lever in reverse: `specific { versions: [1] }` and reload, which is far faster than deleting `2/` because nothing has to be copied. Two costs to name: both versions loaded means both versions' weights resident on the GPU at once, and the reload window transiently holds even more.
code
protobuf · 18 linesname: "ranker"
platform: "onnxruntime_onnx"
max_batch_size: 16
version_policy: { specific: { versions: [ 1, 2 ] } }
input [
{
name: "features"
data_type: TYPE_FP32
dims: [ 128 ]
}
]
output [
{
name: "score"
data_type: TYPE_FP32
dims: [ 1 ]
}
]go deeper
Know that a Triton model can hold several numbered version directories and that by default only the highest-numbered one is served, so adding 2/ is how a new model reaches traffic.
Explain the three version_policy forms and how a client pins a version in the request path, and note that all versions of a model share one config.pbtxt.
Walk the full rollout: atomic copy, an acknowledged reload, both versions loaded for a validation window, metrics broken out by version label, and a rollback that is a policy change rather than a file deletion.
Set the policy: what counts as a version versus a new model name, who owns promotion, how canary traffic is split given Triton has no weighting, and what GPU headroom the fleet reserves so reloads never OOM.
## Versions are part of the filesystem contract Triton does not have a separate model registry with a promotion API — versioning is the numbered directory. That is a deliberately small mechanism, and it pays off because it makes every rollout a file operation plus a configuration line, both of which live in version control and object storage rather than in server state. ``` models/ranker/ config.pbtxt 1/model.onnx 2/model.onnx ``` Both versions are the same *model* — same name, same input and output contract, one `config.pbtxt`. That last point is the constraint that shapes everything else: **versions share a configuration**, so a new version that changes tensor names or shapes is not a version, it is a new model. ## version_policy Three forms: - `version_policy: { latest: { num_versions: N } }` — serve the N highest-numbered versions. This is the default with N = 1. - `version_policy: { all: { } }` — serve every version present. - `version_policy: { specific: { versions: [1, 3] } }` — serve exactly these. Only loaded versions can serve traffic; a version directory that exists but is excluded by the policy costs nothing but disk. ## Rolling forward The sequence for a zero-downtime roll-forward under the default policy: 1. Write `2/` into the repository — atomically, staged elsewhere and moved in, so a poller or an operator cannot see a half-copied artifact. 2. Trigger the reload. In explicit control mode that is `POST /v2/repository/models/ranker/load`, which returns success or failure. In poll mode you wait for the timer and read the log. 3. Triton loads version 2, and once it is ready begins routing unversioned requests to it, then unloads version 1 after its in-flight requests drain. Because the reload holds both versions briefly, plan the GPU memory for the sum, not the maximum. On a GPU sized tightly for one large model, a reload is the moment you discover you had no headroom. ## Serving two versions on purpose When you want a shadow or canary window, load both: ``` version_policy: { specific: { versions: [1, 2] } } ``` Now `/v2/models/ranker/versions/1/infer` and `/v2/models/ranker/versions/2/infer` both work, and an unversioned request goes to the highest loaded version. Triton itself does not split traffic by percentage — there is no weighted routing between versions in the model configuration. Splitting is a caller-side or gateway-side decision that pins the version in the URL. Say that plainly in an interview; candidates often assume a canary weight exists. ## Rolling back The fast rollback is a configuration change, not a file operation: set `specific { versions: [1] }` (or `latest { num_versions: 1 }` after removing `2/`) and reload the model. Version 2's artifact stays on disk, so re-rolling forward later costs nothing. Deleting the directory works too but throws away the artifact and, in a polled repository, races with the poll. A rollback that must also revert `config.pbtxt` — a changed batching cap, say — is a reload of the model configuration, which the load endpoint can also carry as an override in its request body. ## What the metrics should show Triton's Prometheus metrics are labelled by model **and version**, so a canary is observable: the success and failure counters and the latency families carry a version label. That is the evidence for a promote-or-roll-back decision, and it is worth checking that your dashboards actually break out the label rather than summing over it — a summed panel will hide a canary that is failing 5% of requests. ## Judgement Use version directories for changes that keep the contract: retrained weights, a requantized artifact, a recompiled engine. Use a new model name for anything that changes inputs, outputs, or the meaning of the output, and let clients migrate deliberately. The failure mode teams hit is treating the version number as a general-purpose change channel and then discovering one `config.pbtxt` cannot describe both versions.
- Can Triton send 10% of traffic to version 2 and 90% to version 1?Not by itself — the model configuration has no traffic weighting between versions. Both versions can be loaded, and a request either names a version in its URL or gets the highest loaded one. Percentage splitting is done by the caller or by a gateway in front of Triton that rewrites the path to pin a version.
- What is the memory cost of a version_policy that loads two versions of a large model?Both versions' weights are resident simultaneously, plus each one's own runtime allocations, so budget the sum. On a tightly sized GPU this is what turns a validation window into an out-of-memory failure — and note that even a single-version reload holds old and new briefly while the new instance comes up.
- Why is bumping the version directory the wrong tool when a retrained model changes its output tensor?Because all versions of a model share one config.pbtxt, which declares the output names, types and shapes. A changed contract cannot be described by one configuration, and clients would break silently on whichever version they reach. Publish it as a new model name and migrate callers explicitly.
saying these in an interview costs you the question
- Expecting Triton to weight traffic across versions
- Thinking each version can have its own config.pbtxt
- Assuming a new version directory serves immediately without a reload
- Forgetting that two loaded versions double resident weights
- Using a version bump for a changed input or output contract