Should a platform team host 40 models in one Triton server or one per model?
answer
- Deploy independence is already free
- One process, one failure domain
- No per-model GPU quota exists
- Backends live in the image, not the model
- Anything wanting a whole GPU gets its own server
basics
~20 sCo-locate models that are small, share a GPU comfortably and tolerate a shared failure domain; isolate anything that wants a whole GPU or a different container image. Triton's repository and load API already give deploy independence inside one server, so the reason to split is resource and blast radius.
solid answer
~60 sThe pitch for one server is real: one repository, one GPU shared across many small models, per-model load and unload so you can host more models than fit at once, and one metrics endpoint. Triton already gives you *deployment* independence within a single process — a model is a directory, and explicit model control loads and unloads it without touching its neighbours — so "we need separate servers to deploy separately" is not a reason. The reasons that hold are: **blast radius**, since one process means one GPU out-of-memory or one crashing model can take every co-tenant with it; **resource contention**, since models share GPU memory and compute with no per-model quota; and **the image**, since backends are baked into the container, so a vLLM or TensorRT-LLM model forces a specific image variant that your classic ONNX models do not need. In practice teams land on one shared Triton fleet for the long tail of small models, and a dedicated server per LLM, because an LLM engine wants the whole GPU's memory anyway.
go deeper
Know that one Triton server can host many models from one repository, addressed by name, rather than needing a separate service per model.
Explain what co-tenancy shares — one process, one GPU, one image, one metrics endpoint — and that load and unload make deployment per-model even inside a shared server.
Argue the isolation case concretely: no per-model memory quota, a shared failure domain, and a large LLM that wants the whole GPU, then propose which models you would pull out and why.
Own the tiering and the contract: which models share a fleet, who writes into the shared repository, who holds the GPU budget, and how image variants and upgrade cadence shape the grouping.
## The question behind the question An interviewer asking this wants to hear you separate three things that get conflated: deployment independence, failure isolation, and resource isolation. Triton gives you the first for free and neither of the other two. ## What one server genuinely buys **Utilisation.** Forty small models each needing 2 GB do not each deserve a GPU. One server with one repository lets them share, and explicit model control (`--model-control-mode=explicit` plus the load and unload endpoints) lets you host far more models than fit resident at once — cold ones stay unloaded until called, hot ones stay warm. **A single operational surface.** One deployment, one HTTP and gRPC endpoint, one metrics scrape whose series carry model and version labels, so per-model observability survives co-tenancy. One upgrade to test. For a platform team that is not a small saving. **Heterogeneity.** One repository can hold a TensorRT plan, an ONNX ranker and a Python pre-processor, each running on its own backend inside the same process — which is the original reason Triton exists. **Multiple repositories.** `--model-repository` can be repeated, so per-team repositories can be served by one process; `--model-namespacing=true` even allows two repositories to contain a model with the same name. ## What it costs **Blast radius.** The process is the failure domain. A model that exhausts GPU memory takes down inference for its neighbours; memory pressure and CUDA errors do not stay politely local. Restarting the server to fix one model reloads all of them. **Resource contention with no quota.** Triton has no per-model GPU memory limit and no per-model share of compute. A model that suddenly gets a traffic spike consumes what it consumes, and a co-tenant's latency degrades. You can shape this — capping per-model batch ceilings and how many instances run — but you cannot express "this model may never exceed 20% of the GPU". **Noisy startup.** In the default control mode, one server hosting forty models is not ready until all forty load. That is why explicit mode matters as much for co-tenancy as it does for deploys. **The image constraint.** Backends ship in the container: the base `26.07-py3` image carries TensorRT, ONNX Runtime, PyTorch and Python, while vLLM and TensorRT-LLM live in the `26.07-vllm-python-py3` and `26.07-trtllm-python-py3` variants. Wanting one vLLM model alongside your classic models forces the whole co-tenant group onto that variant, and couples their upgrade cadence to a fast-moving LLM engine's. That coupling is often the deciding argument. ## Where LLMs change the calculus An LLM server is designed to claim nearly all of the GPU's memory for weights and KV cache. Co-tenancy with such a model is not a tuning problem, it is a contradiction: whatever you leave for neighbours you take from the KV cache, which is the thing that determines concurrency. So the honest rule is that a model wanting a whole GPU gets its own server, and Triton's multi-model strength applies to the long tail beneath it. ## A decision rule you can defend Co-locate when: models are small relative to the GPU, their traffic is bursty and uncorrelated (so sharing actually raises utilisation), they belong to one team or one criticality tier, and they run on backends present in a single image. Isolate when: the model wants most of a GPU; its availability tier differs from its neighbours' (a revenue-path model should not share a process with an experiment); it needs a different image or a different upgrade cadence; or a compliance boundary demands separate infrastructure. ## The organisational layer The part a principal owns is not the flag, it is the contract. Who may write into the shared repository? Is it pipeline-only, with a promotion step, or can a data scientist drop a directory in? Who is paged when a co-tenant exhausts the GPU? A shared server without an owner for the GPU budget becomes a tragedy of the commons within a quarter — the technically correct co-location fails for organisational reasons. Write down a per-model memory budget, enforce it at load time in the pipeline, and make the tier boundary explicit: shared servers for the lowest tier, dedicated servers above it.
- A co-tenant model spikes and pushes the GPU into out-of-memory failures. What do you change so it cannot happen again?Short term, unload the offending model, cap what it can allocate — its batch ceiling and how many instances run — then reload. Structurally, move it to its own server if its footprint is inherently large, and put a per-model memory budget into the promotion pipeline so a model that does not fit the shared tier never lands there. Triton has no per-model memory quota to fall back on.
- Does hosting many models in one server help or hurt cold-start latency?It helps if you use explicit control: the server starts fast with a hot subset preloaded and brings cold models in on demand, and a warm neighbour keeps the process and its GPU context alive so a load is artifact-read plus allocation rather than a whole pod start. It hurts under the default control mode, where readiness waits for every model in the repository to load.
- How does the container image influence how you group models across servers?Backends are baked into the image, so a group of co-tenants must share one variant. Putting a vLLM model beside classic ONNX models forces everyone onto the vllm-python image and ties their upgrade cadence to a fast-moving engine. Grouping by required image is often a cleaner first cut than grouping by team.
saying these in an interview costs you the question
- Claiming separate servers are needed to deploy models independently
- Assuming Triton enforces per-model GPU memory limits
- Co-locating a large LLM with small models to raise utilisation
- Ignoring that all co-tenants share one container image and one upgrade
- Treating a shared repository as safe without an ownership contract