In Triton, what does raising instance_group count buy, and what does it cost?
answer
- instances are copies, not queue slots
- concurrency knob, not work-per-execution knob
- N copies of weights in VRAM
- helps most when GPU sits idle
- KIND_GPU, KIND_CPU, KIND_AUTO, KIND_MODEL
basics
~20 sEach instance is a separate loaded copy of the model that can execute concurrently, so raising count lets Triton run several executions at once and overlap compute with input/output copies. The cost is one full set of weights per instance plus contention for the same GPU.
solid answer
~50 s`instance_group` in a model's `config.pbtxt` controls how many execution instances Triton creates and where. `count` sets how many, `kind` chooses `KIND_GPU`, `KIND_CPU`, `KIND_AUTO` or `KIND_MODEL`, and `gpus` pins GPU instances to specific device IDs. Each instance is a real copy of the model loaded on that device with its own execution context, and the scheduler dispatches a formed batch to whichever instance is free — so instances give you concurrency, while batching gives you work per execution. Extra instances help when a single execution leaves the GPU idle: CPU-heavy pre/post-processing, small models with launch overhead, or time spent copying inputs and outputs. They help far less on a model already saturating the SMs, where you mostly add VRAM cost and kernel contention. Budget the memory: N instances means roughly N times the weights, plus N sets of activations.
code
protobuf · 12 linesinstance_group [
{
count: 2
kind: KIND_GPU
gpus: [ 0 ]
},
{
count: 1
kind: KIND_GPU
gpus: [ 1 ]
}
]go deeper
Know that instance_group count creates several copies of the model that can run at the same time, and that each copy uses its own GPU memory.
Distinguish concurrency from batch size cleanly: instances add simultaneous executions, batching adds work per execution. Name the kinds and say that memory scales with count.
Reason about where the idle time actually is — copies, CPU pre-processing, kernel launch — and show that you sweep instance count against p95 latency rather than picking a number from the docs.
Own the placement policy: which models get instances versus whole replicas, how the rate limiter protects a latency-critical model sharing a card, and how instance counts interact with your VRAM budget across the fleet.
## What an instance actually is When Triton loads a model it creates one or more **execution instances**. An instance is a fully independent copy of the model, loaded onto a device, with its own execution context and its own backend-side thread. Two instances of the same model can be executing two different batches at the same moment. The default, if you write no `instance_group` at all, is one instance on each visible GPU (or one on CPU for CPU-only backends). You override it with a block like: ``` instance_group [ { count: 2, kind: KIND_GPU, gpus: [ 0 ] } ] ``` You can list several groups to place instances differently across devices — for example two on GPU 0 and one on GPU 1, or a GPU group plus a CPU group for the same model. ## The kinds - `KIND_GPU` — instances on the listed GPUs (all visible GPUs if `gpus` is omitted). - `KIND_CPU` — instances that run on CPU. - `KIND_AUTO` — Triton chooses GPU when the backend supports it, otherwise CPU. This is the default kind. - `KIND_MODEL` — hand device placement to the model/backend itself. This is what you use when the model already spans devices internally, such as a multi-GPU sharded engine that manages its own placement. ## Concurrency versus batch size This is the distinction interviews are testing. Batching increases the *work per execution*; instances increase the *number of simultaneous executions*. They solve different idleness problems. An execution is not pure GPU math. There is input copy (host to device), the kernels, output copy back, and backend-side bookkeeping. While instance A is copying, instance B's kernels can be running. On a small model, kernel launch overhead and copies can be a large share of the wall clock, so a second instance is nearly free throughput. On a large TensorRT engine already at high SM occupancy, a second instance mostly interleaves kernels that were already keeping the GPU busy, and you pay in latency variance. A useful rule: if a single execution leaves the GPU underutilised, add instances; if a single execution already saturates it, add batch size instead — or add a whole replica on another GPU. ## The costs 1. **Memory.** Each instance loads its own weights and its own activation workspace. Two instances of a 6 GB model is 12 GB before you count activations or, for a compiled backend, per-context scratch. This is the constraint that bites first, especially when several models share a card. 2. **Contention.** Concurrent instances compete for SMs, memory bandwidth and PCIe. Aggregate throughput can rise while p99 latency for individual requests gets worse, because any given execution now shares the machine. 3. **Load time and cold start.** More instances means more to initialise on model load, which lengthens the window before the model is ready. ## Instances in a pipeline Instance counts are set per model, which is what makes ensembles powerful: a CPU-bound Python pre-processing step can run with `count: 8` on `KIND_CPU` while the GPU model behind it runs `count: 1` on `KIND_GPU`. Tuning the slow stage independently is usually a bigger win than tuning the whole pipeline as one unit. If you cannot keep the GPU model fed, the pre-processing instance count is often the reason. For the Python backend specifically, each instance is a separate process, so instance count is your only real parallelism knob there — one Python model instance handles its `execute()` calls serially. ## Rate limiting When instances of different models compete for one GPU, Triton's rate limiter lets you declare named resources and per-instance costs inside `instance_group` (`rate_limiter { resources [...] priority: N }`), so the server throttles how many instances may execute at once. Reach for this when a low-priority model's instances are starving a latency-critical one on the same card. ## How to actually pick a number Do not reason about it from the config file. Sweep it: run a load generator at a fixed concurrency against count 1, 2 and 4, and watch throughput against p95 latency and GPU utilisation. Throughput usually climbs then flattens; the point where it flattens while latency keeps climbing is one past your answer. Model Analyzer automates exactly this sweep, jointly over instance count and batching settings.
- Where would you set a different instance count for pre-processing than for the model itself?In an ensemble, each step is its own model with its own config.pbtxt, so you give the CPU pre-processing model a high count on KIND_CPU and the GPU model a low count on KIND_GPU. That is often the cheapest fix for a GPU that sits idle waiting on decode or tokenisation, because you scale only the stage that is actually the bottleneck.
- What does kind: KIND_MODEL mean, and when do you need it?KIND_MODEL tells Triton not to assign devices itself and lets the model or backend decide its own placement. You use it when the model already spans multiple GPUs internally — a sharded or multi-device engine — because Triton pinning an instance to one device would fight the model's own placement logic.
- Two instances doubled throughput on your Python backend model but not on your TensorRT model. Why?The Python model was spending most of its wall clock on CPU work and serialised inference calls, so a second process genuinely ran in parallel. The TensorRT model was already keeping the GPU's SMs busy, so a second execution context only time-slices the same hardware. Concurrency helps where there is idle capacity; it cannot manufacture more compute.
saying these in an interview costs you the question
- Thinking instances share one copy of the weights
- Confusing instance count with max_batch_size
- Assuming throughput scales linearly with instance count
- Adding instances to a GPU already at full utilisation
- Believing instances mean separate server processes or ports