skip to content

Batching, Instances and Ensembles

You will learn Triton's schedulers — dynamic batching for stateless models, sequence batching for stateful ones — plus instance groups for concurrency and ensembles that chain preprocessing, model, and postprocessing in-server. Interviewers ask how you would cut a network hop out of a multi-stage pipeline.

on this pageshow

questions

6

What is a Triton ensemble model, and what cost does it remove from a pipeline?

level: juniorimportance: must knowfreq 52%

answer

  1. platform: "ensemble", no backend, no compute
  2. steps wired by tensor names
  3. input_map and output_map, not an ordered list
  4. one client call, no intermediate round-trips
  5. static DAG: no branching, no loops

basics

~20 s

A Triton ensemble is a scheduling-only model, declared with platform: "ensemble", that wires several real models into a pipeline inside the server. The client sends one request instead of three, and intermediate tensors never leave the server, removing network round-trips and serialization between stages.

solid answer

~50 s

An ensemble has no backend and runs no compute of its own. Its `config.pbtxt` sets `platform: "ensemble"` and an `ensemble_scheduling` block listing `step` entries; each step names a model and version, and uses `input_map` and `output_map` to bind that model's tensor names to shared pipeline names. Triton builds a DAG from those names and runs steps as soon as their inputs exist, so independent branches run concurrently. The win is that a typical preprocess → model → postprocess chain becomes one client call: intermediate tensors stay in server memory rather than being serialized back to the client and re-uploaded, which for image or audio payloads dominates end-to-end latency. Each step remains an ordinary model in the repository, so it keeps its own instance count and batching settings and can be tuned independently. The limits: an ensemble is a static graph with no conditionals or loops, and any step's failure fails the whole request.

code

protobuf · 21 lines
protobuf
name: "image_pipeline"
platform: "ensemble"
max_batch_size: 8
input [ { name: "IMAGE" data_type: TYPE_UINT8 dims: [ -1 ] } ]
output [ { name: "CLASS" data_type: TYPE_FP32 dims: [ 1000 ] } ]
ensemble_scheduling {
  step [
    {
      model_name: "preprocess"
      model_version: -1
      input_map { key: "RAW_IMAGE" value: "IMAGE" }
      output_map { key: "PREPROCESSED" value: "preprocessed_image" }
    },
    {
      model_name: "resnet50"
      model_version: -1
      input_map { key: "input__0" value: "preprocessed_image" }
      output_map { key: "output__0" value: "CLASS" }
    }
  ]
}

go deeper

for a junior

Be able to say that an ensemble chains several models inside the server so the client makes one call, and that platform: "ensemble" plus ensemble_scheduling steps is how it is declared.

for a middle

Explain the wiring: input_map and output_map bind step tensor names to shared pipeline names, and Triton derives the execution DAG from those names rather than from the listed order.

for a senior

Show that you tune stages independently — instance counts and batching per step — and that you diagnose a slow pipeline from per-step metrics rather than the ensemble total.

for a principal

Own the boundary question: which orchestration belongs in-server as an ensemble versus in an external service, given that an ensemble buys latency but couples deployment and versioning of every stage.

## The problem an ensemble solves Real inference is rarely one model call. A vision endpoint decodes a JPEG, resizes and normalises it, runs a network, then turns logits into labels. If each stage is its own deployed service, the client — or an orchestration service — makes three network calls and ships the intermediate tensors over the wire twice. For a 224×224×3 float tensor that is around 600 KB per hop, and it is pure overhead: serialization, HTTP or gRPC framing, and two extra round-trip latencies for work that could have happened a few microseconds apart in the same process. A Triton **ensemble** collapses that into one request. ## How it is declared An ensemble is a directory in the model repository like any other model, with a version directory (usually empty) and a `config.pbtxt`. The config sets `platform: "ensemble"`, declares the pipeline's own `input` and `output` tensors, and contains an `ensemble_scheduling` block with a repeated `step`: - `model_name` — the model in the repository this step calls. - `model_version` — a specific version, or `-1` for the latest available. - `input_map` — for each of that model's input names (the `key`), the pipeline-level tensor name to feed it (the `value`). - `output_map` — for each of that model's output names (the `key`), the pipeline-level name to publish it under (the `value`). Those pipeline-level names are the wiring. If step 1 publishes `preprocessed_image` and step 2 consumes `preprocessed_image`, Triton has learned the dependency. You never list an execution order: the order is inferred from the name graph, and steps whose inputs are ready run concurrently. That means fan-out (two models reading the same preprocessed tensor) is free, and so is fan-in. ## What it costs and what it saves The ensemble scheduler itself is not a backend. It allocates no device memory for compute and runs no kernels — it moves tensor handles between steps and manages the request's lifetime. So the saving is real and the overhead is small: one client round-trip instead of N, no repeated (de)serialization of intermediate tensors, and no client code orchestrating the chain. It does **not** save GPU memory: each step is still a separately loaded model with its own weights and instances. It does not fuse kernels. And it does not remove per-step scheduling: every step goes through its own model's scheduler, so a batched step still waits for its own dynamic batcher. That last point is also a feature. Because each step is an independent model, you tune them independently — eight CPU instances for the Python pre-processing model, one GPU instance with dynamic batching for the network. Sizing the stages separately is usually what actually makes the pipeline fast. ## Batching inside an ensemble If the ensemble declares `max_batch_size` greater than zero, the batch dimension flows through the pipeline, and every step must also support batching. A subtlety worth knowing: batches are re-formed per step. Requests batched together at the ensemble level are handed to step 1, whose own scheduler may group them differently, and so on. The composed models must therefore agree on shapes; a mismatch between what one step emits and what the next expects is the most common ensemble configuration error, and it surfaces at load time or first request rather than at config parse time. ## The limits An ensemble is a **static DAG**. There is no `if`, no loop, no early exit, no retry, no choosing a model based on the content of a tensor. Data-dependent control flow needs Business Logic Scripting in the Python backend instead, where you write ordinary Python and call other models programmatically. Many production pipelines use both: an ensemble for the fixed spine, a Python-backend step inside it for the conditional part. Error handling is all-or-nothing: if any step returns an error, the ensemble request fails with it. There is no partial result and no fallback path. ## Observability Because the steps are real models, they each report their own metrics and their own success and execution counts. When an ensemble is slow, look at the per-step latencies rather than at the ensemble as one number — a pipeline is almost always limited by one stage, and the per-model statistics tell you which.

  • Does an ensemble step run in a fixed order, and how does Triton know the order?
    There is no declared order. Triton builds a dependency graph from the pipeline tensor names in each step's input_map and output_map: a step runs as soon as every tensor it consumes has been produced. Steps with no dependency between them run concurrently, so a fan-out to two models from one preprocessed tensor happens in parallel without any extra configuration.
  • If an ensemble is slow, how do you find which stage is responsible?
    Each step is a normal model, so it has its own statistics and Prometheus metrics — queue duration, compute input, compute infer and compute output durations, execution count. Compare those per step rather than looking at the ensemble total. Typically one stage dominates, and the fix is that stage's instance count or batching settings, not the ensemble config.
  • Can an ensemble step be another ensemble?
    Yes — a step may name an ensemble model, so pipelines compose. That is useful for factoring a shared pre-processing chain out of several endpoints. Keep the nesting shallow though: debugging shape mismatches gets much harder when the failing tensor name lives two graphs down, and each level still resolves and schedules its steps independently.

saying these in an interview costs you the question

  • Thinking an ensemble is model ensembling or voting
  • Believing the ensemble runs on its own backend
  • Expecting conditionals or loops in ensemble_scheduling
  • Assuming steps share one copy of GPU weights
  • Listing steps and assuming they run in written order

context

open as a page

In Triton, how do preferred_batch_size and max_queue_delay_microseconds shape dynamic batching?

level: middleimportance: must knowfreq 70%

basics

~20 s

Triton's dynamic batcher holds arriving requests briefly and merges them into one model execution. preferred_batch_size lists batch sizes worth executing immediately; max_queue_delay_microseconds caps how long a partial batch waits for more requests before it runs anyway.

open as a page

In Triton, what does raising instance_group count buy, and what does it cost?

level: middleimportance: must knowfreq 58%

basics

~20 s

Each instance is a separate loaded copy of the model that can execute concurrently, so raising count lets Triton run several executions at once and overlap compute with input/output copies. The cost is one full set of weights per instance plus contention for the same GPU.

open as a page

When would you use Triton BLS in the Python backend instead of an ensemble?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Use Business Logic Scripting when the pipeline needs data-dependent control flow. Inside a Python-backend model you build a pb_utils.InferenceRequest and call exec() to invoke other loaded models, so you can branch, loop, retry or choose a model at runtime — none of which a static ensemble DAG can express.

open as a page

How do you prove a Triton batching config change actually improved throughput?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Sweep load with perf_analyzer at increasing concurrency and compare throughput against p95 latency before and after. Then confirm the mechanism: Triton's execution count versus successful request count gives the achieved average batch size, and the queue-time component shows what the wait actually cost.

open as a page

How does Triton's sequence batcher keep a stateful model's requests together?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Each request carries a correlation ID plus start and end flags. Triton's sequence batcher routes every request sharing a correlation ID to the same model instance, so the state that instance holds stays valid, and it injects control tensors telling the model when a sequence begins and ends.

open as a page