skip to content

NVIDIA Triton and TensorRT-LLM

You will learn NVIDIA's serving stack: Triton Inference Server as a multi-framework, multi-model front end, and TensorRT-LLM as the compiled backend that squeezes the most out of NVIDIA GPUs. Interviewers raise it for enterprise and on-prem deployments where one server must host LLMs alongside classic ML models.

on this pageshow

questions

18

What is a Triton ensemble model, and what cost does it remove from a pipeline?

level: juniorimportance: must knowfreq 52%

answer

  1. platform: "ensemble", no backend, no compute
  2. steps wired by tensor names
  3. input_map and output_map, not an ordered list
  4. one client call, no intermediate round-trips
  5. static DAG: no branching, no loops

basics

~20 s

A Triton ensemble is a scheduling-only model, declared with platform: "ensemble", that wires several real models into a pipeline inside the server. The client sends one request instead of three, and intermediate tensors never leave the server, removing network round-trips and serialization between stages.

solid answer

~50 s

An ensemble has no backend and runs no compute of its own. Its `config.pbtxt` sets `platform: "ensemble"` and an `ensemble_scheduling` block listing `step` entries; each step names a model and version, and uses `input_map` and `output_map` to bind that model's tensor names to shared pipeline names. Triton builds a DAG from those names and runs steps as soon as their inputs exist, so independent branches run concurrently. The win is that a typical preprocess → model → postprocess chain becomes one client call: intermediate tensors stay in server memory rather than being serialized back to the client and re-uploaded, which for image or audio payloads dominates end-to-end latency. Each step remains an ordinary model in the repository, so it keeps its own instance count and batching settings and can be tuned independently. The limits: an ensemble is a static graph with no conditionals or loops, and any step's failure fails the whole request.

code

protobuf · 21 lines
protobuf
name: "image_pipeline"
platform: "ensemble"
max_batch_size: 8
input [ { name: "IMAGE" data_type: TYPE_UINT8 dims: [ -1 ] } ]
output [ { name: "CLASS" data_type: TYPE_FP32 dims: [ 1000 ] } ]
ensemble_scheduling {
  step [
    {
      model_name: "preprocess"
      model_version: -1
      input_map { key: "RAW_IMAGE" value: "IMAGE" }
      output_map { key: "PREPROCESSED" value: "preprocessed_image" }
    },
    {
      model_name: "resnet50"
      model_version: -1
      input_map { key: "input__0" value: "preprocessed_image" }
      output_map { key: "output__0" value: "CLASS" }
    }
  ]
}

go deeper

for a junior

Be able to say that an ensemble chains several models inside the server so the client makes one call, and that platform: "ensemble" plus ensemble_scheduling steps is how it is declared.

for a middle

Explain the wiring: input_map and output_map bind step tensor names to shared pipeline names, and Triton derives the execution DAG from those names rather than from the listed order.

for a senior

Show that you tune stages independently — instance counts and batching per step — and that you diagnose a slow pipeline from per-step metrics rather than the ensemble total.

for a principal

Own the boundary question: which orchestration belongs in-server as an ensemble versus in an external service, given that an ensemble buys latency but couples deployment and versioning of every stage.

## The problem an ensemble solves Real inference is rarely one model call. A vision endpoint decodes a JPEG, resizes and normalises it, runs a network, then turns logits into labels. If each stage is its own deployed service, the client — or an orchestration service — makes three network calls and ships the intermediate tensors over the wire twice. For a 224×224×3 float tensor that is around 600 KB per hop, and it is pure overhead: serialization, HTTP or gRPC framing, and two extra round-trip latencies for work that could have happened a few microseconds apart in the same process. A Triton **ensemble** collapses that into one request. ## How it is declared An ensemble is a directory in the model repository like any other model, with a version directory (usually empty) and a `config.pbtxt`. The config sets `platform: "ensemble"`, declares the pipeline's own `input` and `output` tensors, and contains an `ensemble_scheduling` block with a repeated `step`: - `model_name` — the model in the repository this step calls. - `model_version` — a specific version, or `-1` for the latest available. - `input_map` — for each of that model's input names (the `key`), the pipeline-level tensor name to feed it (the `value`). - `output_map` — for each of that model's output names (the `key`), the pipeline-level name to publish it under (the `value`). Those pipeline-level names are the wiring. If step 1 publishes `preprocessed_image` and step 2 consumes `preprocessed_image`, Triton has learned the dependency. You never list an execution order: the order is inferred from the name graph, and steps whose inputs are ready run concurrently. That means fan-out (two models reading the same preprocessed tensor) is free, and so is fan-in. ## What it costs and what it saves The ensemble scheduler itself is not a backend. It allocates no device memory for compute and runs no kernels — it moves tensor handles between steps and manages the request's lifetime. So the saving is real and the overhead is small: one client round-trip instead of N, no repeated (de)serialization of intermediate tensors, and no client code orchestrating the chain. It does **not** save GPU memory: each step is still a separately loaded model with its own weights and instances. It does not fuse kernels. And it does not remove per-step scheduling: every step goes through its own model's scheduler, so a batched step still waits for its own dynamic batcher. That last point is also a feature. Because each step is an independent model, you tune them independently — eight CPU instances for the Python pre-processing model, one GPU instance with dynamic batching for the network. Sizing the stages separately is usually what actually makes the pipeline fast. ## Batching inside an ensemble If the ensemble declares `max_batch_size` greater than zero, the batch dimension flows through the pipeline, and every step must also support batching. A subtlety worth knowing: batches are re-formed per step. Requests batched together at the ensemble level are handed to step 1, whose own scheduler may group them differently, and so on. The composed models must therefore agree on shapes; a mismatch between what one step emits and what the next expects is the most common ensemble configuration error, and it surfaces at load time or first request rather than at config parse time. ## The limits An ensemble is a **static DAG**. There is no `if`, no loop, no early exit, no retry, no choosing a model based on the content of a tensor. Data-dependent control flow needs Business Logic Scripting in the Python backend instead, where you write ordinary Python and call other models programmatically. Many production pipelines use both: an ensemble for the fixed spine, a Python-backend step inside it for the conditional part. Error handling is all-or-nothing: if any step returns an error, the ensemble request fails with it. There is no partial result and no fallback path. ## Observability Because the steps are real models, they each report their own metrics and their own success and execution counts. When an ensemble is slow, look at the per-step latencies rather than at the ensemble as one number — a pipeline is almost always limited by one stage, and the per-model statistics tell you which.

  • Does an ensemble step run in a fixed order, and how does Triton know the order?
    There is no declared order. Triton builds a dependency graph from the pipeline tensor names in each step's input_map and output_map: a step runs as soon as every tensor it consumes has been produced. Steps with no dependency between them run concurrently, so a fan-out to two models from one preprocessed tensor happens in parallel without any extra configuration.
  • If an ensemble is slow, how do you find which stage is responsible?
    Each step is a normal model, so it has its own statistics and Prometheus metrics — queue duration, compute input, compute infer and compute output durations, execution count. Compare those per step rather than looking at the ensemble total. Typically one stage dominates, and the fix is that stage's instance count or batching settings, not the ensemble config.
  • Can an ensemble step be another ensemble?
    Yes — a step may name an ensemble model, so pipelines compose. That is useful for factoring a shared pre-processing chain out of several endpoints. Keep the nesting shallow though: debugging shape mismatches gets much harder when the failing tensor name lives two graphs down, and each level still resolves and schedules its steps independently.

saying these in an interview costs you the question

  • Thinking an ensemble is model ensembling or voting
  • Believing the ensemble runs on its own backend
  • Expecting conditionals or loops in ensemble_scheduling
  • Assuming steps share one copy of GPU weights
  • Listing steps and assuming they run in written order

context

open as a page

How must a Triton model repository be laid out on disk for a model to load?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A Triton model repository is a directory of model directories. Each model directory holds an optional config.pbtxt plus one or more numerically named version subdirectories, and the model file itself lives inside a version directory.

open as a page

What does trtllm-build produce in TensorRT-LLM, and why is that artifact not portable?

level: middleimportance: must knowfreq 60%

basics

~20 s

trtllm-build compiles a converted checkpoint into a serialized TensorRT engine: fused kernels, hardware-specific tactics and fixed shape limits baked into a binary. That binary is tied to the GPU architecture, the precision and the TensorRT-LLM version that built it.

open as a page

How do you convert a Hugging Face checkpoint into a TensorRT-LLM engine?

level: middleimportance: must knowfreq 45%

basics

~20 s

Two steps. A per-model convert_checkpoint.py script rewrites the Hugging Face weights into a TensorRT-LLM checkpoint directory — a config.json plus one weight file per rank, sharded by the --tp_size you choose. Then trtllm-build compiles that directory into an engine.

open as a page

In Triton, how do preferred_batch_size and max_queue_delay_microseconds shape dynamic batching?

level: middleimportance: must knowfreq 70%

basics

~20 s

Triton's dynamic batcher holds arriving requests briefly and merges them into one model execution. preferred_batch_size lists batch sizes worth executing immediately; max_queue_delay_microseconds caps how long a partial batch waits for more requests before it runs anyway.

open as a page

In Triton, what does raising instance_group count buy, and what does it cost?

level: middleimportance: must knowfreq 58%

basics

~20 s

Each instance is a separate loaded copy of the model that can execute concurrently, so raising count lets Triton run several executions at once and overlap compute with input/output copies. The cost is one full set of weights per instance plus contention for the same GPU.

open as a page

In a Triton config.pbtxt, what do max_batch_size and dims declare together?

level: middleimportance: must knowfreq 64%

basics

~20 s

When max_batch_size is greater than 0, Triton assumes a variable leading batch dimension the model accepts, and dims lists the shape of a single instance without it. When max_batch_size is 0 the model cannot batch, and dims must be the complete tensor shape.

open as a page

How do you load a new model into a running Triton server without restarting it?

level: middleimportance: must knowfreq 58%

basics

~10 s

Start the server with --model-control-mode=explicit and then call the repository API: POST /v2/repository/models/<name>/load to load and /unload to remove it. The alternative is poll mode, where Triton rescans the repository on a timer.

open as a page

Which tensorrtllm_backend settings does Triton need to serve a TensorRT-LLM engine?

level: middleimportance: should knowfreq 42%

basics

~10 s

The tensorrt_llm model's config needs engine_dir pointing at the built engine and batching_strategy set to inflight_fused_batching, plus triton_max_batch_size, decoupled_mode true for streaming, and a KV-cache memory fraction. The preprocessing and postprocessing models need tokenizer_dir.

open as a page

How does Triton decide which backend executes a model in its repository?

level: middleimportance: should knowfreq 54%

basics

~20 s

Triton picks the backend from the model's configuration: either an explicit backend field such as "python" or "vllm", or a platform field such as "onnxruntime_onnx" or "tensorrt_plan". If neither is set, auto-complete infers it from the artifact filename in the version directory.

open as a page

Your TensorRT-LLM engine rejects a 32k-token request; max_seq_len was 8192. Now what?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Rebuild. max_input_len, max_seq_len, max_batch_size and max_num_tokens are compile-time ceilings serialized into the engine's optimization profiles; the serving layer can only configure values at or below them. No runtime setting raises a limit the engine was not built for.

open as a page

In TensorRT-LLM, how do FP8 or INT4-AWQ get baked into an engine?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A calibration pass runs before the build: the quantization script reads the model with a --qformat such as fp8 or int4_awq, computes scales over a calibration set, and writes a quantized checkpoint. trtllm-build then compiles kernels for that format, so precision is frozen in the engine, not switchable at serve time.

open as a page

When would you use Triton BLS in the Python backend instead of an ensemble?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Use Business Logic Scripting when the pipeline needs data-dependent control flow. Inside a Python-backend model you build a pb_utils.InferenceRequest and call exec() to invoke other loaded models, so you can branch, loop, retry or choose a model at runtime — none of which a static ensemble DAG can express.

open as a page

How do you prove a Triton batching config change actually improved throughput?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Sweep load with perf_analyzer at increasing concurrency and compare throughput against p95 latency before and after. Then confirm the mechanism: Triton's execution count versus successful request count gives the achieved average batch size, and the queue-time component shows what the wait actually cost.

open as a page

How does Triton's sequence batcher keep a stateful model's requests together?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Each request carries a correlation ID plus start and end flags. Triton's sequence batcher routes every request sharing a correlation ID to the same model instance, so the state that instance holds stays valid, and it injects control tensors telling the model when a sequence begins and ends.

open as a page

In Triton, how do you roll out model version 2 and roll back without downtime?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Versions are numbered directories under the same model directory, so rollout is adding a 2/ directory and rollback is pointing version_policy back at 1. Clients keep calling the same model name; only the version Triton loads changes.

open as a page

When is TensorRT-LLM's ahead-of-time engine build worth its cost versus a Python-level engine?

level: principalimportance: should knowfreq 30%

basics

~20 s

When the model set is stable, the GPU fleet is uniform, and volume is high enough that a throughput gain pays for a build-and-validate pipeline producing one artifact per model, precision, parallel degree and GPU SKU. Fast-moving model choice or mixed hardware argues the other way.

open as a page

Should a platform team host 40 models in one Triton server or one per model?

level: principalimportance: should knowfreq 36%

basics

~20 s

Co-locate models that are small, share a GPU comfortably and tolerate a shared failure domain; isolate anything that wants a whole GPU or a different container image. Triton's repository and load API already give deploy independence inside one server, so the reason to split is resource and blast radius.

open as a page