Which tensorrtllm_backend settings does Triton need to serve a TensorRT-LLM engine?
answer
- five models, not one, in the repo
- engine_dir points at the build output
- in-flight batching is a named strategy value
- streaming needs a decoupled model
- tokenizer lives on pre/post, not in the engine
basics
~10 sThe tensorrt_llm model's config needs engine_dir pointing at the built engine and batching_strategy set to inflight_fused_batching, plus triton_max_batch_size, decoupled_mode true for streaming, and a KV-cache memory fraction. The preprocessing and postprocessing models need tokenizer_dir.
solid answer
~40 sThe TensorRT-LLM backend is driven by parameters in the `tensorrt_llm` model's config rather than by anything inside the engine. The must-set ones are **`engine_dir`** (the directory `trtllm-build` produced) and **`batching_strategy`**, which you set to `inflight_fused_batching` for in-flight batching or to `V1` to fall back to static batching. Around those you set `triton_max_batch_size`, `decoupled_mode: true` if clients stream tokens, `max_queue_size` and `max_queue_delay_microseconds` for admission behaviour, `gpu_device_ids` when pinning ranks, and a KV-cache free-memory fraction that decides how much spare VRAM becomes cache. The deployment is normally not one model but the five-model template set — `preprocessing`, `tensorrt_llm`, `postprocessing`, an `ensemble` and a BLS variant — where the pre/post models need `tokenizer_dir` because the engine carries no tokenizer. Those templates now live in the NVIDIA/TensorRT-LLM repository under `triton_backend/all_models/inflight_batcher_llm`.
code
bash · 7 linespython3 tools/fill_template.py -i all_models/inflight_batcher_llm/tensorrt_llm/config.pbtxt \
engine_dir:/engines/llama3-8b-tp2,\
batching_strategy:inflight_fused_batching,\
triton_max_batch_size:64,\
decoupled_mode:true,\
kv_cache_free_gpu_mem_fraction:0.9,\
batch_scheduler_policy:max_utilizationgo deeper
Know that serving a TensorRT-LLM engine in Triton means a small pipeline of models — tokenize, run the engine, detokenize — not a single model file dropped into a repository.
Be able to name engine_dir, batching_strategy set to inflight_fused_batching, triton_max_batch_size, decoupled_mode and tokenizer_dir, and say what each one breaks when it is wrong.
Demonstrate tuning judgment: the KV-cache memory fraction against concurrency, scheduler policy against tail latency, and bounded queues so an overloaded server sheds load instead of hoarding it.
Own the generation of these configs. Argue for one source of truth that emits the build command and the model template together, so a fleet cannot drift into engine-versus-config mismatches across dozens of deployments.
## Why there are five models, not one A TensorRT-LLM engine takes token ids in and produces token ids out. Nothing in it turns text into ids. So the standard deployment is a small pipeline of Triton models: - **`preprocessing`** — a Python-backend model that tokenizes the request text. - **`tensorrt_llm`** — the TensorRT-LLM backend model that actually executes the engine. - **`postprocessing`** — detokenizes generated ids back into text. - **`ensemble`** — wires the three together so a client calls one endpoint. - **`tensorrt_llm_bls`** — a Business Logic Scripting alternative to the ensemble, used when the glue needs real control flow. You pick either the ensemble or the BLS front door; both sit over the same `tensorrt_llm` model. The templates ship as `config.pbtxt` files with `${placeholders}`, and a `fill_template.py` helper substitutes values, which is why deployment guides show a string of `key:value` pairs rather than hand-edited configs. ## The parameters that matter on the tensorrt_llm model **`engine_dir`** — the path to the engine directory. This is the parameter people get wrong from memory: the 0.x-era spelling was different, and current documentation uses `engine_dir`. For encoder-decoder models there is a companion `encoder_engine_dir`. **`batching_strategy`** — `inflight_fused_batching` turns on in-flight batching, where new requests join the running batch at iteration boundaries instead of waiting for the current batch to drain. Setting it to `V1` selects the older static behaviour. Note this is *not* Triton's generic `dynamic_batching` block: the TensorRT-LLM backend schedules internally, at iteration granularity, and the generic batcher is the wrong tool here. **`triton_max_batch_size`** — the batch ceiling Triton enforces for this model. It must not exceed the `--max_batch_size` compiled into the engine; you can go lower, never higher. **`decoupled_mode`** — set `true` when the model should emit many responses for one request, which is what token-by-token streaming requires. Leave it false and clients get a single response at the end. A deployment that "doesn't stream" is usually this flag. **`max_num_tokens`** — the runtime token budget per iteration, again bounded by what the engine was built with. **`max_queue_size` and `max_queue_delay_microseconds`** — how many requests may wait and how long the backend will hold a request before scheduling it. A bounded queue is what lets an overloaded server reject fast instead of accumulating an unbounded backlog whose latency nobody can meet. **`batch_scheduler_policy`** — `max_utilization` packs aggressively and accepts that a request may be paused and resumed when cache pressure spikes; `guaranteed_no_evict` only admits work it can carry to completion, trading some throughput for predictability. **`kv_cache_free_gpu_mem_fraction`** — the share of free VRAM after weights are loaded that becomes KV cache. Raise it for more concurrency, lower it if something else on the GPU needs room. **`gpu_device_ids`** — which GPUs this model's ranks occupy, which matters when several models share a node. **`instance_count`** — how many instances of the backend model run. For a multi-GPU engine, the mapping between ranks and instances is not the naive one; treat instance count for a TensorRT-LLM model as something to verify against the template guidance rather than something to increase reflexively. **`logits_datatype`** — the datatype of returned logits, when a client asks for them. It is not a knob for compute precision, which is fixed in the engine. **`tokenizer_dir`** — set on the preprocessing and postprocessing models, pointing at the original Hugging Face directory. Mismatch it with the checkpoint the engine was built from and you get garbled text out of a perfectly healthy engine. **`return_perf_metrics`** — turns on per-request performance metrics, useful when you want the backend's own view of queueing versus execution rather than only client-side timings. ## What the engine does not tell Triton None of the build ceilings are auto-discovered into your config. If the engine was built with `--max_batch_size 64` and you fill in `triton_max_batch_size: 128`, you have created a mismatch that surfaces as a load or runtime error rather than as graceful clamping. Keeping the build command and the config template beside each other in the same repository — generated from one source of truth — is the practice that prevents this class of bug. ## Diagnosing a deployment Three symptoms cover most first deployments. The model fails to load: check `engine_dir`, the world size versus `gpu_device_ids`, and that the engine's library version matches the container. Responses arrive all at once: `decoupled_mode`. Text comes back mangled: `tokenizer_dir` on pre/post. After that, the interesting questions become tuning ones — cache fraction against concurrency, scheduler policy against tail latency, queue bounds against overload behaviour. ## The interview shape Interviewers are checking whether you have deployed this rather than read about it. Naming `engine_dir` and `inflight_fused_batching` correctly, knowing that the tokenizer lives outside the engine, and knowing that `decoupled_mode` is what makes streaming work are the three signals that separate hands-on from hearsay.
- A client gets one big response instead of streamed tokens. What is misconfigured?`decoupled_mode` on the `tensorrt_llm` model is false, so the backend emits a single response per request. Set it to true and use a streaming-capable client call; the ensemble or BLS front door must be invoked in streaming mode too. Nothing about the engine build changes — this is purely a serving-layer setting.
- How do max_utilization and guaranteed_no_evict differ as scheduler policies?`max_utilization` admits as much work as it can and accepts that a running request may be paused and resumed when KV-cache pressure spikes, maximising throughput at the cost of tail-latency variance. `guaranteed_no_evict` only admits requests it can carry to completion with the cache it has, giving steadier per-request latency and lower peak throughput. Pick by whether your SLO is on tails or on tokens per dollar.
- Why does the deployment need tokenizer_dir at all if the engine already encodes the model?Because the engine consumes and produces token ids only — tokenizer files are not serialized into it. The preprocessing and postprocessing models load the tokenizer from `tokenizer_dir`, and it must match the exact model revision the weights came from. A mismatch yields healthy-looking inference with wrong text at the boundaries.
saying these in an interview costs you the question
- Names gpt_model_path, the removed 0.x-era parameter
- Adds a generic dynamic_batching block instead of setting batching_strategy
- Sets triton_max_batch_size above the engine's built ceiling
- Expects streaming to work with decoupled_mode left false
- Assumes the engine carries its own tokenizer