In Triton, how do preferred_batch_size and max_queue_delay_microseconds shape dynamic batching?
answer
- server-side merge of independent requests
- two knobs: preferred sizes, max wait
- delay is a ceiling, not a cost
- default delay is zero microseconds
- low QPS: pure latency, no batch
basics
~20 sTriton's dynamic batcher holds arriving requests briefly and merges them into one model execution. preferred_batch_size lists batch sizes worth executing immediately; max_queue_delay_microseconds caps how long a partial batch waits for more requests before it runs anyway.
solid answer
~50 sDynamic batching is server-side: independent clients each send a single request, and Triton's scheduler combines them into one execution so the GPU does more work per kernel launch. You turn it on by adding a `dynamic_batching { }` block to the model's `config.pbtxt`. `preferred_batch_size` is a list of sizes the scheduler will dispatch the moment it can form one — if you omit it, Triton picks sizes automatically. `max_queue_delay_microseconds` is the only latency knob: it is the maximum time a not-yet-preferred batch will linger in the queue waiting for company. Its default is 0, meaning "dispatch whatever is queued as soon as an instance is free". Raising it trades tail latency for batch size, and it only pays when arrival rate is high enough that requests actually show up during the wait. Related knobs in the same block: `preserve_ordering`, `priority_levels`, and `default_queue_policy` with `max_queue_size` and `timeout_action`. This scheduler is for stateless, non-generative models; the TensorRT-LLM backend uses in-flight batching instead.
code
protobuf · 12 linesname: "resnet50"
platform: "tensorrt_plan"
max_batch_size: 32
dynamic_batching {
preferred_batch_size: [ 8, 16 ]
max_queue_delay_microseconds: 2000
default_queue_policy {
max_queue_size: 128
timeout_action: REJECT
default_timeout_microseconds: 50000
}
}go deeper
Know that dynamic batching happens on the server: separate clients send single requests and Triton merges them into one execution. Say that you enable it with a dynamic_batching block in the model's config.pbtxt.
Explain the two knobs precisely — preferred_batch_size dispatches immediately when reachable, max_queue_delay_microseconds bounds how long a partial batch waits — and state that the delay default is 0.
Show that you measure rather than guess: tie the delay setting to observed arrival rate, prove the achieved batch size moved, and know when to reject via queue policy instead of queueing forever.
Own the tradeoff across a fleet: whether to spend latency budget on batching or buy more replicas, when priority levels beat a second deployment, and which models should not share a scheduler at all.
## What the dynamic batcher is A GPU is far more efficient running one kernel over 16 inputs than 16 kernels over one input each. But your clients are independent — a browser, a mobile app, a batch job — and none of them knows about the others. Triton's **dynamic batcher** closes that gap on the server: it accepts single-request inferences, queues them for a few microseconds, and hands the backend one batched execution. This is purely a scheduler. It does not change the model, and clients need no code change. It is enabled per model, in that model's `config.pbtxt`, by adding a `dynamic_batching` section. The model must actually support batching — its first tensor dimension is the batch dimension and `max_batch_size` must be greater than zero — otherwise the scheduler has nothing to combine. ## preferred_batch_size `preferred_batch_size` is a repeated field listing batch sizes that are "good" for this model, for example `[ 8, 16 ]`. When the queue contains enough requests to form one of those sizes, the scheduler dispatches immediately rather than waiting for anything better. If you do not specify it, Triton chooses preferred sizes on its own based on the model configuration, so it is an optimisation hint, not a requirement. Why would some sizes be better than others? Compiled backends often have kernels or optimisation profiles tuned around particular shapes, and a TensorRT engine built with an optimisation profile centred on batch 8 will run batch 8 more efficiently than batch 7. Padding, tiling and occupancy all favour round numbers. If you have no such structure, leaving the field out is fine. ## max_queue_delay_microseconds This is the knob interviewers actually probe, because it is the one that spends latency. It is the maximum time the scheduler will hold a batch that has not reached a preferred size, hoping more requests arrive. The default is **0**: form a batch from whatever is queued and dispatch as soon as a model instance is free. The key asymmetry: the delay is an upper bound, not a fixed cost. Under heavy load, batches fill before the timer expires and the added latency is near zero. Under light load, nothing arrives during the wait, so you pay the full delay on every request and gain almost no batch size. That is why blindly setting a 50 ms delay on a low-QPS endpoint makes the service strictly worse — a very common wrong answer. A reasonable starting point is a delay on the order of one model execution time (single-digit milliseconds for most vision models), then measure. The thing you are measuring is not just throughput: watch the queue-time component of latency and the achieved average batch size. If batch size did not move, the delay bought nothing. ## How it interacts with instances Batches are dispatched to a free model instance. Batching and instance count are complementary but pull in different directions on a saturated GPU: more instances add concurrent executions and memory pressure, larger batches add work per execution. On a GPU-bound model, a bigger batch is usually the cheaper win. ## Ragged inputs and queue policy Requests with different input shapes cannot always be stacked. Triton pads or refuses depending on the backend; for backends that support it, `allow_ragged_batch` on an input lets variable-length inputs be concatenated instead of padded, with a `batch_input` telling the model where each element starts. Separately, `default_queue_policy` inside the `dynamic_batching` block bounds the queue itself: `max_queue_size` caps depth, and `timeout_action` decides whether an over-aged request is rejected or delayed. Under overload, rejecting fast is often better than letting a queue grow beyond any useful latency. `priority_levels` plus per-level queue policies let you run interactive traffic ahead of bulk traffic through the same model, which is a cheap alternative to standing up a second deployment. ## Where it does not apply The dynamic batcher forms a batch, runs it, and returns — a request/response shape. Generative LLM serving needs a batch whose membership changes every decoding step, which is a different scheduler entirely: Triton's TensorRT-LLM and vLLM backends do their own in-flight/continuous batching internally, and you do not configure it with `dynamic_batching`. Similarly, stateful models where request N depends on request N-1 need the sequence batcher, not this one.
- You set max_queue_delay_microseconds to 5000 and p99 latency rose 5 ms with no throughput gain. What happened?Arrival rate is too low for the wait to collect anything. The timer expires before a second request shows up, so every request pays the full delay and still executes at batch size one. Confirm it by comparing execution count against successful request count — if the ratio is near one, batching is not happening. Either lower the delay back toward zero or consolidate traffic onto fewer replicas so each one sees enough concurrency to batch.
- Does dynamic batching help a model that is already CPU-bound in its pre-processing step?Not much. Batching amortises GPU kernel launch and improves GPU occupancy; if the bottleneck is Python-side or CPU pre-processing that scales linearly with batch elements, a bigger batch just does the same serial work in one call. The fix there is more model instances, or moving pre-processing into its own model in an ensemble so it can be scaled and batched independently.
- How would you serve latency-sensitive and bulk traffic through one batched model?Use the dynamic batcher's `priority_levels` with `default_priority_level`, and have clients tag requests with a priority. Triton keeps a queue per level and drains higher-priority queues first, so bulk work does not sit in front of interactive work. Give each level its own queue policy so bulk traffic can carry a deeper queue and a longer timeout while interactive traffic is rejected quickly rather than queued.
saying these in an interview costs you the question
- Claiming the client must batch requests itself
- Treating max_queue_delay as a fixed added latency
- Setting a large delay on low-QPS endpoints
- Thinking dynamic batching handles LLM token generation
- Believing preferred_batch_size is mandatory to enable batching