skip to content

In a Triton config.pbtxt, what do max_batch_size and dims declare together?

level: middleimportance: must knowfreq 64%

answer

  1. dims describes one instance, not the batch
  2. The leading dimension is implied, not listed
  3. Zero means the model cannot batch
  4. -1 marks a per-request variable dimension
  5. Duplicating the batch dim fails at execute time

basics

~20 s

When max_batch_size is greater than 0, Triton assumes a variable leading batch dimension the model accepts, and dims lists the shape of a single instance without it. When max_batch_size is 0 the model cannot batch, and dims must be the complete tensor shape.

solid answer

~50 s

`max_batch_size` is the promise that the model's first tensor dimension is a variable batch dimension, and its value caps how large a batch Triton will assemble. Because that dimension is implied, the `dims` in each `input`/`output` block describe **one instance only**: with `max_batch_size: 8` and `dims: [3, 224, 224]`, a four-request batch arrives at the model as `[4, 3, 224, 224]`. If a model was exported with a fixed batch dimension, or genuinely takes a whole-shape tensor, you set `max_batch_size: 0`; then `dims` is the full shape exactly as the model expects it and Triton batches nothing. A `-1` entry in `dims` marks a dimension that varies per request. The commonly seen error is duplicating the batch dimension — `max_batch_size: 8` with `dims: [8, 3, 224, 224]` — which makes the model see a rank-5 tensor and fail at execute time, not at load time.

code

protobuf · 17 lines
protobuf
name: "resnet50"
platform: "onnxruntime_onnx"
max_batch_size: 8
input [
  {
    name: "input"
    data_type: TYPE_FP32
    dims: [ 3, 224, 224 ]
  }
]
output [
  {
    name: "logits"
    data_type: TYPE_FP32
    dims: [ 1000 ]
  }
]

go deeper

for a junior

Know that config.pbtxt declares each input and output with a name, a data_type such as TYPE_FP32, and dims, and that max_batch_size is where batching is switched on or off.

for a middle

Be ready to state the rule out loud: with max_batch_size greater than zero the leading batch dimension is implied and must not appear in dims, and with zero dims is the full shape.

for a senior

Diagnose the failure mode from a symptom — a model that works with one request and fails under concurrency is usually a batch-dimension mismatch — and justify the max_batch_size ceiling from activation memory and tail latency.

for a principal

Treat the shape contract as a platform standard: whether auto-complete is allowed at all, whether --disable-auto-complete-config is enforced in production, and who reviews config.pbtxt before an artifact is promoted.

## Two fields, one shape contract A Triton model configuration describes each tensor with a `name`, a `data_type` and `dims`. What makes `dims` subtle is that it does not stand alone: it is interpreted relative to `max_batch_size`, and the two together are the shape contract between Triton and the backend. Getting this pair wrong is the single most common reason a model loads cleanly and then fails on the first request. ## max_batch_size > 0: the implicit leading dimension Setting `max_batch_size: N` asserts two things. First, the model accepts a **variable-size first dimension** — this is what "the model supports batching" means in Triton's vocabulary. Second, Triton will never assemble a batch larger than `N` for that model. Because the batch dimension is implied, you must **omit it from `dims`**. A classifier that takes `[batch, 3, 224, 224]` is configured as: ``` max_batch_size: 8 input [ { name: "data_0" data_type: TYPE_FP32 dims: [ 3, 224, 224 ] } ] ``` A client sending one request supplies a tensor of shape `[1, 3, 224, 224]` — the batch dimension is present on the wire, it is only absent from the configuration. If Triton combines four such requests, the backend sees `[4, 3, 224, 224]`. The same rule applies to outputs: `dims: [1000]` with `max_batch_size: 8` means the model returns `[batch, 1000]`, and Triton splits that back into per-request responses. ## max_batch_size: 0 — batching off `max_batch_size: 0` says the model does **not** support a variable leading dimension. Now `dims` is the literal, complete shape the model expects, batch dimension included if the model has a fixed one: ``` max_batch_size: 0 input [ { name: "input" data_type: TYPE_FP32 dims: [ 1, 3, 224, 224 ] } ] ``` This is the correct configuration for a graph exported with a hard-coded batch of 1, and for models whose first dimension means something other than batch — sequence length, for instance. The trade-off is real: with `max_batch_size: 0` Triton executes exactly one request per model execution, and the scheduling features that combine requests are unavailable to that model, because they all require a variable batch dimension to combine along. ## Variable dimensions Any entry in `dims` may be `-1`, meaning "variable, supplied per request". Text and sequence models lean on this heavily: `dims: [-1]` for a token-id vector of unknown length. Use it deliberately rather than everywhere — a fully variable shape gives Triton and the backend less to validate, so genuine client mistakes surface later, inside the framework, with a worse error message. The `reshape` field exists for the mismatch case: when the model's tensor has no batch dimension at all but you still want Triton to batch, `reshape { shape: [ ] }` lets Triton insert and remove the dimension around execution. ## Data types and names `data_type` uses Triton's own spelling — `TYPE_FP32`, `TYPE_FP16`, `TYPE_INT32`, `TYPE_INT64`, `TYPE_BOOL`, `TYPE_STRING`, `TYPE_UINT8` — which is then mapped onto the backend's native type. `name` must match the tensor name in the artifact exactly (the ONNX graph input name, or the Python backend's `pb_utils.get_input_tensor_by_name` argument). A typo here fails at load for backends that can introspect their model, and at execute for those that cannot. ## An input may be optional An input block with `optional: true` is one clients may omit; the backend must handle its absence. This is useful for models that take an optional mask or per-request parameter, and it is preferable to inventing a sentinel value. ## When you can skip the whole file For TensorRT, ONNX, TensorFlow SavedModel and OpenVINO models the artifact already carries its shape metadata, so Triton's auto-complete can generate the configuration and `config.pbtxt` becomes optional; the Python backend can do the same by implementing `auto_complete_config`. Auto-complete is on by default and `--disable-auto-complete-config` turns it off. Note what auto-complete infers: shapes and types, and a batching setting derived from whether the artifact's first dimension is variable. It cannot guess your throughput intent, so in production most teams write the file explicitly anyway — it is the place where shape, batching cap and scheduling live where a reviewer can see them. ## Choosing the number `max_batch_size` is a ceiling, not a target: it bounds memory per execution and bounds the worst-case latency of a batched execution. Set it from what one execution's activations cost on your GPU and from the latency you can accept for the last request in a full batch — then verify under load rather than trusting the arithmetic.

  • A model exported with a hard-coded batch size of 1 is configured with max_batch_size: 8. What breaks?
    The model will refuse the batched tensor: Triton sends `[N, ...]` where the graph only accepts `[1, ...]`, so execution fails as soon as two requests are combined — and often passes in single-request testing, which is why it reaches production. Either set `max_batch_size: 0` and put the full shape in `dims`, or re-export the model with a dynamic batch dimension.
  • Does raising max_batch_size on its own make a model batch more?
    No. It only raises the ceiling. Whether requests are actually combined depends on the scheduler configured for the model and on whether requests are in flight at the same time; with a low arrival rate you keep executing batches of one. It also does not shrink memory: a higher ceiling means the worst-case activation footprint of one execution is larger.
  • When would you set a dims entry to -1 rather than a fixed size?
    When the dimension genuinely varies per request — token counts, sequence lengths, variable-length audio. Fixed sizes are better when you know them, because Triton rejects malformed requests up front with a clear shape error instead of letting the framework fail deeper. Use -1 where variation is real, not as a way to silence shape mismatches.

saying these in an interview costs you the question

  • Listing the batch dimension in dims alongside max_batch_size > 0
  • Thinking max_batch_size: 0 means unlimited batching
  • Believing a higher max_batch_size by itself increases achieved batch size
  • Assuming dims must match the client's wire tensor exactly
  • Setting every dimension to -1 to avoid shape errors

context