How does Triton's sequence batcher keep a stateful model's requests together?
answer
- state lives in one instance, not the server
- correlation ID plus start and end flags
- boundaries arrive as control input tensors
- direct pins a slot, oldest pins an instance
- idle timeout reclaims abandoned slots
basics
~20 sEach request carries a correlation ID plus start and end flags. Triton's sequence batcher routes every request sharing a correlation ID to the same model instance, so the state that instance holds stays valid, and it injects control tensors telling the model when a sequence begins and ends.
solid answer
~50 sYou enable it with a `sequence_batching` block in `config.pbtxt` instead of `dynamic_batching`. The client sets a sequence ID and the start/end flags on each inference request; the batcher then guarantees all requests for that ID reach the same model instance, in order, for the life of the sequence. It communicates sequence boundaries to the model through `control_input` entries — `CONTROL_SEQUENCE_START`, `CONTROL_SEQUENCE_END`, `CONTROL_SEQUENCE_READY` and `CONTROL_SEQUENCE_CORRID` — which appear to the model as ordinary input tensors, so the backend knows when to reset its state. Two strategies exist: `direct` pins a sequence to a fixed batch slot in the instance, which is what implicit-state and slot-indexed models need; `oldest` only pins it to an instance and lets requests from different sequences be batched together. Slots are finite, so a client that vanishes without sending the end flag would pin one forever — `max_sequence_idle_microseconds` reclaims it.
code
protobuf · 22 linessequence_batching {
max_sequence_idle_microseconds: 5000000
direct { }
control_input [
{
name: "START"
control [ { kind: CONTROL_SEQUENCE_START fp32_false_true: [ 0, 1 ] } ]
},
{
name: "END"
control [ { kind: CONTROL_SEQUENCE_END fp32_false_true: [ 0, 1 ] } ]
},
{
name: "READY"
control [ { kind: CONTROL_SEQUENCE_READY fp32_false_true: [ 0, 1 ] } ]
},
{
name: "CORRID"
control [ { kind: CONTROL_SEQUENCE_CORRID data_type: TYPE_UINT64 } ]
}
]
}go deeper
Know that stateful models use sequence_batching rather than dynamic_batching, and that the client tags requests with a correlation ID plus start and end flags.
Explain that the batcher routes a correlation ID to one instance and surfaces sequence boundaries to the model as control input tensors such as START, END, READY and CORRID.
Distinguish direct from oldest by what each pins, and name the operational traps: finite slots, idle timeout reclaiming abandoned sequences, and sticky routing required above multiple replicas.
Own the capacity model: sequences in flight, not requests per second, is the scaling unit for stateful serving, and that changes autoscaling signals, replica sizing and whether the state belongs in the server at all.
## Why a separate scheduler exists The dynamic batcher assumes requests are independent: any request can go to any instance in any batch. Stateful models break that assumption. A streaming speech recogniser, an online tracker or any recurrent model carries state from one call to the next, and that state lives inside a particular execution instance. Send call 2 of a session to a different instance and the state is gone. Triton's **sequence batcher** exists to preserve that affinity while still batching across *different* sessions. ## What the client must send A sequence is identified by a **correlation ID** — an unsigned integer or a string, chosen by the client and unique among in-flight sequences. On each inference request the client sets: - the sequence ID, - a **start** flag on the first request of the sequence, - an **end** flag on the last request. In the Python client these are the `sequence_id`, `sequence_start` and `sequence_end` arguments to `infer`. Without those flags, the batcher cannot tell where a session begins or ends, and it will hold resources longer than necessary. ## What the model receives The model does not read the correlation ID out of some side channel; Triton **materialises it as an input tensor**. In `config.pbtxt` you declare `control_input` entries mapping tensor names to control kinds: - `CONTROL_SEQUENCE_START` — 1 on the first request of a sequence, 0 otherwise. This is the model's cue to reset state. - `CONTROL_SEQUENCE_END` — 1 on the last request, telling the model it may release state. - `CONTROL_SEQUENCE_READY` — 1 when this batch slot actually contains a request this execution. Because a batch of slots may be partially filled, the model needs to know which rows are real. - `CONTROL_SEQUENCE_CORRID` — the correlation ID itself, for models that key their own state by session. You choose the tensor names and the false/true values (`fp32_false_true`, `int32_false_true`, `bool_false_true`, or `data_type` for CORRID). The backend then treats them like any other input. Alternatively, Triton can carry the state for you: an `state` section inside `sequence_batching` declares input/output state tensor pairs and an `initial_state`, and Triton feeds each execution's output state back as the next execution's input state for that sequence. This **implicit state management** keeps the state on the server side so the backend does not have to own a per-slot store. ## direct versus oldest This is the distinction that separates a candidate who has run a stateful model from one who has read the page. - **`direct`** — all requests of a sequence go to the same instance *and the same batch slot* within it. Slot index is stable for the sequence's whole life. This is what you need when the model or Triton's implicit state keys state by position in the batch. Cost: a slot is reserved for the sequence even while that client is thinking, so a burst of new sequences queues behind idle ones. - **`oldest`** — the batcher routes to an instance but does not reserve a fixed slot. It gathers the oldest candidate sequences and dynamically batches their requests together, with `max_candidate_sequences`, `preferred_batch_size` and `max_queue_delay_microseconds` controlling the batching. Better utilisation, but the model must not assume a stable slot index — it has to key its state off the correlation ID. ## Failure modes worth naming **Slot exhaustion.** Capacity in sequences is instances × slots. With `direct`, an idle-but-open sequence still holds its slot, so a chat-like workload with long think times can exhaust capacity at very low request rates. New sequences then wait. **Clients that never end.** A crashed or disconnected client never sends the end flag. `max_sequence_idle_microseconds` bounds this: if no request for a correlation ID arrives within that window, Triton considers the sequence dead, releases the slot, and errors subsequent requests for that ID. Tune it against your real inter-request think time — too short and legitimate slow clients get evicted mid-session; too long and dead sessions hoard slots. **Correlation ID collisions.** IDs are chosen by clients. Two clients picking the same ID interleave into one sequence and corrupt each other's state. Generate them from something unique per session. **Load balancing.** Sequence affinity is to an instance *inside one server*. Put two Triton replicas behind a round-robin load balancer and request 2 can land on the replica that has never seen the session. Stateful models need sticky routing at the load balancer, keyed on the same session identifier. ## Relationship to the other schedulers A model uses exactly one scheduler. `sequence_batching` replaces `dynamic_batching`; the `oldest` strategy is how you get dynamic-batching-like behaviour back while keeping instance affinity. And note that generative LLM backends do not use this at all — they manage multi-turn context their own way.
- Why does a stateful Triton model behind a round-robin load balancer break?Sequence affinity is enforced only within one server process: the batcher can route a correlation ID to a consistent instance inside its own server, but it has no say over which replica the load balancer picks. The second request can hit a replica that never saw the sequence start, so its state does not exist. You need sticky routing keyed on the session identifier above the servers.
- What does CONTROL_SEQUENCE_READY tell the model that START and END do not?It marks which rows of the executed batch actually carry a request this time. With direct batching, slots are reserved per sequence but not every sequence has a pending request at every execution, so the batch can be partially populated. READY is how the model knows to compute only the live rows and leave the others' state untouched.
- How would you choose max_sequence_idle_microseconds for a conversational workload?Base it on the real distribution of client think time, then add headroom — long enough that a user pausing mid-session is not evicted, short enough that abandoned sessions free their slot before capacity suffers. Measure how many sequences are open versus how many are actually sending, since a large gap under direct batching means idle sequences are the constraint, not compute.
saying these in an interview costs you the question
- Thinking the server stores session state centrally
- Believing dynamic batching can serve stateful models
- Confusing direct and oldest strategies
- Forgetting sticky routing across Triton replicas
- Never sending the sequence end flag from the client