skip to content

How does TGI's --waiting-served-ratio decide when to pause decoding for queued requests?

level: seniorimportance: should knowfreq 34%

answer

  1. A ratio, not an absolute queue length
  2. Queued versus running requests
  3. Interrupting decode has a victim
  4. Second flag is the anti-starvation backstop
  5. Steps convert to milliseconds of TTFT

basics

~20 s

TGI's --waiting-served-ratio is the ratio of waiting requests to running requests at which the router will interrupt decoding, prefill the queued requests, and fold them into the running batch. A lower value admits newcomers sooner and protects their time-to-first-token.

solid answer

~50 s

Once TGI has a batch decoding, new arrivals face a scheduling question: interrupt the running batch to prefill them, or let the current requests keep producing tokens? Two launcher flags answer it. `--waiting-served-ratio` compares queued requests to running ones. When that ratio crosses the threshold (0.3 in current TGI 3.x releases), the router considers pausing decode to prefill the waiting set and concatenate it into the running batch. `--max-waiting-tokens` (default 20) is the backstop: after that many decode steps, waiting requests are admitted regardless of ratio, so a queued request cannot starve behind a long generation. The tradeoff is symmetric. Admitting eagerly protects newcomers' TTFT but interrupts everyone already streaming, adding a visible stall to their inter-token latency. Admitting lazily keeps existing streams smooth and raises decode efficiency, but queue time - and therefore p95 TTFT - climbs.

go deeper

for a junior

Know that a TGI server keeps a queue and that new requests are not always added to the batch immediately. You are not expected to name the flags at this level.

for a middle

Explain that admitting a waiting request requires pausing decode to run a prefill, and that these flags decide when that pause is taken. Say which side each direction favours.

for a senior

Translate --max-waiting-tokens into a worst-case TTFT contribution in milliseconds, and describe using TGI's queue-duration versus inference-duration metrics to decide which way to move.

for a principal

Own the position that one scheduler cannot serve two SLOs: argue for separate deployments for interactive and batch traffic, and connect the admission policy to goodput and cost per served request rather than raw tokens per second.

## The scheduling decision this flag answers TGI's router keeps a running batch that is producing one token per sequence per step, plus a queue of requests that have arrived but not yet been prefilled. Prefill and decode cannot run in the same forward pass on the same batch, so admitting a newcomer means stopping decode, running a prefill, and concatenating the result into the batch. That pause is paid by every request already streaming. So the router needs a rule for when the pause is worth it. TGI's rule has two parts. ## --waiting-served-ratio This is a ratio of *queued* requests to *running* requests. If it is 0.3 and eight requests are decoding, the router starts considering an admission pause once roughly three are waiting. The intuition: when the queue is small relative to the batch, the newcomers are a minor share of demand and the interruption cost outweighs their gain; when the queue grows relative to the batch, those waiting requests represent enough of your traffic that leaving them queued is the bigger latency problem. The ratio form is deliberate. An absolute queue threshold would behave completely differently on a server running batches of four versus batches of sixty-four. A ratio scales with whatever batch size the memory budget currently supports. Note also that admission is still bounded by memory: there must be room in the batch under `--max-batch-total-tokens` for the new sequences' cache. Crossing the ratio expresses intent, not permission. ## --max-waiting-tokens Ratio alone has a starvation hole. Imagine one long generation running alone and a single request queued behind it: one waiting over one running is above the threshold, but if the batch were larger the same lone request would sit indefinitely. `--max-waiting-tokens` (default 20) forces the issue - after that many decode steps have elapsed, waiting requests get admitted if the batch has room. The useful way to read this number is in wall-clock terms. If a decode step takes about 20 ms, 20 steps is roughly 0.4 seconds, which is the worst-case extra queue delay this backstop imposes. That converts directly into a TTFT budget line: if your SLO says first token within one second and your model decodes at 25 ms per step, a `--max-waiting-tokens` of 20 spends half a second of that budget in the worst case. ## Tuning by workload shape - **Interactive chat, TTFT-sensitive**: lower both. Admit early and often. You are deliberately paying inter-token jitter for existing streams to keep new users from staring at an empty box. Chunked prefill in recent TGI releases makes that stall shorter than it used to be. - **Batch/offline scoring**: raise them. Nobody is watching a stream, so long uninterrupted decode runs maximise tokens per second per GPU. Queue delay is irrelevant when the job is measured in total wall clock. - **Mixed traffic on one endpoint**: this is where candidates usually get stuck, and the honest senior answer is that a single scheduler cannot satisfy both SLOs. Split into separate deployments with separate settings rather than searching for a ratio that pleases everyone. ## How you observe the effect Don't tune these blind. TGI's Prometheus endpoint exposes queue-time and inference-time series separately, so you can see which side of the tradeoff you are on: if queue duration dominates end-to-end latency, admission is too lazy; if queue duration is near zero but inter-token latency is spiky, you are interrupting decode too often. Change one flag at a time and re-measure under a fixed load profile - these settings interact with arrival patterns, so an A/B under synthetic uniform load can mislead you about bursty production traffic. ## What interviewers listen for The signal is whether you can state the tradeoff in both directions and name who pays. Weak answers describe the flag as 'batching efficiency' with no mention that the cost lands on requests that are already streaming. Strong answers translate `--max-waiting-tokens` into milliseconds against a TTFT SLO and admit that mixed workloads want separate deployments.

  • Why express the threshold as a ratio rather than an absolute number of queued requests?
    Because the meaningful question is how much of current demand is unserved, and that depends on batch size. Three queued requests behind a batch of four is a backlog; three behind a batch of sixty-four is noise. A ratio scales automatically as the memory budget lets the running batch grow or shrink, so one setting behaves sensibly across model sizes and GPUs.
  • Your p95 TTFT is bad but inter-token latency is excellent. Which direction do you move these flags?
    Toward eager admission: lower `--waiting-served-ratio` and lower `--max-waiting-tokens`. That profile says requests are sitting in the queue while the running batch decodes undisturbed. Confirm first with TGI's queue-duration metric versus inference-duration - if queue time is not the dominant term, the fix is capacity or prefill batching, not the admission policy.
  • How would you serve interactive chat and an overnight scoring job on the same model?
    Two deployments of the same weights with different scheduler settings: chat gets eager admission and tight TTFT-oriented flags, the batch job gets lazy admission and long uninterrupted decode runs. Sharing one endpoint means one scheduling policy, and whichever way you tune it, one workload's SLO loses. Separate replicas also isolate a traffic spike in one from harming the other.

saying these in an interview costs you the question

  • Thinking it counts queued requests absolutely
  • Ignoring that admission stalls existing streams
  • Treating a lower ratio as strictly better
  • Assuming crossing the ratio guarantees admission regardless of memory
  • Never converting max-waiting-tokens into milliseconds

context