skip to content

Output tokens/s went flat but TTFT p95 keeps climbing — what is happening?

level: seniorimportance: must knowfreq 55%

answer

  1. flat throughput, rising latency
  2. the extra requests are waiting, not running
  3. queue wait lands on first-token time
  4. past the knee, concurrency is not capacity
  5. operating point sits below peak throughput

basics

~20 s

The server is past its saturation knee. Decode throughput is capped by memory bandwidth and KV-cache capacity, so extra concurrency no longer produces more tokens — it only waits in the scheduler queue, and that queue time lands entirely in time to first token.

solid answer

~50 s

That shape is the classic saturation curve of an inference server. Below the knee, adding concurrency fills the batch and both throughput and utilisation rise while latency grows gently. At the knee the GPU is producing tokens as fast as memory bandwidth and available KV-cache blocks allow. Past it, additional requests cannot be admitted to the running batch, so they sit in the waiting queue; queue time is pure additive delay on **TTFT**, which is why TTFT explodes while token throughput is flat. Inter-token latency for already-running requests degrades more gently, because they are still being decoded — until KV pressure forces the scheduler to preempt some of them, at which point throughput actually falls. The operating point is *below* the knee, at the highest load where your TTFT and per-token SLO still hold. Past that, the answer is more replicas, not more concurrency: no amount of queueing creates GPU capacity.

go deeper

for a junior

Know that an inference server has a saturation point, and that beyond it extra requests wait in a queue rather than being served faster. Waiting time shows up as slower first tokens.

for a middle

Explain why queue wait lands on TTFT while inter-token latency degrades more gently, and why flat throughput with rising latency means the offered load exceeds the service rate.

for a senior

Diagnose it on a live server: distinguish a configured concurrency cap from a KV-cache memory limit, spot preemption, and pick an operating point from the SLO rather than from peak throughput.

for a principal

Own the headroom policy. Decide how far below the knee production runs, what the load-shedding behaviour is when demand exceeds capacity, and which signal autoscaling reacts to — knowing utilisation cannot separate healthy from saturated.

## What the curve looks like Sweep concurrency (or arrival rate) against a served model and plot two lines: output tokens per second, and TTFT p95. You get a shape that recurs on every engine. - **Under-loaded region.** Few requests in flight. The GPU spends decode steps reading the entire model weights to produce a handful of tokens, so per-token cost is dominated by weight loading and hardware is badly under-used. Throughput climbs nearly linearly with concurrency; latency barely moves. - **The knee.** The batch is large enough that the fixed cost of a decode step is amortised across many sequences. Throughput growth flattens. Latency starts to rise noticeably. - **Saturated region.** Throughput is flat — the ceiling is set by memory bandwidth per decode step and by how many sequences fit in KV cache. Additional arrivals queue. TTFT rises roughly linearly with the backlog. - **Collapse.** If KV-cache capacity binds hard, the scheduler starts preempting running sequences (recomputing or swapping their cache), and throughput can *fall* while latency continues to climb. ## Why the queue lands on TTFT specifically TTFT is the interval between the request arriving and its first token appearing. It decomposes into: network, admission/queue wait, and the prefill forward pass. Once the server is saturated, queue wait dominates and it grows with backlog. Every second a request spends waiting for a batch slot is a second added to TTFT and to nothing else. Inter-token latency behaves differently. A request that has been admitted is being decoded alongside its batch mates; it degrades as the batch grows, but it degrades smoothly. This asymmetry is the diagnostic signal: **TTFT climbing steeply while inter-token latency rises mildly and throughput is flat means a queueing problem, not a decode-speed problem.** When inter-token latency *also* jumps and throughput drops, you are into preemption or memory pressure territory, which is a different and worse failure. ## Confirming the diagnosis on a live server - Watch queue depth / number of waiting requests. Growing monotonically under steady arrivals means arrivals exceed service rate — the queue will never drain on its own. - Watch running-batch size. If it is pinned at its cap while requests wait, you are concurrency-limited by configuration. If it is *below* the cap while requests still wait, you are memory-limited: there are no free KV-cache blocks to admit anyone. - Watch preemption counters. Non-zero preemption under steady load means the engine is thrashing. - Check GPU utilisation. Near 100% with flat throughput confirms you have found the hardware ceiling rather than a software stall. Those two sub-cases matter because the remedies differ: a configuration limit can sometimes be raised; a memory limit means you need more KV cache, a smaller context budget, a smaller model footprint, or another GPU. ## Little's law, applied In steady state, `average in-flight requests = arrival rate x average time in system`. If throughput is capped at T requests/s and you offer more than T, the excess accumulates: time in system rises without bound and there is no equilibrium. This is the formal reason "just raise concurrency" is not a fix. Concurrency is not a capacity knob past the knee; it is a decision about *where the waiting happens* — in your queue, or in the client's connection pool. ## Choosing the operating point Do not run at the knee. The knee is peak throughput, which means peak latency-per-token and zero headroom: a traffic burst puts you straight into the collapse region. Instead: 1. Define the latency SLO first (TTFT p95, per-token p95). 2. From the sweep, find the highest load at which those percentiles still hold — that is the SLO-bounded capacity, usually meaningfully below peak throughput. 3. Provision replicas so expected peak traffic sits below that point, keeping headroom for bursts and for a replica being unavailable during a rollout. 4. Autoscale on a signal that leads the cliff — queue depth or TTFT — rather than on GPU utilisation, which is high across both the healthy and the saturated regions and therefore cannot distinguish them. ## The wrong answers "Increase the batch size" — past the knee the batch is already as large as memory allows, and forcing it larger buys preemption. "Scale up the GPU" — helps only if you were memory-capacity-bound, and it is slower and often costlier than another replica. "Raise the client timeout" — hides the symptom while the backlog keeps growing. The correct response to demand above capacity is more capacity, or shedding load with an explicit rejection so clients fail fast instead of waiting behind an unbounded queue.

  • How do you tell a scheduler-configuration limit from a KV-cache memory limit?
    Compare the running batch size against its configured cap while requests are waiting. Pinned at the cap with memory to spare means the configured concurrency limit is binding and can be raised. Below the cap with requests still queued means there were no free KV-cache blocks to admit them — a memory limit, fixed only by more cache room (shorter contexts, smaller footprint, more GPU) or more replicas.
  • Why is GPU utilisation a poor autoscaling signal for this?
    Utilisation is already near 100% just below the knee, where latency is fine, and stays near 100% deep into saturation, where it is not. It cannot distinguish healthy from overloaded. Queue depth or TTFT p95 both move sharply exactly at the transition, so they lead the cliff and make usable scaling triggers.
  • What should the server do with requests it cannot serve within the SLO?
    Shed them explicitly. An unbounded queue converts an overload into a latency disaster for every waiting client, including ones that would otherwise have been served fine. A bounded queue with a fast rejection lets callers retry elsewhere, fall back to a smaller model, or surface a clear error, and it keeps the admitted requests inside their SLO.
  • Why can throughput actually decrease past a certain load?
    Because the scheduler starts preempting running sequences when KV-cache blocks run out. A preempted sequence's cached state is either recomputed later or swapped out and back, so the GPU spends cycles redoing work it had already done. That is wasted capacity, and it shows up as throughput falling while latency keeps climbing.

saying these in an interview costs you the question

  • Reading flat throughput as the server being idle or under-used
  • Adding concurrency to fix latency past the saturation point
  • Running production at the peak-throughput point with no headroom
  • Autoscaling on GPU utilisation, which is high on both sides of the knee
  • Raising client timeouts instead of adding capacity or shedding load

context