skip to content

Latency, Throughput and Batching

Inference economics come down to keeping an expensive accelerator busy without making any single user wait. Be able to name the metrics, say why decode is bandwidth-bound rather than compute-bound, and defend where you sit on the latency-throughput curve.

on this pageshow

questions

4

In LLM serving, what do time-to-first-token and inter-token latency each measure?

level: juniorimportance: must knowfreq 78%

answer

  1. Two clocks, not one
  2. First token versus every token after
  3. Prompt length moves only one of them
  4. TTFT plus rate times output tokens
  5. Streaming changes perception, not generation time

basics

~20 s

Time-to-first-token is the wait from sending a request until the first output token arrives, and it grows with prompt length. Inter-token latency is the gap between successive tokens after that, and it sets how fast the answer streams.

solid answer

~40 s

Three separate clocks matter. **Time-to-first-token (TTFT)** covers queueing plus processing the whole prompt before any output exists, so it scales with prompt length and with server load. **Inter-token latency (ITL)**, also called time-per-output-token, is the steady-state gap between output tokens; its reciprocal is the per-user streaming rate, and it barely depends on prompt length. **End-to-end completion time** is roughly TTFT plus ITL times the number of output tokens. They trade against each other and against server throughput, so a real SLO names them separately — for a consumer chat product, something like a 400 ms p95 TTFT and a 30 tokens-per-second floor, rather than one blended latency number that hides which stage regressed.

go deeper

for a junior

Know the two names and what each measures: time-to-first-token is the wait before anything appears, inter-token latency is the gap between tokens once it starts. Be able to say which one a longer prompt affects.

for a middle

Explain why the two respond to different levers, why both degrade under load, and how completion time is built from TTFT, streaming rate and output length. Say plainly that streaming changes perception, not generation time.

for a senior

Show you would define these as separate tail objectives per route class, measured at the client and sliced by prompt length, and that you can diagnose a TTFT tail by separating queue wait from prompt processing.

for a principal

Own the argument that latency targets are product decisions with a capacity price attached, and that different traffic classes deserve different objectives rather than one fleet-wide number that overprovisions for everyone.

## The three clocks Every generation request has three latency numbers, and conflating them is the most common source of useless service-level objectives. **Time-to-first-token (TTFT)** is the wall-clock gap between sending the request and receiving the first output token. It contains two things: how long the request waited in the server's queue, and the forward pass over the entire prompt that must finish before any token can be emitted. Because that work grows with the number of input tokens, TTFT grows with prompt length — a 200-token prompt and a 100,000-token prompt do not have the same first-token wait on the same hardware. **Inter-token latency (ITL)**, often called time-per-output-token (TPOT), is the steady-state gap between consecutive output tokens once generation is under way. Its reciprocal is the per-user streaming rate in tokens per second. For a given model, hardware and server load it is roughly constant, and it is largely independent of how long the prompt was. **End-to-end completion time** is what a non-streaming caller experiences: approximately TTFT + ITL x (number of output tokens). Note the third term — output length is a latency variable. A prompt change that makes the model twice as verbose doubles the tail of every response even though nothing about the server changed. ## Why the two front metrics move independently The first token and every token after it are produced by work with different shapes. Producing the first token requires consuming the whole prompt; producing token N+1 requires one more step on top of state the server already holds. That is why the two numbers respond to different levers. Shortening the prompt, or splitting a giant document into a retrieval step, moves TTFT and leaves ITL alone. Choosing a smaller or faster-decoding model moves ITL. Nothing about a shorter prompt makes tokens stream faster once streaming has begun. Load couples them from the other direction. More concurrent traffic means requests wait longer before they start, which inflates TTFT, and means each generation step is shared with more sequences, which inflates ITL. This is why both numbers must be quoted at a stated load, and why a benchmark run against an idle server tells you almost nothing about production. ## Perceived versus measured latency Streaming changes none of these numbers — the model does not generate faster because you show tokens as they arrive — but it changes the experience completely. With streaming, perceived latency collapses to TTFT plus whether tokens outrun reading speed. Comfortable reading is on the order of a handful of tokens per second, so a stream at 30 tokens per second feels instantaneous after the first token appears; the user is never waiting on the model. Without streaming, the same request is a 12-second blank screen, which reads as a hang even though the total generation time is identical. This is the first-token illusion, and it cuts both ways. A regression from 400 ms to 2 s TTFT makes a chat product feel broken while total completion time is unchanged, so a team watching only end-to-end time will see a flat graph and a support queue full of complaints. ## What belongs in the SLO Write them as separate objectives, at tails, per route class: - p95 (and p99) TTFT, because the tail is what users remember and the mean hides queueing spikes entirely. - p95 ITL, or equivalently a minimum tokens-per-second floor per stream. - A completion-time budget only where a caller genuinely blocks on the full answer — an API integration, a synchronous backend step, a non-streamed job. Measure at the client, not inside the server, so network and any intermediary are included. Slice TTFT by prompt-length bucket, because one route with enormous prompts will otherwise drag the whole distribution and look like a general slowdown. And keep the interactive and non-interactive routes on separate objectives: a background summarization path has no meaningful TTFT requirement at all, and holding it to the chat target wastes capacity. ## Common mistakes Reporting a single mean latency number is the classic one — it is the average of a bimodal mixture of short and long prompts, and it moves for reasons nobody can attribute. Assuming faster hardware fixes ITL is another; per-token generation speed is bounded by how fast weights can be read from memory, not by raw arithmetic throughput. And treating output length as free ignores that it multiplies straight into completion time.

  • How does a much longer prompt change TTFT and inter-token latency differently?
    TTFT grows roughly with prompt length, because the whole prompt must be processed before any token can be emitted. Inter-token latency grows only slightly, since each later step attends over a longer history but performs the same weight reads. So a long-document request is slow to start and then streams at close to normal speed.
  • Your p50 TTFT is 300 ms but p99 is 6 seconds. Where do you look first?
    Split TTFT into queue wait and prompt-processing time in the trace. A tail driven by queue wait points at bursty arrivals or too little headroom; a tail driven by processing points at a route with very large prompts sharing the same pool. Then slice by prompt-length bucket and by route before touching any tuning knob.
  • Why can end-to-end completion time stay flat while users report the product got slower?
    Because a streamed product is judged on TTFT and streaming rate, not total time. A shift from 400 ms to 2 s before the first token feels broken even when the answer finishes at the same moment. The reverse also happens: a prompt change that makes answers longer degrades completion time while TTFT looks perfect.

saying these in an interview costs you the question

  • Reports one latency number without separating first token from the rest
  • Thinks streaming makes the model generate faster
  • Uses mean latency instead of p95 or p99 for the SLO
  • Assumes TTFT is unaffected by prompt length
  • Believes inter-token latency stays fixed regardless of server load

context

open as a page

Why is single-stream LLM decoding limited by memory bandwidth rather than FLOPs?

level: middleimportance: must knowfreq 60%

basics

~20 s

Generating one token for one request reads every parameter that token's forward pass needs out of GPU memory but does only a couple of arithmetic operations per parameter. The accelerator therefore sits waiting on memory, so tokens per second track memory bandwidth, not peak FLOPs.

open as a page

Raising server batch size from 8 to 64 triples throughput but slows each user — how do you choose?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Fix the user-facing latency target first, then sweep concurrency and pick the largest batch size whose p95 first-token and per-token latencies still meet it. Optimise for throughput that satisfies the target, not raw tokens per second, and re-check the tail rather than the mean.

open as a page

How do you serve one LLM for interactive chat and an overnight 900K-listing scoring job?

level: principalimportance: should knowfreq 40%

basics

~20 s

Treat them as two deployments of the same model tuned to opposite ends of the latency-throughput frontier: the chat path optimised for tail first-token latency at modest concurrency, the scoring job optimised for tokens per accelerator-hour at maximum concurrency, with the bulk work scheduled into off-peak capacity.

open as a page