skip to content

For a streamed LLM endpoint, which latency numbers do you track besides total time?

level: middleimportance: must knowfreq 58%

answer

  1. decompose, do not average
  2. server first byte is not what the user sees
  3. think time is its own line
  4. mind the gaps, not the mean
  5. p50 and p99, measured on the client

basics

~20 s

Track three separately: time to the first token the user can actually see, inter-token latency once the stream is flowing, and total time to the last token. Measure the first one on the client, at p50 and p99.

solid answer

~50 s

Total time is the weakest of the numbers because it hides where the wait went. The headline metric is **time to first visible token** — the moment content appears in front of the user. That is deliberately different from server-side time to first byte, because with a reasoning model the stream can open immediately and then emit nothing readable while the model thinks, and because a buffering hop can hold bytes the server already flushed. The second number is **inter-token latency** (or its inverse, tokens per second) — once text is flowing, the gaps between chunks decide whether it reads smoothly or judders; track the tail, not the mean, since a single two-second stall is what users report. The third is time to last token, which matters for cost and capacity but not much for how the feature feels. Report each at p50 and p99, and instrument on the client.

go deeper

for a junior

Know that streaming has more than one latency number: how long until text appears, how fast it flows after that, and how long the whole answer takes.

for a middle

Be able to name and define the decomposition — think time, time to first visible token, inter-token latency, total — and explain why the first visible token is the one users judge you on.

for a senior

Show how you would instrument this in production: client-side timestamps, p50/p99 reporting, and using the gap between server and client first-token numbers as a diagnostic for buffering.

for a principal

Own the targets. Decide which percentile of which metric the team is accountable for, how the reasoning-time budget trades against answer quality, and when further speed stops being worth engineering effort.

## Why one number is not enough For a buffered request, total latency is the whole story: the user waits, then gets the answer. Streaming splits that single wait into phases with different causes, different owners and different fixes, and a single average hides all of it. Two endpoints can both average nine seconds, one showing text after 500 ms and streaming smoothly, the other sitting blank for eight seconds and then dumping everything. Only a decomposition tells them apart. ## The four numbers **Think time.** With a reasoning model, the run begins with an internal reasoning phase before any answer token exists. This is real elapsed time during which the response is open but nothing readable is coming. It can be seconds to tens of seconds, and it scales with how much reasoning effort the request asks for. As of mid-2026 this is a large, separately-budgeted component rather than a rounding error, so it deserves its own line. **Time to first token (TTFT).** The classic serving metric: how long until the first token of output is produced. Measured at the server or the provider boundary. **Time to first *visible* token (TTFVT).** How long until something appears in front of the user. This is the metric that governs whether the feature feels alive, and it is not the same as TTFT for two reasons. First, with extended thinking the first produced token may be a reasoning token the user never sees, so the answer's first visible token comes much later — unless you deliberately surface a thinking summary, which resets TTFVT to something small. Second, everything between your server and the user's screen can add delay: a buffering proxy, a compression layer, a client that waits for a complete parse before rendering. TTFT is a server number; TTFVT is a user number, and only client-side instrumentation can measure it. **Inter-token latency (ITL).** Once the stream is running, the gap between successive chunks — usually reported as its inverse, output tokens per second. This decides smoothness. The key discipline is to look at the distribution, not the mean: a stream averaging 40 tokens/s that contains one 2.5-second stall reads as broken, and the mean will not show it. Track the p99 gap, or simply count stalls over a threshold. **Time to last token / total.** Still worth tracking, because it drives connection lifetime, capacity, and the cost side of the ledger. But it is the wrong headline for a streamed UI, since the user has usually been reading for most of it. ## Where to measure Instrument on the client wherever the metric is about perception. A server-side histogram cannot see a proxy that buffered your flushes, a slow network, or a renderer that batches DOM updates. The usual shape is: the client records timestamps for request start, first visible chunk, each chunk arrival, and stream end, then reports a compact summary; the server records its own TTFT and total for capacity work. When the two disagree, the gap between them is your infrastructure problem, and that gap is the single most useful diagnostic this decomposition gives you. ## Percentiles, not averages Every number here should be reported at p50 and p99. Streaming latency distributions are heavy-tailed — a queue at the provider, a cold path, or an unusually long reasoning phase produces occasional very slow runs, and the mean absorbs them silently. A p99 TTFVT of 20 seconds against a p50 of 900 ms says something a nine-second average never will. ## Turning the numbers into targets The decomposition maps to different levers. A bad TTFVT with a good TTFT is a delivery problem: buffering somewhere in the path, or a client that renders too late. A bad TTFVT with a matching bad TTFT on a reasoning model is a think-time problem: the fix is either to surface progress during the thinking phase or to ask for less reasoning. A good TTFVT with bad inter-token latency is a serving or network problem: contention, throttling, or chunk framing overhead. Bad total time with everything else healthy usually means the answer is simply long, which is a prompt or output-budget question, not a latency bug. One more calibration point: past a certain output rate, faster stops helping. People read a few words a second, so once the model comfortably outruns reading speed, additional tokens per second buys nothing perceptible, while first-visible-token improvements keep paying all the way down.

  • Your server reports a 400 ms time to first token but users report a four-second blank screen. Where do you look?
    The gap is entirely between your flush and their screen. Check for buffering in the path — a reverse proxy or CDN accumulating the body, a compression layer waiting to fill a block — and then check the client: is it awaiting a complete parse, or batching renders before painting? Reproduce by hitting the app directly with an unbuffered client and comparing chunk arrival timestamps hop by hop.
  • Why report inter-token latency at the tail rather than as an average?
    Because smoothness is destroyed by individual stalls, not by the mean rate. A stream averaging 40 tokens per second that pauses once for two seconds reads as frozen, yet the average looks excellent. Tracking the p99 inter-chunk gap, or simply counting gaps over a threshold like one second, surfaces the thing users actually complain about.
  • Once a stream is flowing faster than the user reads, is there value in increasing tokens per second further?
    Very little perceptually. Human reading runs at a few words per second, so once output comfortably outpaces it the extra speed is invisible and only shortens the tail of the response. The remaining payoff is operational — shorter connections, freed serving capacity — so further effort is usually better spent on time to first visible token.

saying these in an interview costs you the question

  • Reports one average latency number for a streamed endpoint
  • Treats server time to first byte as the user's wait
  • Measures streaming latency only on the server
  • Uses mean inter-token latency and misses multi-second stalls
  • Ignores reasoning time as a separate component of the wait

context