skip to content

TTFT, Throughput and SLOs

You will learn the metrics that actually describe an LLM endpoint — time to first token, inter-token latency, end-to-end p95, and tokens per second — and how tuning for one sacrifices another. Interviewers ask you to set an SLO and defend the batch-size choice that follows from it.

on this pageshow

questions

6

For an LLM server, why is requests/s a misleading capacity metric?

level: middleimportance: must knowfreq 62%

answer

  1. a request is not a fixed unit of work
  2. prefill scales on input, decode on output
  3. two numbers, not one
  4. length distribution must travel with the number
  5. tokens/s is the honest denominator

basics

~20 s

Requests per second hides how much work each request does: one call may emit 40 output tokens, another 2,000. LLM capacity tracks output tokens per second, so report token throughput together with the input and output length distribution it was measured at.

solid answer

~40 s

A request is not a fixed unit of work on an inference server. Prefill cost scales with **input** tokens and decode cost scales with **output** tokens, so two requests/s of 4k-in/1k-out is a completely different load than two requests/s of 200-in/50-out. Quoting requests/s alone lets anyone reproduce your number with shorter prompts and shorter answers. Report **output token throughput** as the primary capacity number, **total token throughput** (input + output) when you want to include prefill work, and always attach the length distribution the run used. Load generators do this for you: `vllm bench serve` prints request throughput, output token throughput and total token throughput side by side, and its `--random-input-len` / `--random-output-len` flags pin the shape of the load. Requests/s only becomes meaningful once that shape is fixed and stated.

go deeper

for a junior

Be able to say that LLM requests vary hugely in size, so counting requests hides the real work. Name output tokens per second as the metric people actually quote for a served model.

for a middle

Explain that prefill cost tracks input tokens and decode cost tracks output tokens, so capacity must be denominated in tokens. Show that you would always publish the input/output length distribution alongside any throughput number.

for a senior

Demonstrate that you sample the real length distribution from production logs before benchmarking, and that you can spot a capacity plan built on short synthetic answers before it becomes an incident.

for a principal

Own the unit the organisation plans in. Decide whether capacity, quotas and chargeback are denominated in tokens or requests, and make sure the benchmark harness, the cost model and the SLO all use the same denominator so numbers compose across teams.

## The unit-of-work problem Every capacity metric assumes requests are roughly interchangeable. That assumption holds for a stateless REST service where a request is a database lookup and a JSON serialization. It fails badly for an LLM server, where one request may be a 200-token prompt with a 30-token answer and the next may be a 32k-token document summarised in 1,500 tokens. Those two differ by two orders of magnitude in GPU work, yet both count as "one request". A capacity claim of "we do 20 requests/s" therefore carries no information until somebody says what a request was. ## Prefill and decode scale on different inputs An LLM request has two phases with different cost drivers. **Prefill** processes the whole prompt in one pass. Its cost grows with the number of **input** tokens (super-linearly once attention over long contexts dominates). Prefill is compute-heavy: it runs large matrix multiplications over many tokens at once and keeps the GPU's math units busy. **Decode** produces output one token at a time, each step reading the full model weights and the accumulated key/value cache. Its cost grows with the number of **output** tokens, and per token it is memory-bandwidth-heavy rather than compute-heavy. So a request's cost is roughly `a x input_tokens + b x output_tokens`, with very different constants. A workload that is all long prompts and one-word answers stresses prefill; a workload that is short prompts and long generations stresses decode. Both can be described as "10 requests/s", and they will saturate the same GPU at wildly different rates. ## What to report instead Three numbers, together: 1. **Output token throughput (tokens/s)** — the headline capacity number. It is what users perceive as "the server is producing text", it is what your decode-phase capacity is spent on, and it is the denominator for cost per million tokens. 2. **Total token throughput (input + output tokens/s)** — includes prefill work. Useful when prompts are long, because output-only throughput understates how loaded the GPU is on a RAG or long-document workload. 3. **Request throughput (requests/s)** — still worth reporting, but only as a derived number, and only next to the length distribution. And one qualifier that is not optional: **the input/output length distribution**. "1,024 in / 256 out, 500 requests, concurrency 32" turns a bare number into a reproducible claim. ## Reading a real benchmark summary vLLM's built-in load generator (`vllm bench serve` in vLLM 0.27, which replaced the older `benchmarks/benchmark_serving.py` script) prints exactly this trio — request throughput, output token throughput, total token throughput — above the TTFT / TPOT / ITL percentile tables. That layout is deliberate: the throughput block tells you how much work the server did, the latency block tells you what each user felt, and neither is interpretable without the other. GenAI-Perf, the load generator now living in the `triton-inference-server/perf_analyzer` repository, reports the same split (output token throughput plus per-request latency statistics). ## Where requests/s is still the right unit Requests/s is fine — and often better — in three places: - **Admission and rate limiting.** Quotas and queue depth are naturally per request, and a client's fair share is easier to express that way (usually alongside a token quota). - **Autoscaling signals**, where queue depth in requests is a direct measure of backlog. - **Fixed-shape workloads**: classification, embedding, or reranking traffic where every request really is the same size. There, requests/s is a legitimate capacity metric because the length distribution is a constant. For open-ended chat or summarisation, none of those conditions hold. ## The failure this causes in practice The common production incident is a capacity plan built from a benchmark whose outputs were short. The team measures 25 requests/s on 128-in/128-out synthetic traffic, provisions for 20 requests/s of real traffic, and then discovers that real answers average 600 tokens. Decode work is roughly five times what was budgeted, the scheduler starts queueing, and time to first token degrades for everyone — because the capacity number was denominated in a unit that did not track the work. The fix is procedural, not clever: measure with the length distribution your production traffic actually has (sample it from real logs), publish capacity in output tokens/s, and treat any requests/s figure without an attached length profile as unverified.

  • Why report total token throughput as well as output token throughput?
    Output token throughput only counts decode work. On a long-prompt workload such as RAG or document summarisation, most of the GPU time goes into prefill, so a server can look under-utilised on output tokens/s while actually being saturated. Total token throughput (input + output) captures that prefill work and keeps long-prompt and short-prompt runs comparable.
  • How would you get a realistic length distribution to benchmark with?
    Sample it from production. Log the input and output token counts your provider or server already reports per request, take the p50/p90/p99 of each, and replay that mix. If you have no traffic yet, use a public conversational dataset with a similar shape rather than fixed-length synthetic prompts, and re-measure once real traffic exists.
  • Does this change how you express cost?
    Yes. Cost per request is unstable for the same reason capacity per request is: it moves with answer length. Price in dollars per million tokens, split input and output, then multiply by your measured length distribution to get a cost per request for a specific feature.

saying these in an interview costs you the question

  • Quoting requests per second with no prompt or output length attached
  • Treating an LLM request as a fixed unit of work like a REST call
  • Assuming input tokens and output tokens cost the same
  • Benchmarking with short answers, then planning capacity for long ones
  • Reporting throughput without the concurrency or arrival rate it was measured at

context

open as a page

Output tokens/s went flat but TTFT p95 keeps climbing — what is happening?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The server is past its saturation knee. Decode throughput is capped by memory bandwidth and KV-cache capacity, so extra concurrency no longer produces more tokens — it only waits in the scheduler queue, and that queue time lands entirely in time to first token.

open as a page

Why does a fixed-length synthetic benchmark overstate an LLM server's capacity?

level: middleimportance: should knowfreq 47%

basics

~10 s

Identical prompts hit the server's prefix cache, uniform lengths remove the long-request tail that clogs real batches, and steady closed-loop pacing never produces bursts. The result is a throughput number production traffic cannot reproduce.

open as a page

How do you load-test a self-hosted LLM endpoint, and what do you report?

level: middleimportance: should knowfreq 56%

basics

~20 s

Drive the server with a token-aware load generator such as vLLM's vllm bench serve or GenAI-Perf, pinning prompt and output lengths and either an arrival rate or a concurrency cap. Report TTFT, per-output-token latency and inter-token latency percentiles alongside output token throughput.

open as a page

What is goodput for an LLM server, and why report it over raw throughput?

level: seniorimportance: should knowfreq 43%

basics

~20 s

Goodput counts only the requests that finished inside the latency SLO — for example TTFT under 500 ms and per-output-token latency under 50 ms. Raw throughput counts a request that took 30 seconds as a success, so it can rise while every user's experience gets worse.

open as a page

What belongs in an SLO for a streaming LLM endpoint, and why not just e2e p95?

level: principalimportance: should knowfreq 44%

basics

~20 s

End-to-end latency scales with answer length, which the model and caller choose, so an e2e p95 target is mostly an SLO on how long answers are. Bound time to first token and per-output-token latency instead, each at a stated load level and length regime.

open as a page