For an LLM server, why is requests/s a misleading capacity metric?
answer
- a request is not a fixed unit of work
- prefill scales on input, decode on output
- two numbers, not one
- length distribution must travel with the number
- tokens/s is the honest denominator
basics
~20 sRequests per second hides how much work each request does: one call may emit 40 output tokens, another 2,000. LLM capacity tracks output tokens per second, so report token throughput together with the input and output length distribution it was measured at.
solid answer
~40 sA request is not a fixed unit of work on an inference server. Prefill cost scales with **input** tokens and decode cost scales with **output** tokens, so two requests/s of 4k-in/1k-out is a completely different load than two requests/s of 200-in/50-out. Quoting requests/s alone lets anyone reproduce your number with shorter prompts and shorter answers. Report **output token throughput** as the primary capacity number, **total token throughput** (input + output) when you want to include prefill work, and always attach the length distribution the run used. Load generators do this for you: `vllm bench serve` prints request throughput, output token throughput and total token throughput side by side, and its `--random-input-len` / `--random-output-len` flags pin the shape of the load. Requests/s only becomes meaningful once that shape is fixed and stated.
go deeper
Be able to say that LLM requests vary hugely in size, so counting requests hides the real work. Name output tokens per second as the metric people actually quote for a served model.
Explain that prefill cost tracks input tokens and decode cost tracks output tokens, so capacity must be denominated in tokens. Show that you would always publish the input/output length distribution alongside any throughput number.
Demonstrate that you sample the real length distribution from production logs before benchmarking, and that you can spot a capacity plan built on short synthetic answers before it becomes an incident.
Own the unit the organisation plans in. Decide whether capacity, quotas and chargeback are denominated in tokens or requests, and make sure the benchmark harness, the cost model and the SLO all use the same denominator so numbers compose across teams.
## The unit-of-work problem Every capacity metric assumes requests are roughly interchangeable. That assumption holds for a stateless REST service where a request is a database lookup and a JSON serialization. It fails badly for an LLM server, where one request may be a 200-token prompt with a 30-token answer and the next may be a 32k-token document summarised in 1,500 tokens. Those two differ by two orders of magnitude in GPU work, yet both count as "one request". A capacity claim of "we do 20 requests/s" therefore carries no information until somebody says what a request was. ## Prefill and decode scale on different inputs An LLM request has two phases with different cost drivers. **Prefill** processes the whole prompt in one pass. Its cost grows with the number of **input** tokens (super-linearly once attention over long contexts dominates). Prefill is compute-heavy: it runs large matrix multiplications over many tokens at once and keeps the GPU's math units busy. **Decode** produces output one token at a time, each step reading the full model weights and the accumulated key/value cache. Its cost grows with the number of **output** tokens, and per token it is memory-bandwidth-heavy rather than compute-heavy. So a request's cost is roughly `a x input_tokens + b x output_tokens`, with very different constants. A workload that is all long prompts and one-word answers stresses prefill; a workload that is short prompts and long generations stresses decode. Both can be described as "10 requests/s", and they will saturate the same GPU at wildly different rates. ## What to report instead Three numbers, together: 1. **Output token throughput (tokens/s)** — the headline capacity number. It is what users perceive as "the server is producing text", it is what your decode-phase capacity is spent on, and it is the denominator for cost per million tokens. 2. **Total token throughput (input + output tokens/s)** — includes prefill work. Useful when prompts are long, because output-only throughput understates how loaded the GPU is on a RAG or long-document workload. 3. **Request throughput (requests/s)** — still worth reporting, but only as a derived number, and only next to the length distribution. And one qualifier that is not optional: **the input/output length distribution**. "1,024 in / 256 out, 500 requests, concurrency 32" turns a bare number into a reproducible claim. ## Reading a real benchmark summary vLLM's built-in load generator (`vllm bench serve` in vLLM 0.27, which replaced the older `benchmarks/benchmark_serving.py` script) prints exactly this trio — request throughput, output token throughput, total token throughput — above the TTFT / TPOT / ITL percentile tables. That layout is deliberate: the throughput block tells you how much work the server did, the latency block tells you what each user felt, and neither is interpretable without the other. GenAI-Perf, the load generator now living in the `triton-inference-server/perf_analyzer` repository, reports the same split (output token throughput plus per-request latency statistics). ## Where requests/s is still the right unit Requests/s is fine — and often better — in three places: - **Admission and rate limiting.** Quotas and queue depth are naturally per request, and a client's fair share is easier to express that way (usually alongside a token quota). - **Autoscaling signals**, where queue depth in requests is a direct measure of backlog. - **Fixed-shape workloads**: classification, embedding, or reranking traffic where every request really is the same size. There, requests/s is a legitimate capacity metric because the length distribution is a constant. For open-ended chat or summarisation, none of those conditions hold. ## The failure this causes in practice The common production incident is a capacity plan built from a benchmark whose outputs were short. The team measures 25 requests/s on 128-in/128-out synthetic traffic, provisions for 20 requests/s of real traffic, and then discovers that real answers average 600 tokens. Decode work is roughly five times what was budgeted, the scheduler starts queueing, and time to first token degrades for everyone — because the capacity number was denominated in a unit that did not track the work. The fix is procedural, not clever: measure with the length distribution your production traffic actually has (sample it from real logs), publish capacity in output tokens/s, and treat any requests/s figure without an attached length profile as unverified.
- Why report total token throughput as well as output token throughput?Output token throughput only counts decode work. On a long-prompt workload such as RAG or document summarisation, most of the GPU time goes into prefill, so a server can look under-utilised on output tokens/s while actually being saturated. Total token throughput (input + output) captures that prefill work and keeps long-prompt and short-prompt runs comparable.
- How would you get a realistic length distribution to benchmark with?Sample it from production. Log the input and output token counts your provider or server already reports per request, take the p50/p90/p99 of each, and replay that mix. If you have no traffic yet, use a public conversational dataset with a similar shape rather than fixed-length synthetic prompts, and re-measure once real traffic exists.
- Does this change how you express cost?Yes. Cost per request is unstable for the same reason capacity per request is: it moves with answer length. Price in dollars per million tokens, split input and output, then multiply by your measured length distribution to get a cost per request for a specific feature.
saying these in an interview costs you the question
- Quoting requests per second with no prompt or output length attached
- Treating an LLM request as a fixed unit of work like a REST call
- Assuming input tokens and output tokens cost the same
- Benchmarking with short answers, then planning capacity for long ones
- Reporting throughput without the concurrency or arrival rate it was measured at