What is goodput for an LLM server, and why report it over raw throughput?
answer
- throughput minus the requests that missed
- two conditions, both must hold
- the flat line that hides an incident
- peaks below the raw throughput peak
- borrowed from networking
basics
~20 sGoodput counts only the requests that finished inside the latency SLO — for example TTFT under 500 ms and per-output-token latency under 50 ms. Raw throughput counts a request that took 30 seconds as a success, so it can rise while every user's experience gets worse.
solid answer
~50 sRaw throughput answers "how many tokens did the GPU emit?" Goodput answers "how many of those went to users who got an acceptable experience?" You define an SLO — typically a TTFT threshold plus a per-output-token or inter-token latency threshold — and then measure the request or token rate that satisfies *both* conditions. The two curves separate exactly where it matters: as you push load past the saturation knee, raw throughput stays flat or dips slightly while goodput falls off a cliff, because a growing share of requests now breach the TTFT bound while waiting in the queue. That makes goodput the right capacity-planning number. Provision replicas so peak traffic sits below the load where goodput peaks, not where raw throughput peaks. It also cleanly settles configuration arguments: a setting that raises tokens/s while pushing a third of requests past their TTFT bound is a regression, and only a goodput measurement shows that.
go deeper
Know that goodput counts only requests that met the latency target, so it is throughput with the too-slow ones removed. Say why a plain token count can look healthy during a latency incident.
State the SLO as a conjunction of a first-token bound and a per-output-token bound, and explain why goodput and raw throughput diverge past the saturation point.
Show that you measure it with an open-loop sweep, compute it from exported per-request data rather than assuming a flag exists, and use its peak instead of peak throughput for capacity.
Own the thresholds and the traffic classes they apply to. Decide which SLO each workload gets, tie the error budget to goodput, and make it the metric that settles tuning and vendor-comparison arguments across teams.
## The definition **Goodput** is throughput restricted to requests that met their service-level objective. On an LLM server the SLO is usually a conjunction: > a request counts if `TTFT <= A` **and** `per_output_token_latency <= B` where per-output-token latency is the request's generation time divided by its output token count (some teams use the inter-token latency percentile within the request instead). Both conditions must hold. Goodput is then the rate of such requests per second, or, if you prefer a token-denominated number, the output tokens per second belonging to them. The term is borrowed from networking, where goodput excludes retransmitted and dropped packets. The intuition transfers directly: work that did not usefully arrive is not capacity. ## Why raw throughput lies at exactly the wrong moment Throughput and goodput agree while the server is healthy — every request meets the SLO, so all output counts. They diverge past saturation: - Raw output tokens/s plateaus. The GPU is still busy, still emitting tokens at its hardware ceiling. - Meanwhile queue time grows. Requests sit waiting, breach the TTFT bound, and their tokens stop counting toward goodput. - Goodput therefore *falls* while throughput is flat. So a dashboard showing token throughput alone reports a perfectly steady green line through an outage-grade latency incident. That is the whole argument for goodput as a headline metric. ## Why a conjunction, not one number Using only TTFT rewards a server that admits everyone instantly and then streams at a crawl — first tokens are fast, the rest of the answer takes a minute. Using only per-token latency rewards the opposite: a server that queues aggressively so admitted requests decode in a small, fast batch, while everyone else waits. Only requiring both closes both loopholes. It also mirrors what a user perceives: a fast start *and* a readable streaming pace. ## Measuring it Run an open-loop sweep: for each arrival rate, record every request's TTFT and per-output-token latency, apply the SLO predicate, and count the passing rate. Plot goodput against offered load. The curve rises with offered load, peaks, and then declines — the peak is your SLO-bounded capacity, and it sits meaningfully below the raw-throughput peak. Standard load generators give you the raw material even when they do not compute goodput directly: `vllm bench serve` reports TTFT, TPOT and ITL percentile tables plus throughput, and can write per-request results out with `--save-result` for you to post-process. GenAI-Perf similarly exports per-request records. Computing goodput is a filter over that export, not a separate experiment. ## Where it changes decisions **Capacity planning.** Provision against peak goodput. Sizing to the raw-throughput ceiling guarantees you operate in the region where a large fraction of users are outside SLO. **Configuration tuning.** Almost every serving knob trades first-token latency against aggregate token throughput. Larger batches, more aggressive admission, and bigger prefill chunks all raise tokens/s and can push TTFT up. Evaluated on throughput alone every one of them looks like a win. Evaluated on goodput, the ones that broke the SLO show up as losses. This is the single most valuable use of the metric. **Autoscaling targets.** Setting a scale-out trigger at the load where goodput starts declining is defensible in a way that a CPU- or GPU-utilisation target is not. **Multiple traffic classes.** Interactive chat and background batch jobs deserve different SLOs, so they deserve different goodput definitions. A batch scoring job with no human waiting has effectively no TTFT bound, and measuring it against a chat SLO produces a meaningless number. ## Limits and honest caveats Goodput is only as good as the SLO you chose; a loose threshold makes any server look fine. It is a step function, so a request that misses by one millisecond counts the same as one that missed by ten seconds — pair it with the latency percentile distribution rather than replacing that distribution. And it is not a native metric on any of the common engines: you compute it from exported per-request data. Being able to say that out loud, rather than implying a flag turns it on, is part of the credible answer.
- Why require both a TTFT bound and a per-token bound rather than just one?Each alone has a degenerate optimum. A TTFT-only SLO rewards admitting every request instantly and then streaming at a crawl. A per-token-only SLO rewards heavy queueing so the admitted batch stays small and fast while everyone else waits. The conjunction closes both gaps and matches what a user actually perceives — a fast start and a readable pace.
- How does goodput change how you evaluate a tuning change?It reverses some verdicts. Most serving knobs buy aggregate tokens per second at the cost of first-token latency, so on a throughput metric they always look like improvements. Measured as goodput, a change that raises tokens/s by 15% while pushing a third of requests past their TTFT bound registers as the regression it is.
- Is goodput something the engine reports directly?No. vLLM's benchmark output and GenAI-Perf both give you per-request TTFT and per-token latency, and vLLM can export per-request results for post-processing, but applying the SLO predicate and counting passing requests is something you compute yourself. Treat it as a derived metric over exported benchmark data or over production traces.
saying these in an interview costs you the question
- Reporting token throughput as health while latency SLOs are breaching
- Defining goodput on first-token latency alone
- Sizing capacity from the peak-throughput point rather than peak goodput
- Applying one SLO to both interactive chat and background batch traffic
- Claiming an engine flag reports goodput natively