skip to content

How do you load-test a self-hosted LLM endpoint, and what do you report?

level: middleimportance: should knowfreq 56%

answer

  1. generic HTTP tools cannot see first token
  2. pin the lengths, pin the load level
  3. arrival rate or in-flight cap
  4. percentiles, never means
  5. warm up and discard

basics

~20 s

Drive the server with a token-aware load generator such as vLLM's vllm bench serve or GenAI-Perf, pinning prompt and output lengths and either an arrival rate or a concurrency cap. Report TTFT, per-output-token latency and inter-token latency percentiles alongside output token throughput.

solid answer

~50 s

Use a generator that understands streaming tokens, not a generic HTTP tool — `ab` or `wrk` will time the whole response body and tell you nothing about time to first token. Two standard choices: **`vllm bench serve`**, the load generator shipped in the vLLM CLI (vLLM 0.27; it replaced the older `benchmarks/benchmark_serving.py`), and **GenAI-Perf**, which now lives in the `triton-inference-server/perf_analyzer` repository rather than the Triton server repo. Both hit an OpenAI-compatible endpoint, both let you pin synthetic input and output lengths, and both report the four numbers that matter: **TTFT** and **inter-token latency** percentiles, end-to-end request latency, and **output token throughput**. The critical choice is open-loop versus closed-loop. `vllm bench serve` defaults to `--request-rate inf` (fire everything at once, a closed-loop stress test); set a finite `--request-rate` to model Poisson arrivals, or `--max-concurrency` to cap in-flight requests. Report percentiles, never means — LLM latency distributions have long tails.

code

bash · 9 lines
bash
vllm bench serve \
  --base-url http://localhost:8000 \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --dataset-name random \
  --random-input-len 1024 \
  --random-output-len 256 \
  --num-prompts 500 \
  --max-concurrency 32 \
  --ignore-eos

go deeper

for a junior

Know that LLM endpoints need a streaming-aware load tool, and be able to name one such as vLLM's vllm bench serve or GenAI-Perf. Say that you report first-token latency separately from total time.

for a middle

Explain the run configuration you would pin — input and output token lengths, forced output length, and either an arrival rate or a concurrency cap — and why percentiles rather than means are reported.

for a senior

Demonstrate open-loop versus closed-loop judgment, warmup discipline, and awareness that server-side prefix caching or a saturated client can manufacture numbers the production system will never reproduce.

for a principal

Own the benchmark as a durable artifact: a versioned harness with a production-derived length distribution, run on every engine or model upgrade, with results stored so regressions are detectable rather than argued about.

## Why a generic HTTP load tool is the wrong instrument Tools like `ab`, `wrk`, `hey` and k6 (without a custom script) measure one thing: how long the full response took. For a streaming LLM endpoint that number is close to useless on its own. It merges the time the user waited before seeing anything (**time to first token**) with the time spent streaming the rest of the answer, and it moves with output length, which the model chose. You need a generator that parses server-sent events, timestamps each token as it arrives, and reports the first-token wait separately from the per-token cadence. ## The two standard generators **`vllm bench serve`** is part of the vLLM CLI as of vLLM 0.27, where the `vllm bench` subcommand absorbed what used to be the `benchmarks/benchmark_serving.py` script. It targets a running OpenAI-compatible server over `--base-url`, builds load from a dataset (`--dataset-name random` for synthetic, or ShareGPT/HF datasets for realistic conversational shapes), and prints request throughput, output token throughput, total token throughput, and percentile tables for TTFT, TPOT and ITL. **GenAI-Perf** comes from NVIDIA's perf-analyzer project. Note the currency detail: `genai-perf` and `perf_analyzer` were moved out of the `triton-inference-server/server` repository into `triton-inference-server/perf_analyzer`, so instructions that send you to the server container to find them are stale. `genai-perf profile` speaks the OpenAI chat protocol (`--endpoint-type chat --streaming`), so it works against vLLM and TGI, not only Triton. ## The knobs that decide what you are actually measuring **Input and output length.** Pin both. `--random-input-len` / `--random-output-len` in vLLM, `--synthetic-input-tokens-mean` / `--output-tokens-mean` in GenAI-Perf. Match them to your production distribution. **Force the output length.** Models stop when they emit an end-of-sequence token, so a run configured for 256 output tokens may average 80. `vllm bench serve --ignore-eos` suppresses that so every request generates exactly the requested count, which is what makes runs comparable. **Open loop vs closed loop.** This is the choice people get wrong most often. - *Closed loop* fixes the number of in-flight requests (`--max-concurrency N`): every completion immediately launches a replacement. It answers "at concurrency N, what latency and throughput do I get?" and it can never build an unbounded queue, because the load backs off automatically when the server slows down. - *Open loop* fixes the arrival rate (`--request-rate R`), independent of how the server is coping. `vllm bench serve` draws exponential inter-arrival times, i.e. a Poisson process, and `--burstiness` changes the shape away from Poisson. Open loop is the one that reveals queue explosion, because arrivals do not slow down when the server does. Use closed loop to map the latency/throughput curve, and open loop to test whether a specific target rate is survivable. **Warmup.** The first requests after startup pay for lazy CUDA kernel loading, graph capture and compilation. Send a warmup batch and discard it, or you will benchmark startup instead of steady state. ## What to report Four numbers, all as percentiles (p50, p95, p99) rather than means: 1. **TTFT** — how long before the user sees the first character. 2. **Per-output-token latency / inter-token latency** — the streaming cadence after that. 3. **End-to-end request latency** — only meaningful next to the output length it was measured at. 4. **Output token throughput** — the server-side capacity number. Plus the run configuration: model, quantization, GPU SKU and count, engine version, input/output lengths, and the arrival rate or concurrency. A latency number without its load level is not a measurement; it is a screenshot. ## Client-side measurement pitfalls Run the generator close to the server (same VPC/zone) unless you are deliberately measuring wide-area latency, or network round-trip contaminates TTFT. Watch the client's own CPU: a Python generator at high concurrency can become the bottleneck and manufacture latency that the server never caused. And be aware of caching on the server side — vLLM enables automatic prefix caching by default in its V1 engine, so replaying identical prompts measures the cache, not the model. ## What good looks like in an interview The strong answer names a token-aware generator, states the length distribution it drove, distinguishes open from closed loop, reports percentiles, and attaches the hardware and engine version. The weak answer is "I ran a load test and it did 500 ms average latency."

  • When would you choose an open-loop arrival rate over a fixed concurrency?
    When you are testing whether a specific production rate is survivable. Closed loop self-throttles — if the server slows, fewer requests are launched, so queues never explode and the test always looks stable. Open loop keeps arriving at the configured rate regardless, which is what real users do, and it is the only setup that exposes runaway queueing and the latency cliff past saturation.
  • Why does forcing the output length matter for comparability?
    Without it, the model stops at its own end-of-sequence token, so the realised output length varies per run and per configuration. Two runs that generated different average token counts are not comparable on either latency or throughput. `vllm bench serve --ignore-eos` pins the count so the only variable is the server configuration you are testing.
  • What run metadata must accompany a published latency number?
    Engine and version, model and quantization, GPU SKU and count, parallelism settings, input and output token lengths, and the load level (arrival rate or concurrency) with the percentile reported. Any latency figure missing the load level is unreproducible, because latency on an inference server is a function of load, not a property of the model.

saying these in an interview costs you the question

  • Using ab or wrk, which cannot separate TTFT from total response time
  • Reporting mean latency instead of p95 and p99
  • Benchmarking without warmup, so startup compilation lands in the numbers
  • Letting the model stop early, so runs generated different output lengths
  • Running the load generator across the internet and blaming the server for the round-trip

context