Why does a fixed-length synthetic benchmark overstate an LLM server's capacity?
answer
- the same prompt twice is not two prompts
- real answers are heavy-tailed
- closed loop cannot queue
- bursts, not a metronome
- warm up before you count
basics
~10 sIdentical prompts hit the server's prefix cache, uniform lengths remove the long-request tail that clogs real batches, and steady closed-loop pacing never produces bursts. The result is a throughput number production traffic cannot reproduce.
solid answer
~40 sFour things make synthetic runs flattering. First, **repeated identical prompts** land on the prefix cache — vLLM enables automatic prefix caching by default in its V1 engine, so a replayed prompt skips prefill entirely and prefill cost silently vanishes from the measurement. Second, **uniform output lengths** mean every request in a batch finishes together; real traffic mixes 40-token and 2,000-token answers, and the long ones occupy KV-cache blocks and batch slots for far longer, dragging everyone's inter-token latency. Third, **closed-loop pacing** self-throttles: if the server slows, the generator launches fewer requests, so queues never build. Real arrivals are bursty and do not back off. Fourth, **no warmup** cuts the other way and can understate. The fix is to replay a production-derived length distribution, randomise prompt prefixes, and use an open-loop arrival process with burstiness.
code
bash · 8 linesvllm bench serve \
--base-url http://localhost:8000 \
--model meta-llama/Llama-3.1-8B-Instruct \
--dataset-name sharegpt \
--dataset-path /data/sharegpt.json \
--num-prompts 2000 \
--request-rate 15 \
--burstiness 0.5go deeper
Know that benchmark traffic must resemble real traffic — same prompt lengths, same answer lengths — or the resulting capacity number will not hold in production.
Explain the specific mechanisms: prefix cache hits on repeated prompts, missing long-request tail in a uniform batch, and closed-loop pacing that self-throttles instead of queueing.
Show that you derive the length distribution and prefix-sharing rate from production logs, run open-loop at target rate to expose the latency cliff, and re-measure after every engine, model or quantization change.
Treat the benchmark harness as owned infrastructure: versioned, fed by production telemetry, run in CI on model and engine upgrades, and the agreed source of truth when capacity or vendor claims are disputed.
## The gap between a benchmark and production A benchmark is a model of your traffic, and like any model it is wrong in specific, predictable ways. On an inference server the errors nearly all point the same direction — the synthetic run looks better than reality — which is why teams routinely provision from a benchmark and then get paged. ## Error 1: the prefix cache eats your prefill Modern engines cache the key/value state of prompt prefixes so a repeated prefix does not have to be recomputed. vLLM turns automatic prefix caching on by default in its V1 engine. If your load generator sends the same prompt 500 times, request 2 through 500 skip most or all of prefill. That is a real production feature and it genuinely helps workloads with shared system prompts. The problem is measurement honesty: your run reports a throughput that assumes a cache hit rate near 100%, while your production hit rate might be 20%. Fix it by generating prompts with random content (`--dataset-name random` in `vllm bench serve` produces randomised token content rather than one repeated string) or by using a real conversational dataset. If your production traffic does share a big system prompt, then keep that shared prefix in the benchmark deliberately — but know that you did, and know your real hit rate. ## Error 2: uniform lengths hide the long-request tail The most consequential simplification. Real chat traffic has a heavy-tailed output length distribution: most answers are short, a few are enormous. On a continuously-batched server, a request that generates 2,000 tokens sits in the batch for the whole of those 2,000 decode steps, holding KV-cache blocks and a scheduler slot. It raises the memory pressure and the per-step cost that every concurrently decoding request pays. A uniform 256-token benchmark never creates that situation. Every request enters and leaves together, memory usage is flat, and the scheduler never has to make an eviction or preemption decision. Then production arrives with a p99 output length of 3,000 tokens, KV-cache capacity binds, the engine starts preempting, and observed throughput is a fraction of the benchmark's. Uniform *input* lengths mislead too: long prompts make prefill expensive and can force the scheduler to interleave prefill work with ongoing decode, which raises inter-token latency for everyone already streaming. ## Error 3: closed-loop pacing cannot show you queue collapse A closed-loop generator keeps exactly N requests in flight. When the server slows, completions slow, so the generator issues fewer requests per second. The load automatically adapts to the server's capacity, which means the queue can never run away and latency degrades gracefully no matter how overloaded you are. Real users do not adapt. They arrive at whatever rate they arrive at. Model that with an open-loop arrival process: `vllm bench serve --request-rate R` uses exponential inter-arrival times (a Poisson process), and `--burstiness` reshapes the distribution to be more or less bursty than Poisson. Only under open loop does the latency cliff past saturation appear, and that cliff is what your on-call rotation actually experiences. ## Error 4: no warmup (the error in the other direction) The first requests after server start pay one-time costs: lazy kernel loading, CUDA graph capture, compilation. Including them drags the tail percentiles up and makes the server look worse than steady state. Send a warmup batch and discard it. Note that this error and the previous three point opposite ways, so "my benchmark is pessimistic in one place" does not cancel out the optimism elsewhere. ## Building a benchmark you can provision from 1. **Sample the shape from production.** Log input and output token counts per request. Take the empirical distribution — not the mean — and replay it. 2. **Match the prefix-sharing rate.** If 90% of your traffic shares one system prompt, keep that in the benchmark. If none does, randomise. 3. **Use open loop at your target rate**, with burstiness, and check whether the queue is stable rather than only reading the average latency. 4. **Warm up and discard.** 5. **Run long enough** to reach steady state — a 30-second run on a workload with 2,000-token answers has barely filled the pipeline. 6. **Re-measure after every change** to engine version, model, quantization, or parallelism, because all four move the curve. ## The one-sentence version A synthetic benchmark measures the server; a production-shaped benchmark measures the system. Provision from the second.
- If your production traffic really does share one large system prompt, should the benchmark keep it?Yes — then prefix cache hits are a genuine property of your workload and excluding them would understate capacity. The rule is to match the production hit rate, not to eliminate caching. Measure the real hit rate from server metrics first, then construct the benchmark's prompt mix so it reproduces roughly that rate, and state the rate alongside the result.
- How long should a benchmark run be?Long enough to reach steady state and to sample the tail of your length distribution. If p99 answers take 60 seconds to generate, a 30-second run never contains one. A practical rule: run until the throughput and p95 latency readings stop drifting, then keep running for several multiples of your longest expected request, and discard the warmup window.
- What makes bursty arrivals harder on an inference server than a steady rate at the same average?A burst arrives faster than the scheduler can admit it, so requests queue and every queued request's time to first token inherits the wait. Because KV-cache capacity is finite, a large burst can also push the engine into preempting in-flight sequences, which throws away decode work and costs throughput on top of latency.
saying these in an interview costs you the question
- Replaying one identical prompt and calling the cache hits capacity
- Using a fixed output length when real answers are heavy-tailed
- Closed-loop testing only, so queue explosion is invisible
- Including cold-start compilation in the latency percentiles
- Running a 30-second test and provisioning from it