skip to content

Which vLLM /metrics series show whether the server is queueing or just slow?

level: seniorimportance: should knowfreq 50%

answer

  1. running versus waiting, first
  2. occupancy plus preemptions means memory
  3. histograms for what the user feels
  4. one flag empties every dashboard
  5. the cache gauge was renamed

basics

~10 s

vllm:num_requests_waiting against vllm:num_requests_running separates queueing from execution; vllm:kv_cache_usage_perc and vllm:num_preemptions say whether memory is the cause. vllm:time_to_first_token_seconds and vllm:inter_token_latency_seconds give the latency that clients feel.

solid answer

~40 s

The OpenAI-compatible server exposes Prometheus text at `/metrics` on the API port. Read it in three layers. **Queueing:** `vllm:num_requests_running` versus `vllm:num_requests_waiting` — a persistently non-zero waiting count means arrivals exceed what the scheduler can admit. **Memory pressure:** `vllm:kv_cache_usage_perc` near 1.0 together with a climbing `vllm:num_preemptions` counter means the block pool, not the GPU's compute, is the binding constraint. **Client-visible latency:** the `vllm:time_to_first_token_seconds` and `vllm:inter_token_latency_seconds` histograms. Waiting near zero with high TTFT points at long prompts or an undersized GPU rather than saturation. Two operational notes: `--disable-log-stats` turns statistics collection off and leaves the dashboards empty, and the old V0 name `vllm:gpu_cache_usage_perc` no longer exists — alert rules copied from older material silently never fire.

code

bash · 7 lines
bash
# Scrape once by hand after a version bump and check the series you alert on
curl -s localhost:8000/metrics | grep -E '^vllm:(num_requests_(running|waiting)|kv_cache_usage_perc|num_preemptions)'

# vllm:num_requests_running 3.0
# vllm:num_requests_waiting 41.0
# vllm:kv_cache_usage_perc 0.98
# vllm:num_preemptions_total 512.0

go deeper

for a junior

Know that a vLLM server publishes Prometheus metrics at /metrics on its own API port, with no separate exporter, and that the series are prefixed vllm:.

for a middle

Name the pairs and say what each answers: running versus waiting for queueing, kv_cache_usage_perc with num_preemptions for memory pressure, the TTFT and inter-token histograms for what the client feels.

for a senior

Show that you read them in combination to separate memory-bound from compute-bound saturation, and that you re-scrape and diff series names after a version bump because a renamed metric makes an alert silently stop firing.

for a principal

Own the observability contract for the fleet: which vLLM series are the standard capacity signals, how dashboards and alerts are versioned with the engine, and where device-level telemetry from DCGM fits alongside engine-level metrics.

## Where the numbers come from A vLLM server started with `vllm serve` (or the `vllm/vllm-openai` image) publishes Prometheus-format metrics at `/metrics` on the same port as the API. No sidecar or exporter is involved; you point a scrape at the pod. Collection is on by default and is switched off by `--disable-log-stats`, which is a real trap: the endpoint still responds, but the `vllm:` series are simply absent and every dashboard built on them goes blank. ## The three layers worth wiring up **Scheduler state.** `vllm:num_requests_running` is how many sequences the engine is decoding this iteration; `vllm:num_requests_waiting` is how many are admitted but not yet running. `vllm:num_requests_waiting_by_reason` breaks the wait down further. These two gauges are the single most useful pair on the dashboard, because they distinguish "the server has more work than it can start" from "the work it started is slow". **Memory pressure.** `vllm:kv_cache_usage_perc` reports how much of the KV block pool is occupied, on a 0-to-1 scale. `vllm:num_preemptions` counts sequences the scheduler had to evict to free blocks — a counter that should be flat and whose slope is a direct measure of over-subscription. When usage sits near 1.0 and preemptions are climbing, the fix is memory-shaped: fewer concurrent sequences, a shorter context ceiling, an fp8 KV cache, or more GPUs. There are also `vllm:kv_offload_cpu_cache_usage_perc` for CPU-offloaded blocks and `vllm:external_prefix_cache_hits` / `vllm:external_prefix_cache_queries` for external prefix-cache tiers. **Latency and work.** `vllm:time_to_first_token_seconds` covers everything before the first token — queue wait plus prefill. `vllm:inter_token_latency_seconds` and `vllm:request_time_per_output_token_seconds` cover the streaming phase. `vllm:iteration_tokens_total` shows how much work each scheduler step actually did, which is the honest measure of batching efficiency. `vllm:prompt_tokens` and `vllm:generation_tokens` are the token counters that feed cost accounting, and `vllm:prompt_tokens_cached` shows how many prompt tokens were served from the prefix cache rather than recomputed. `vllm:request_success` counts completed requests, labelled by finish reason. ## Reading them together The diagnostic value is in the combination, not any single series: - Waiting high, cache usage high, preemptions rising: **capacity-bound on memory.** More replicas, less context, or a cheaper cache dtype. - Waiting high, cache usage moderate, iteration tokens flat: **compute-bound.** The GPU is doing all it can per step; a faster card or a smaller model is the lever. - Waiting near zero, TTFT high: **not saturation.** Look at prompt length distribution, prefix-cache hit rate, and whether one enormous request is dominating each prefill. - Everything calm inside the server, users complaining: the problem is outside vLLM — the ingress, the routing layer, or a client that is not consuming the stream. ## Naming is a versioned claim The metric namespace has changed over vLLM's life, and copying a dashboard from a two-year-old blog post produces panels that quietly display nothing. The specific one to know: the KV-cache occupancy gauge was `vllm:gpu_cache_usage_perc` under the V0 engine and is `vllm:kv_cache_usage_perc` in current versions — the old name does not exist any more. Because a Prometheus query for a missing series returns empty rather than erroring, this failure mode is invisible until an incident, when the panel everyone trusts is blank. The habit that prevents it is to scrape `/metrics` once by hand after any version bump and diff the series names against what your dashboards and alert rules reference. ## What these metrics are not They are per-replica engine internals. They do not tell you about GPU temperature, ECC errors or device utilization — that is DCGM's job — and they say nothing about what is happening between replicas. GPU device memory in particular is useless as a load signal for vLLM, because the engine claims its budget at startup and holds it flat regardless of traffic; the occupancy of the block pool is the number that actually moves.

  • Your dashboards are empty although /metrics returns 200. What do you check first?
    The launch arguments, for --disable-log-stats. With statistics collection off the endpoint still answers but publishes none of the vllm: series, so every panel is blank while the server is perfectly healthy. If that is not it, check the series names: a version upgrade can rename a metric, and Prometheus returns an empty result rather than an error for a query naming a series that no longer exists.
  • Why isn't GPU memory used a good load signal for a vLLM replica?
    Because vLLM claims its memory budget at startup and allocates the KV block pool immediately, so device memory reads near-constant whether the server is idle or saturated. The quantity that actually varies is how much of that pool is occupied, which vLLM exports as vllm:kv_cache_usage_perc, alongside vllm:num_requests_waiting for queue depth. Those move with load; device memory does not.
  • How do you tell a slow prefill from a long queue when TTFT p95 rises?
    Compare TTFT against vllm:num_requests_waiting over the same window. TTFT includes both the wait before admission and the prefill itself, so a rise with a growing waiting gauge is queueing, while a rise with a flat, near-zero waiting gauge is prefill cost — usually longer prompts, or a drop in prefix-cache hits visible in vllm:prompt_tokens_cached relative to vllm:prompt_tokens.

saying these in an interview costs you the question

  • Alerts on the old gpu_cache_usage_perc name
  • Uses GPU memory used as the load signal
  • Runs with --disable-log-stats then wonders why panels are blank
  • Reads TTFT alone without checking queue depth
  • Treats preemption count as harmless noise

context