Why is GPU utilization a poor autoscaling signal for an LLM inference server?
answer
- busy is not the same as full
- batching hides the headroom
- the queue is what saturates
- utilization pins at 100% either way
- vllm:num_requests_waiting, not GPU percent
basics
~20 sGPU utilization only reports that a kernel was running during the sampling window. A server that batches continuously sits near 100% at low load and looks identical when its queue is backing up. Scale on waiting-request count, KV-cache occupancy and measured time-to-first-token instead.
solid answer
~50 sGPU utilization, as reported by nvidia-smi or DCGM, is an occupancy-over-time number: it says a kernel was resident, not how much work was in flight. An LLM server doing continuous batching keeps the GPU busy whether the batch holds one sequence or sixty, so utilization saturates far below capacity and reads the same at 5% load as at overload. GPU memory used is worse still — vLLM pre-allocates its KV cache up to `--gpu-memory-utilization` at startup, so the number is essentially constant regardless of traffic. The signals that actually track saturation come from the engine: how many requests are waiting, how full the KV cache is, and the TTFT you are measuring. vLLM exposes `vllm:num_requests_waiting`, `vllm:num_requests_running` and `vllm:kv_cache_usage_perc` on `/metrics`; TGI and Triton expose their own queue and batch series. Scale out when queued requests stay above zero for a sustained window, or when TTFT p95 eats the share of your latency budget you allocated to queueing. Keep utilization as a health signal, not a scaling one.
code
promql · 8 lines# leading signal: requests admitted but not yet in the running batch
avg(vllm:num_requests_waiting)
# early warning: how close the paged KV-cache pool is to forcing preemption
max(vllm:kv_cache_usage_perc)
# context: how full the running batch already is
avg(vllm:num_requests_running)go deeper
Know that an LLM server batches many requests into one GPU pass, so a busy GPU does not mean a full one. Be able to say that queued-request count is the signal to look at.
Explain why continuous batching decouples utilization from load, and name the engine metrics you would scrape instead — queued requests, running batch size, KV-cache occupancy — attributing the spellings to the engine that uses them.
Show you have set a real threshold: which gauge, what target, how long a stabilization window, and why you left headroom because reinforcements arrive minutes late. Mention preemption and KV-cache occupancy as the early-warning signal.
Own the policy across a fleet: which signal is contractual for every team, how thresholds relate to the TTFT SLO and to cold-start time, and the cost of the permanent headroom that slow scale-up forces you to buy.
## What the utilization number actually measures The familiar `utilization.gpu` from `nvidia-smi`, and the equivalent `DCGM_FI_DEV_GPU_UTIL` field from NVIDIA's DCGM exporter, report the fraction of the sampling window during which at least one kernel was executing on the device. It is a duty-cycle figure, not a work figure. One small kernel looping continuously reads 100%; a fully packed matmul that saturates every streaming multiprocessor also reads 100%. Nothing in that number tells you how much of the machine the kernel used, and nothing tells you whether anyone is waiting. On a stateless web service, CPU utilization works as a scaling signal because the queue and the CPU are coupled: more concurrent requests means more CPU time consumed, and the number climbs smoothly toward a saturation point. On an LLM server that coupling is broken by batching. ## Why continuous batching decouples the two An inference engine runs a scheduling loop: on every iteration it forms a batch of the sequences it is currently serving and executes one forward pass for all of them together. Decoding is memory-bandwidth-bound — the pass mostly reads model weights out of HBM — so adding more sequences to the batch costs very little extra time. That is exactly why batching pays. The consequence for monitoring is that the GPU is *busy* the moment there is a single active request, and it stays busy as concurrency climbs by two orders of magnitude. Utilization pins near its ceiling at low load and stops moving. It is equally uninformative in the other direction: when the queue is deep and new arrivals are waiting tens of seconds for admission, the GPU is still just "running kernels" and still reads 100%. A signal that is identical at 5% of capacity and at 300% of capacity cannot drive a scaling decision. Profiling-grade fields such as SM activity or achieved occupancy are closer to real work than the plain utilization gauge, but they still describe the device, not the queue, and they still say nothing about whether requests are being admitted promptly. ## Why GPU memory used is even worse A natural fallback is memory pressure — but LLM engines pre-allocate. vLLM sizes its paged KV-cache pool at startup against the `--gpu-memory-utilization` fraction (0.9 by default) and holds that allocation for the life of the process. From the driver's point of view the memory is used whether the server is idle or thrashing. Scaling on device memory would either never fire or fire immediately and never recover. ## The signals that do track saturation What saturates on an LLM server is the *scheduler*: the number of sequences it can keep in the running batch, bounded by the concurrent-sequence cap and by free KV-cache blocks. Everything beyond that waits. So the honest signals are: - **Queue depth.** vLLM's `vllm:num_requests_waiting` gauge counts requests admitted to the server but not yet in the running batch. Sustained non-zero means demand exceeds what this replica can hold. - **Running batch size.** `vllm:num_requests_running` tells you whether you are near the concurrency cap. - **KV-cache occupancy.** `vllm:kv_cache_usage_perc` (this is the 0.27 spelling; the older `gpu_`-prefixed name belonged to the V0 engine and is gone) shows how close the block pool is to forcing preemption. - **TTFT.** Measured time-to-first-token is the user-visible consequence of queueing and the thing your SLO is written against. TGI and Triton publish their own queue-time and batch-size series; the names differ per engine, which is precisely why you should read the engine's own `/metrics` rather than assume a shared vocabulary. ## Choosing a target Queue depth is a step function, not a smooth ratio, so pick a target that tolerates normal jitter: scale out when the per-replica average of waiting requests stays above a small threshold for 30–60 seconds, and require a much longer window before scaling in. KV-cache occupancy above roughly 80–90% is a useful second trigger because it predicts preemption before latency degrades. TTFT p95 works as a scaling signal too, but it is a lagging indicator — by the time it crosses your threshold, users have already felt it — so prefer queue depth as the leading signal and keep TTFT as the SLO you validate against. Whichever you pick, the signal has to reach the autoscaler as a custom or external metric; the engine's Prometheus endpoint is the source, and the collection path is ordinary cluster plumbing. ## The catch that makes thresholds hard All of this assumes the autoscaler can act quickly, and on GPU fleets it cannot — a new replica needs a node, a multi-gigabyte image, tens of gigabytes of weights and a warmup pass before it serves anything. That is why targets on an LLM fleet are set with far more headroom than on a web tier: you must trigger while the current replicas still have slack, because the reinforcement arrives minutes later.
- If queue depth is the leading signal, what do you do with TTFT?Keep it as the SLO and the validation signal rather than the trigger. TTFT p95 tells you whether the scaling policy is actually protecting users, and it is what you alert on. But it only rises after queueing has already happened, so driving scale-out from it guarantees you react late. Use queue depth to act and TTFT to judge whether the threshold was right.
- Why is requests-per-second a weak scaling signal here too?Because request cost varies enormously. A 50-token classification and a 4,000-token summarization both count as one request, but differ by two orders of magnitude in GPU work. A fleet at a fixed RPS can be idle or overloaded depending on prompt and output length, so an RPS target has to be re-tuned every time traffic mix shifts.
- Does GPU memory used ever tell you anything useful on a vLLM server?Only about configuration, not load. Because the KV-cache pool is pre-allocated against `--gpu-memory-utilization`, the number is roughly constant at runtime, so it is useful for confirming the process claimed the headroom you intended and for catching a co-tenant stealing memory. Runtime pressure shows up as KV-cache occupancy and preemption counters instead.
saying these in an interview costs you the question
- Treat GPU utilization like CPU on a web tier
- Claims 100% GPU utilization means the server is out of capacity
- Scales on GPU memory used, ignoring KV-cache pre-allocation
- Uses requests-per-second as a fixed capacity target
- Waits for TTFT alerts before adding replicas