Your LLM server's waiting queue keeps growing under load — how do you diagnose the bottleneck?
answer
- running set versus waiting set
- which admission condition is failing?
- cap pinned with spare cache means throttled
- cache pinned means capacity, not config
- separate queue time from service time
basics
~20 sCompare the running-sequence count against the configured concurrency cap and the KV-cache utilization gauge. Pinned at the cap with spare cache means the cap is throttling you; pinned cache means memory is the limit and no scheduler knob will fix it.
solid answer
~50 sA continuously batched engine keeps a running set and a waiting set, and a request sits in the waiting set for exactly one of a few reasons. Read them off the server's own metrics rather than guessing. vLLM exposes `vllm:num_requests_running` and `vllm:num_requests_waiting` on `/metrics`, alongside a KV-cache utilization gauge. If running is pinned at your concurrency cap while cache utilization is comfortable, admission is artificially throttled — raise the cap. If cache utilization is pinned near 100%, memory is the limiter: every admitted sequence needs blocks, and the scheduler will refuse new work and may evict running sequences. That is a capacity conversation — shorter contexts, a smaller or quantized cache, more replicas — not a scheduler setting. A third case: long prompts consuming the per-step token budget slow ingestion even with capacity free. Growing waiting depth with all limits comfortable usually means arrival rate simply exceeds one replica's service rate.
code
bash · 1 linecurl -s http://localhost:8000/metrics | grep -E 'vllm:num_requests_(running|waiting)'go deeper
Know that requests wait in a queue when the server has no room for them, and that the server publishes counts of running and waiting requests you can look at.
Explain the admission conditions — concurrency cap, KV-cache space, per-step token budget — and which metric distinguishes each one from the others.
Show the diagnosis as a decision tree over real metrics, name the case where no knob helps because capacity is exhausted, and explain why an overloaded engine degrades silently rather than erroring.
Own the policy question: decide in advance whether excess load is queued or shed, set the queue-depth or timeout threshold that encodes that choice, and align it with the scaling signal and the latency contract you sold.
## The mental model to start from A continuously batched server is a queueing system with an unusual service discipline. There is a **waiting set** of accepted-but-unscheduled requests, a **running set** participating in the current forward pass, and an admission decision taken before every iteration. A growing waiting set means admission is being denied step after step. Diagnosis is therefore the question: *which* admission condition is failing? There are only a handful of candidates, which is what makes this tractable in an interview. ## Candidate 1 — the concurrency cap binds The engine will not run more than its configured maximum number of sequences at once. If the running count is sitting exactly at that number, that is your answer. The tell that it is *artificial* rather than real is the KV-cache utilization gauge sitting comfortably below full: you have memory you are not allowed to use. Raising the cap converts that headroom into throughput, up to the point where cache becomes the binding constraint instead. This is the happy case, because it is a one-line fix. ## Candidate 2 — KV-cache memory binds Every running sequence holds cache blocks proportional to its current length, and the pool is fixed at startup. When utilization pins near full, the scheduler stops admitting, and under pressure it will evict running sequences to keep the ones it has making progress. The signature is a full cache gauge, a running count *below* the configured cap, and a waiting set that grows. No scheduler tuning fixes this, and raising the concurrency cap actively makes it worse. The real levers are elsewhere: reduce the maximum context you promise, quantize the cache to a smaller dtype, shard the model so more memory is available per replica, or add replicas. This case is worth naming precisely because it is where candidates reach for the wrong knob. "The queue is long, so let more requests in" is exactly backwards when the cache is the constraint. ## Candidate 3 — prefill is eating the step With chunked prefill on, each iteration's token budget is shared between the decode tokens of running sequences and a slice of some waiting request's prompt. A workload dominated by very long prompts spends most of its budget on prefill, so requests take many iterations to become fully prefilled and start decoding. Throughput in *output* tokens looks poor while the GPU is genuinely busy. The tell is a workload profile — median prompt length in the tens of thousands — combined with healthy cache and a running count below the cap. Here the lever is the token budget, plus anything that avoids the prefill entirely, such as prefix reuse for shared system prompts. ## Candidate 4 — you are simply out of capacity If the running count is at the cap *and* the cache is full *and* step times are what you would expect for this model and hardware, then nothing is misconfigured: arrival rate exceeds this replica's service rate. Queueing theory then guarantees the waiting set grows without bound; there is no setting that makes a saturated server keep up. The response is horizontal — more replicas, routed by a signal that reflects the real state — or it is admission control: shed or reject work above a threshold rather than letting queue time inflate first-token latency for everyone. ## Getting the data Use the server's own metrics endpoint before reaching for GPU-level tools. vLLM's `/metrics` publishes `vllm:num_requests_running` and `vllm:num_requests_waiting`, plus gauges for cache utilization and histograms for request timings; TGI and Triton each expose their own Prometheus surface. A GPU utilization percentage is close to useless for this — a decode step keeps the GPU "busy" while doing very little math, so utilization can read high on a badly under-occupied server. The single most useful derived signal is the split between **queue time** and **service time** for a request. If first-token latency is mostly time spent waiting to be admitted, you have a capacity or admission problem. If it is mostly time spent being prefilled, you have a prompt-size or scheduling problem. Those two lead to completely different fixes, and a candidate who separates them is demonstrating the thing the question is testing. ## What overload looks like to the client Worth stating explicitly, because it distinguishes an LLM server from a stateless service. There is usually no fast failure. Requests are accepted, sit in the waiting set, and eventually stream slowly; the user's experience degrades as a long silence before the first token rather than an error. If you want an error instead — and often you should, for interactive traffic — you have to impose it yourself with a queue-depth threshold or a request timeout, so that the system fails visibly rather than turning everyone's latency into minutes.
- Why is a GPU utilization percentage a poor signal for whether an LLM server is saturated?Because a decode step keeps the GPU occupied while doing very little arithmetic — it is bound by streaming weights out of memory, not by math units. A server running one sequence per step can show high utilization while wasting almost all of its capability. The meaningful signals are the running-sequence count against its cap, KV-cache utilization, waiting-queue depth, and step duration.
- If capacity is genuinely exhausted, what should the client see?Whatever you decide it should — the default is silence. An overloaded engine accepts requests into its waiting set, so users experience a long delay before the first token rather than a failure. For interactive traffic that is usually the wrong behaviour: impose a queue-depth limit or request timeout so excess load is rejected quickly, protecting the latency of admitted requests instead of degrading everyone equally.
- How would you tell a long queue caused by a traffic spike from one caused by a bad deploy?Correlate arrival rate with service rate. A spike shows a rising request rate with unchanged step duration and unchanged per-request service time. A bad deploy shows flat arrivals with worse step times or a lower sustainable running-sequence count — typically from a changed model, precision, context setting or scheduler budget. Comparing step duration and running count against the previous version separates them immediately.
saying these in an interview costs you the question
- Raises the concurrency cap when the KV cache is already full
- Reads a high GPU utilization number as proof of saturation
- Assumes overloaded requests get rejected automatically
- Never separates queue time from prefill time in first-token latency
- Blames the model size before checking admission metrics