An LLM server ran fine for an hour, then died with CUDA out of memory. How do you triage?
answer
- Healthy servers queue, they do not OOM
- read the per-process table, not the total
- who else is on this card?
- the outlier prompt, not the elapsed hour
- headroom is for kernels and collectives
basics
~20 sSuspect something outside the engine's own budget. Serving engines claim their KV pool at startup and then queue or preempt rather than OOM, so a steady-state failure usually means another process on the card, an activation spike from an unusually long prompt, or a utilization fraction set too high.
solid answer
~50 sStart from the fact that a correctly sized server should not OOM at runtime: it profiles memory at startup, reserves a fixed KV pool, and under overload queues or preempts instead of allocating more. So an hour-in OOM points somewhere outside that pool. Check `nvidia-smi` for a **second process** on the same GPU — a stray notebook, an eval job, a sidecar, or another replica scheduled onto the same device. Check whether the failure correlates with an **unusually long prompt**: prefill activations scale with prompt length, and the profiled peak assumed a cap you may not be enforcing. Check whether **headroom was squeezed** — vLLM's `--gpu-memory-utilization` pushed to 0.97 leaves nothing for kernel workspace and NCCL buffers. Then check for late allocations: runtime-loaded LoRA adapters, a multimodal encoder, or a metrics exporter. The fixes are lower the utilization fraction, cap max context and prompt length, and give the process the GPU to itself.
go deeper
Know that GPU memory is claimed up front by the serving process, and that nvidia-smi shows which processes hold it. Checking for a second process is the first thing to try.
Explain why a full KV pool causes queueing rather than an OOM, and name the two knobs that change the budget: the memory utilization fraction and the maximum served context length.
Run the triage as an ordered checklist and produce a root cause — the colocated job, the outlier prompt, the squeezed headroom — plus the guardrail that stops recurrence and the metric that would have warned you.
Own the isolation policy. Deciding that inference GPUs are never shared with batch or notebook workloads costs utilization and buys predictability; make that call explicitly rather than discovering it through incidents.
## The insight that shapes the whole triage An inference engine does not allocate KV memory on demand. At startup it measures free memory, loads weights, runs a profiling pass to find the activation peak, and then claims essentially all of the remainder as a **fixed pool of cache blocks**. Under load, when the pool is full, the scheduler's job is to *queue* new requests and *preempt* running ones — not to ask CUDA for more memory. That is precisely why a healthy, correctly sized server gets slow under overload instead of crashing. So the first question is not "how do I give it more memory", it is **"what allocated memory that the engine did not account for?"** ## The checklist, in order of likelihood **1. Another process on the device.** Run `nvidia-smi` and read the per-process table, not just the total. Common culprits: a debugging notebook someone left attached, an evaluation or embedding job scheduled to the same GPU, a second replica placed on the same card because the scheduler saw memory as fungible, or a profiler. The engine profiled its budget against whatever was free at *its* startup; anything that arrives later is stealing from a pool it already promised away. **2. An outlier prompt.** Prefill activation memory scales with prompt length, and the startup profile assumed the maximum you configured. If your configured cap is 32k but you profiled behaviour on 2k prompts, the first genuinely long document is a memory event, not just a slow request. This is the classic "fine for an hour" signature: the cause is a rare input, not elapsed time. Look for the request that was in flight at the crash, and its input token count. **3. Headroom set too aggressively.** vLLM's `--gpu-memory-utilization` defaults to 0.9 for a reason: the last 10% covers the CUDA context, kernel workspaces, NCCL collective buffers on sharded deployments, CUDA graph capture, and allocator fragmentation. Someone raising it to 0.97 to squeeze in more cache converts a comfortable server into one that fails on the first unusual allocation. TensorRT-LLM's Triton backend has the same knob as `kv_cache_free_gpu_mem_fraction`; TGI instead bounds the budget through per-request token caps such as `--max-total-tokens`. **4. Something loaded after startup.** Runtime-loaded LoRA adapters, a vision encoder that only allocates when the first image arrives, a lazily initialized tokenizer or reranker in the same process — all of these are allocations the startup profile never saw. **5. Fragmentation.** Repeated variable-sized allocations outside the block pool can leave enough free bytes in total but no contiguous run large enough. This is the least common cause on modern engines precisely because the KV pool is block-structured, but it does happen in the workspace region under mixed request shapes. ## What to actually do - **Isolate the GPU.** One serving process per device. On Kubernetes that means whole-GPU requests, not best-effort colocation, unless you are deliberately using a partitioning feature. - **Lower the utilization fraction** back toward the default and take the capacity hit. A server that survives is worth more cache than a server that does not. - **Enforce the caps you profiled with.** Set the maximum model length and maximum input length to what you actually intend to serve, and reject longer requests at the edge with a clear error rather than letting them reach the GPU. - **Reproduce with the offending shape.** Replay the longest prompt you saw against a staging replica; if it OOMs deterministically, you have found it in one step. - **Alert on the right signal.** Watch the engine's reported KV-cache utilization and the queue depth, and alert on sustained saturation. A pool that sits at 100% for minutes is telling you that you are one unusual request away from trouble, before the crash. ## The answer to avoid "Add more GPU memory" and "restart it on a schedule" are both non-answers, and the second is worse because it hides a bug that will reappear at higher traffic. So is blaming a memory leak — engines with a preallocated block pool do not leak KV memory in the usual sense; the memory is claimed once and reused. Find the allocation that was not in the budget.
- Why does a well-configured server slow down under overload instead of crashing?Because the KV cache is a fixed pool claimed at startup. When it is full the scheduler stops admitting new sequences and evicts or pauses running ones, so the failure mode is queueing latency, not allocation failure. Crashing means something consumed memory the scheduler did not know about, or the pool was sized against a smaller activation peak than the traffic actually produces.
- How would you stop a single 100k-token prompt from taking down a replica?Enforce a maximum input length at the gateway and at the engine, so oversized requests get a clean 4xx rather than reaching the GPU. Configure the served context to what you profiled with — vLLM caps it with `--max-model-len`, TGI with `--max-total-tokens`. Then verify by replaying your longest realistic prompt against a staging replica.
- The OOM only happens on a multi-GPU replica, never on a single GPU. What does that suggest?Sharded execution adds memory the single-GPU profile never needed: NCCL communication buffers, per-device duplicated activations and additional workspace for collectives. Pushing the utilization fraction near 1.0 is far less survivable when several gigabytes per device belong to the communication layer. Lower the fraction and re-profile on the actual topology.
- What do you monitor so this is caught before the crash?The engine's KV-cache utilization and pending-queue depth, plus per-process GPU memory from the node exporter. Sustained cache saturation with a growing queue is the pre-crash signature; a second process appearing on a device is the other one. Alert on both, because either predicts the failure minutes to hours ahead of it.
saying these in an interview costs you the question
- Blaming a KV-cache memory leak in a preallocated pool
- Fixing it by raising the memory utilization fraction
- Scheduling nightly restarts instead of finding the allocation
- Assuming free VRAM shown at startup stays free
- Ignoring input length as a memory variable