skip to content

Serving Concepts

You will learn the engine-agnostic mechanics every inference server implements — batching, KV-cache management, quantization, speculation, latency budgets, GPU sizing, sharding, and scaling. Interviewers use these to test whether you understand why an LLM server behaves nothing like a stateless REST service.

on this pageshow

explore

questions

page 1 of 2

Why is GPU utilization a poor autoscaling signal for an LLM inference server?

level: middleimportance: must knowfreq 60%

answer

  1. busy is not the same as full
  2. batching hides the headroom
  3. the queue is what saturates
  4. utilization pins at 100% either way
  5. vllm:num_requests_waiting, not GPU percent

basics

~20 s

GPU utilization only reports that a kernel was running during the sampling window. A server that batches continuously sits near 100% at low load and looks identical when its queue is backing up. Scale on waiting-request count, KV-cache occupancy and measured time-to-first-token instead.

solid answer

~50 s

GPU utilization, as reported by nvidia-smi or DCGM, is an occupancy-over-time number: it says a kernel was resident, not how much work was in flight. An LLM server doing continuous batching keeps the GPU busy whether the batch holds one sequence or sixty, so utilization saturates far below capacity and reads the same at 5% load as at overload. GPU memory used is worse still — vLLM pre-allocates its KV cache up to `--gpu-memory-utilization` at startup, so the number is essentially constant regardless of traffic. The signals that actually track saturation come from the engine: how many requests are waiting, how full the KV cache is, and the TTFT you are measuring. vLLM exposes `vllm:num_requests_waiting`, `vllm:num_requests_running` and `vllm:kv_cache_usage_perc` on `/metrics`; TGI and Triton expose their own queue and batch series. Scale out when queued requests stay above zero for a sustained window, or when TTFT p95 eats the share of your latency budget you allocated to queueing. Keep utilization as a health signal, not a scaling one.

code

promql · 8 lines
promql
# leading signal: requests admitted but not yet in the running batch
avg(vllm:num_requests_waiting)

# early warning: how close the paged KV-cache pool is to forcing preemption
max(vllm:kv_cache_usage_perc)

# context: how full the running batch already is
avg(vllm:num_requests_running)

go deeper

for a junior

Know that an LLM server batches many requests into one GPU pass, so a busy GPU does not mean a full one. Be able to say that queued-request count is the signal to look at.

for a middle

Explain why continuous batching decouples utilization from load, and name the engine metrics you would scrape instead — queued requests, running batch size, KV-cache occupancy — attributing the spellings to the engine that uses them.

for a senior

Show you have set a real threshold: which gauge, what target, how long a stabilization window, and why you left headroom because reinforcements arrive minutes late. Mention preemption and KV-cache occupancy as the early-warning signal.

for a principal

Own the policy across a fleet: which signal is contractual for every team, how thresholds relate to the TTFT SLO and to cold-start time, and the cost of the permanent headroom that slow scale-up forces you to buy.

## What the utilization number actually measures The familiar `utilization.gpu` from `nvidia-smi`, and the equivalent `DCGM_FI_DEV_GPU_UTIL` field from NVIDIA's DCGM exporter, report the fraction of the sampling window during which at least one kernel was executing on the device. It is a duty-cycle figure, not a work figure. One small kernel looping continuously reads 100%; a fully packed matmul that saturates every streaming multiprocessor also reads 100%. Nothing in that number tells you how much of the machine the kernel used, and nothing tells you whether anyone is waiting. On a stateless web service, CPU utilization works as a scaling signal because the queue and the CPU are coupled: more concurrent requests means more CPU time consumed, and the number climbs smoothly toward a saturation point. On an LLM server that coupling is broken by batching. ## Why continuous batching decouples the two An inference engine runs a scheduling loop: on every iteration it forms a batch of the sequences it is currently serving and executes one forward pass for all of them together. Decoding is memory-bandwidth-bound — the pass mostly reads model weights out of HBM — so adding more sequences to the batch costs very little extra time. That is exactly why batching pays. The consequence for monitoring is that the GPU is *busy* the moment there is a single active request, and it stays busy as concurrency climbs by two orders of magnitude. Utilization pins near its ceiling at low load and stops moving. It is equally uninformative in the other direction: when the queue is deep and new arrivals are waiting tens of seconds for admission, the GPU is still just "running kernels" and still reads 100%. A signal that is identical at 5% of capacity and at 300% of capacity cannot drive a scaling decision. Profiling-grade fields such as SM activity or achieved occupancy are closer to real work than the plain utilization gauge, but they still describe the device, not the queue, and they still say nothing about whether requests are being admitted promptly. ## Why GPU memory used is even worse A natural fallback is memory pressure — but LLM engines pre-allocate. vLLM sizes its paged KV-cache pool at startup against the `--gpu-memory-utilization` fraction (0.9 by default) and holds that allocation for the life of the process. From the driver's point of view the memory is used whether the server is idle or thrashing. Scaling on device memory would either never fire or fire immediately and never recover. ## The signals that do track saturation What saturates on an LLM server is the *scheduler*: the number of sequences it can keep in the running batch, bounded by the concurrent-sequence cap and by free KV-cache blocks. Everything beyond that waits. So the honest signals are: - **Queue depth.** vLLM's `vllm:num_requests_waiting` gauge counts requests admitted to the server but not yet in the running batch. Sustained non-zero means demand exceeds what this replica can hold. - **Running batch size.** `vllm:num_requests_running` tells you whether you are near the concurrency cap. - **KV-cache occupancy.** `vllm:kv_cache_usage_perc` (this is the 0.27 spelling; the older `gpu_`-prefixed name belonged to the V0 engine and is gone) shows how close the block pool is to forcing preemption. - **TTFT.** Measured time-to-first-token is the user-visible consequence of queueing and the thing your SLO is written against. TGI and Triton publish their own queue-time and batch-size series; the names differ per engine, which is precisely why you should read the engine's own `/metrics` rather than assume a shared vocabulary. ## Choosing a target Queue depth is a step function, not a smooth ratio, so pick a target that tolerates normal jitter: scale out when the per-replica average of waiting requests stays above a small threshold for 30–60 seconds, and require a much longer window before scaling in. KV-cache occupancy above roughly 80–90% is a useful second trigger because it predicts preemption before latency degrades. TTFT p95 works as a scaling signal too, but it is a lagging indicator — by the time it crosses your threshold, users have already felt it — so prefer queue depth as the leading signal and keep TTFT as the SLO you validate against. Whichever you pick, the signal has to reach the autoscaler as a custom or external metric; the engine's Prometheus endpoint is the source, and the collection path is ordinary cluster plumbing. ## The catch that makes thresholds hard All of this assumes the autoscaler can act quickly, and on GPU fleets it cannot — a new replica needs a node, a multi-gigabyte image, tens of gigabytes of weights and a warmup pass before it serves anything. That is why targets on an LLM fleet are set with far more headroom than on a web tier: you must trigger while the current replicas still have slack, because the reinforcement arrives minutes later.

  • If queue depth is the leading signal, what do you do with TTFT?
    Keep it as the SLO and the validation signal rather than the trigger. TTFT p95 tells you whether the scaling policy is actually protecting users, and it is what you alert on. But it only rises after queueing has already happened, so driving scale-out from it guarantees you react late. Use queue depth to act and TTFT to judge whether the threshold was right.
  • Why is requests-per-second a weak scaling signal here too?
    Because request cost varies enormously. A 50-token classification and a 4,000-token summarization both count as one request, but differ by two orders of magnitude in GPU work. A fleet at a fixed RPS can be idle or overloaded depending on prompt and output length, so an RPS target has to be re-tuned every time traffic mix shifts.
  • Does GPU memory used ever tell you anything useful on a vLLM server?
    Only about configuration, not load. Because the KV-cache pool is pre-allocated against `--gpu-memory-utilization`, the number is roughly constant at runtime, so it is useful for confirming the process claimed the headroom you intended and for catching a co-tenant stealing memory. Runtime pressure shows up as KV-cache occupancy and preemption counters instead.

saying these in an interview costs you the question

  • Treat GPU utilization like CPU on a web tier
  • Claims 100% GPU utilization means the server is out of capacity
  • Scales on GPU memory used, ignoring KV-cache pre-allocation
  • Uses requests-per-second as a fixed capacity target
  • Waits for TTFT alerts before adding replicas

context

open as a page

How does continuous batching differ from static batching in an LLM inference server?

level: middleimportance: must knowfreq 85%

basics

~20 s

Static batching forms a batch, runs it to completion, and only then accepts new work, so every slot waits for the longest generation. Continuous batching re-forms the batch before each decode step: finished sequences leave and queued requests join immediately.

open as a page

Will a 70B model fit on 2xA100-80GB at 32k context? How do you check?

level: middleimportance: must knowfreq 72%

basics

~20 s

Add the four terms up. 70B weights at bf16 are about 140 GB of the 160 GB on the node, so after CUDA context and activations only a few GB are left for KV cache. It loads, but it serves almost no concurrency.

open as a page

How does a PagedAttention block table let one sequence's KV cache be non-contiguous?

level: middleimportance: must knowfreq 66%

basics

~20 s

PagedAttention splits the KV cache into fixed-size blocks, each holding a set number of tokens. Every sequence keeps a block table listing which physical blocks hold its logical positions, and the attention kernel follows that table instead of striding through one range.

open as a page

Why does a contiguous per-request KV cache waste most of an LLM server's VRAM?

level: middleimportance: must knowfreq 62%

basics

~20 s

A contiguous allocator must reserve one slab per request sized for the longest output that request could produce. Most requests never grow into it, so the reserved tail sits idle, and the leftover gaps between slabs are too small for the next request to use.

open as a page

When does automatic prefix caching cut TTFT, and when does it silently miss?

level: middleimportance: must knowfreq 58%

basics

~20 s

Prefix caching reuses already-computed KV blocks when a new request starts with the exact same tokens as an earlier one, skipping that part of prefill. It misses whenever the prompt differs at the front, when the shared run is shorter than one block, or when those blocks have already been evicted.

open as a page

For an LLM server, why is requests/s a misleading capacity metric?

level: middleimportance: must knowfreq 62%

basics

~20 s

Requests per second hides how much work each request does: one call may emit 40 output tokens, another 2,000. LLM capacity tracks output tokens per second, so report token throughput together with the input and output length distribution it was measured at.

open as a page

In multi-GPU LLM inference, how do tensor and pipeline parallelism differ?

level: middleimportance: must knowfreq 72%

basics

~20 s

Tensor parallelism splits each layer's weight matrices across GPUs, so all of them work on every token and must all-reduce partial results at each layer. Pipeline parallelism gives each GPU a different block of layers and passes activations forward once per stage.

open as a page

Why can 4-bit weight-only quantization make an LLM server slower at large batch?

level: middleimportance: must knowfreq 60%

basics

~20 s

Weight-only 4-bit helps only while decoding is memory-bandwidth-bound. At large batch the server becomes compute-bound, and every matmul must first dequantize weights back to 16-bit — extra work fp16 never pays, so throughput can fall.

open as a page

How does speculative decoding get several tokens out of one target-model forward pass?

level: middleimportance: must knowfreq 65%

basics

~20 s

A cheap drafter guesses the next few tokens. The expensive target model then scores all of those guessed positions in a single forward pass, keeps the longest prefix it agrees with, and emits one corrected token where it disagrees.

open as a page

Why does a new GPU replica take minutes to serve LLM traffic?

level: seniorimportance: must knowfreq 55%

basics

~20 s

A new replica must wait for a GPU node, pull a multi-gigabyte container image, fetch tens of gigabytes of weights, load and shard them into VRAM, and run a warmup pass before it is useful. Each stage is minutes, and none of them is code you wrote.

open as a page

How do you compute the cost per million tokens of a self-hosted LLM endpoint?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Divide the hourly cost of the GPU by the tokens it actually produces in that hour. GPU $/hr divided by (tokens per second times 3600), times one million. The number that matters is measured throughput at your real utilization, not the card's peak.

open as a page

What happens to a running request when the KV cache runs out of free blocks?

level: seniorimportance: must knowfreq 52%

basics

~20 s

The scheduler preempts a running sequence: it frees that sequence's blocks and returns it to the waiting queue, either discarding its cache to recompute later or copying it to host memory. The client sees a stall, not an error, and tokens already streamed are never taken back.

open as a page

Output tokens/s went flat but TTFT p95 keeps climbing — what is happening?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The server is past its saturation knee. Decode throughput is capped by memory bandwidth and KV-cache capacity, so extra concurrency no longer produces more tokens — it only waits in the scheduler queue, and that queue time lands entirely in time to first token.

open as a page

Why does tensor parallelism scale poorly over PCIe compared with NVLink?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Tensor parallelism runs an all-reduce roughly twice per transformer layer, on the critical path of every token. NVLink moves those messages at hundreds of GB/s between GPUs; PCIe offers roughly an order of magnitude less bandwidth and higher latency, so the collectives dominate the step.

open as a page

Which quantization scheme suits an A10G versus an H100 inference server?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Hardware decides. FP8 tensor cores exist only on Ada and Hopper generations and newer, so an H100 should serve FP8 weights and activations. An Ampere card like the A10G or A100 has no FP8 math and should serve 4-bit GPTQ or AWQ through a Marlin-class kernel, or INT8.

open as a page

Before switching a production endpoint to a quantized model, what do you validate?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Validate both halves of the trade. On quality, compare against the full-precision baseline on real task evals and the failure modes quantization hits hardest. On serving, measure TTFT, inter-token latency and throughput at production concurrency — then canary traffic with the old replica warm for rollback.

open as a page

Speculative decoding drafts 5 tokens but latency barely moved — how do you diagnose and tune it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Measure the acceptance rate and the per-position acceptance curve first, then the draft cost as a fraction of a target step. Low acceptance means a mismatched or weak drafter; high acceptance with no win means the drafter is too expensive or the draft length overshoots.

open as a page

Why does round-robin load balancing misbehave across LLM replicas?

level: middleimportance: should knowfreq 50%

basics

~20 s

Round-robin equalizes request counts, but LLM requests differ in cost by orders of magnitude and run for seconds to minutes. One replica ends up holding all the long generations while another idles, so tail latency blows up even though every replica got the same number of requests.

open as a page

What does chunked prefill change about how an LLM server schedules a long prompt?

level: middleimportance: should knowfreq 55%

basics

~20 s

Chunked prefill splits a long prompt's prefill into pieces that fit the scheduler's per-step token budget, so each step can carry prefill tokens and ongoing decode tokens together instead of stalling every generating sequence behind one huge prompt.

open as a page

How do you pick a GPU SKU (L4, A10, A100, H100) for serving an 8B LLM?

level: middleimportance: should knowfreq 55%

basics

~20 s

Capacity gates the choice, bandwidth sets the speed, and price decides between what is left. An 8B model at bf16 needs ~16 GB, so a 24 GB L4 or A10 holds it but leaves thin cache room; bigger cards buy concurrency and far faster decoding.

open as a page

Why does a fixed-length synthetic benchmark overstate an LLM server's capacity?

level: middleimportance: should knowfreq 47%

basics

~10 s

Identical prompts hit the server's prefix cache, uniform lengths remove the long-request tail that clogs real batches, and steady closed-loop pacing never produces bursts. The result is a throughput number production traffic cannot reproduce.

open as a page

How do you load-test a self-hosted LLM endpoint, and what do you report?

level: middleimportance: should knowfreq 56%

basics

~20 s

Drive the server with a token-aware load generator such as vLLM's vllm bench serve or GenAI-Perf, pinning prompt and output lengths and either an arrival rate or a concurrency cap. Report TTFT, per-output-token latency and inter-token latency percentiles alongside output token throughput.

open as a page

What limits which tensor-parallel sizes an LLM checkpoint supports?

level: middleimportance: should knowfreq 45%

basics

~20 s

The degree must divide the model's structure: the attention head count above all, plus the MLP intermediate size, and with grouped-query attention the much smaller KV-head count becomes the binding limit. Hardware adds its own cap — the GPUs must share one fast interconnect domain.

open as a page

Should a server load a prequantized checkpoint or quantize weights at startup?

level: middleimportance: should knowfreq 45%

basics

~20 s

Prefer a prequantized checkpoint for production: calibration is already done, the layout matches a fast kernel, and startup stays short. Startup quantization is convenient for experiments but adds load time and usually lands on slower kernels — the exception is FP8, which quantizes online cheaply.

open as a page

When would you use n-gram prompt-lookup speculation instead of a separate draft model?

level: middleimportance: should knowfreq 42%

basics

~20 s

Use prompt lookup when the output largely copies the input — summarisation, document QA, code edits, structured rewrites. It costs no extra weights or VRAM and drafts almost instantly, but on open-ended generation with nothing to copy its acceptance rate collapses to near zero.

open as a page

Why is speculative decoding called lossless, and what exactly does it preserve?

level: middleimportance: should knowfreq 48%

basics

~20 s

The verification rule accepts a drafted token with probability capped by the ratio of the target's probability to the drafter's, and resamples from a corrected distribution otherwise. The result is distributed exactly as the target model alone — the drafter changes speed, not output quality.

open as a page

How do you route LoRA-adapter requests across a fleet of LLM replicas?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Route by adapter identity, not round-robin. A replica can only serve an adapter it currently has resident, and engines cap how many distinct adapters may appear in one batch, so hash adapters to replicas, keep each replica's hot set stable, and give a heavy tenant its own group.

open as a page

How does prefix-aware routing across LLM replicas raise cache hit rate?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A prefix cache lives inside one replica's own KV memory, so a shared preamble or a continuing conversation only skips prefill if the request lands on the replica that already holds it. Prefix-aware routing hashes the leading tokens and steers matching requests to the same replica.

open as a page

Your LLM server's waiting queue keeps growing under load — how do you diagnose the bottleneck?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Compare the running-sequence count against the configured concurrency cap and the KV-cache utilization gauge. Pinned at the cap with spare cache means the cap is throttling you; pinned cache means memory is the limit and no scheduler knob will fix it.

open as a page

showing 1–30 of 45