Serving Concepts
You will learn the engine-agnostic mechanics every inference server implements — batching, KV-cache management, quantization, speculation, latency budgets, GPU sizing, sharding, and scaling. Interviewers use these to test whether you understand why an LLM server behaves nothing like a stateless REST service.
on this pageshowhide
explore
- Continuous Batching5 questions
- KV Cache and Paged Attention6 questions
- Quantization for Serving5 questions
- Speculative Decoding6 questions
- TTFT, Throughput and SLOs6 questions
- GPU Memory Sizing and Cost5 questions
- Multi-GPU Parallelism6 questions
- Autoscaling and Model Routing6 questions
questions
page 2 of 2How do you tune an LLM server's per-iteration token budget and concurrent-sequence cap?
basics
~20 sSize the per-iteration token budget by how long a step may take, since step duration is every user's inter-token latency, and size the concurrent-sequence cap by how much KV cache you actually have. Then measure at your target load rather than guessing.
An LLM server ran fine for an hour, then died with CUDA out of memory. How do you triage?
basics
~20 sSuspect something outside the engine's own budget. Serving engines claim their KV pool at startup and then queue or preempt rather than OOM, so a steady-state failure usually means another process on the card, an activation spike from an unusually long prompt, or a utilization fraction set too high.
In a paged KV cache, what does raising the block size gain and cost?
basics
~20 sBigger blocks mean fewer table entries and longer contiguous reads, so kernels and bookkeeping get cheaper. The cost is coarser granularity: more wasted space in each sequence's partly-filled tail block, and prefix sharing that only matches in larger chunks.
What is goodput for an LLM server, and why report it over raw throughput?
basics
~20 sGoodput counts only the requests that finished inside the latency SLO — for example TTFT under 500 ms and per-output-token latency under 50 ms. Raw throughput counts a request that took 30 seconds as a success, so it can rise while every user's experience gets worse.
How does expert parallelism shard an MoE model for serving across GPUs?
basics
~20 sExpert parallelism places whole experts on different GPUs instead of slicing every matrix. Each MoE layer then does an all-to-all: tokens are dispatched to the GPUs owning their chosen experts and the results are gathered back. Uneven routing makes one GPU the straggler.
How do you split one large LLM across two 8-GPU nodes for serving?
basics
~20 sKeep tensor parallelism inside each node — TP=8 over the intra-node fabric — and cross the network with pipeline parallelism, PP=2. The node boundary is the slow link, and only pipeline parallelism sends little enough traffic to survive it. First check the model cannot simply fit on one node.
Why does speculative decoding stop paying off as an inference server's batch size grows?
basics
~20 sAt small batch a decode step has idle arithmetic capacity, so scoring extra drafted positions is nearly free. As concurrency rises the GPU's matrix units fill up, and the extra positions — most of the rejected ones pure waste — start costing real time and stealing throughput from other requests.
When is scale-to-zero right for a self-hosted LLM endpoint?
basics
~20 sScale to zero when idle hours dominate the bill and the first caller after an idle period can tolerate minutes — internal tools, batch pipelines, dev environments, rarely used per-tenant models. Never for interactive traffic with a time-to-first-token target, or where GPU capacity may not be reacquirable.
When does renting GPUs to self-host an open model beat paying per token?
basics
~20 sWhen sustained volume is high enough to keep GPUs genuinely busy, and when something other than price — data control, a custom checkpoint, predictable latency — is also on the table. Below that, per-token pricing wins because you are not paying for idle hardware.
Between chat turns, should the server keep a session's KV blocks resident or recompute?
basics
~20 sPinning blocks through a user's think time blocks capacity for every idle session, so it only pays for short gaps and high-value sessions. The usual answer is neither extreme: free the blocks but leave them in the prefix cache, where they are reused if still present and recomputed if not.
What belongs in an SLO for a streaming LLM endpoint, and why not just e2e p95?
basics
~20 sEnd-to-end latency scales with answer length, which the model and caller choose, so an e2e p95 target is mostly an SLO on how long answers are. Bound time to first token and per-output-token latency instead, each at a stated load level and length regime.
When do you add replicas instead of raising the tensor-parallel degree?
basics
~20 sDefault to replicas whenever the model fits on one GPU: independent copies add throughput linearly with no cross-GPU collectives and no shared failure domain. Shard only when the model does not fit, when per-token latency must drop, or when one replica needs a bigger KV-cache pool.
Do you standardize one quantization scheme across a serving fleet, or choose per model?
basics
~20 sStandardize per GPU generation, not across the whole fleet — hardware decides which schemes have fast kernels. Allow per-model exceptions only with evidence, because every model-and-scheme pair carries its own validation, artifact and re-validation-on-upgrade cost.
When would you split prefill and decode onto separate LLM server instances?
basics
~20 sWhen one continuously batched replica can no longer serve two SLOs at once. Disaggregation runs prefill workers and decode workers separately and ships the KV cache between them, so each can be sized, tuned and scaled independently — worth it only at cluster scale.
When is speculative decoding the wrong way to buy inference latency for a fleet?
basics
~20 sWhen the constraint is cost or capacity rather than per-request latency. Speculation raises total GPU work and consumes memory that limits concurrency, so a throughput-bound or budget-bound fleet is usually better served by more replicas, a smaller or quantized model, or better batching.
showing 31–45 of 45