skip to content

Serving Concepts

You will learn the engine-agnostic mechanics every inference server implements — batching, KV-cache management, quantization, speculation, latency budgets, GPU sizing, sharding, and scaling. Interviewers use these to test whether you understand why an LLM server behaves nothing like a stateless REST service.

on this pageshow

explore

questions

page 2 of 2

How do you tune an LLM server's per-iteration token budget and concurrent-sequence cap?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Size the per-iteration token budget by how long a step may take, since step duration is every user's inter-token latency, and size the concurrent-sequence cap by how much KV cache you actually have. Then measure at your target load rather than guessing.

open as a page

An LLM server ran fine for an hour, then died with CUDA out of memory. How do you triage?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Suspect something outside the engine's own budget. Serving engines claim their KV pool at startup and then queue or preempt rather than OOM, so a steady-state failure usually means another process on the card, an activation spike from an unusually long prompt, or a utilization fraction set too high.

open as a page

In a paged KV cache, what does raising the block size gain and cost?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Bigger blocks mean fewer table entries and longer contiguous reads, so kernels and bookkeeping get cheaper. The cost is coarser granularity: more wasted space in each sequence's partly-filled tail block, and prefix sharing that only matches in larger chunks.

open as a page

What is goodput for an LLM server, and why report it over raw throughput?

level: seniorimportance: should knowfreq 43%

basics

~20 s

Goodput counts only the requests that finished inside the latency SLO — for example TTFT under 500 ms and per-output-token latency under 50 ms. Raw throughput counts a request that took 30 seconds as a success, so it can rise while every user's experience gets worse.

open as a page

How does expert parallelism shard an MoE model for serving across GPUs?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Expert parallelism places whole experts on different GPUs instead of slicing every matrix. Each MoE layer then does an all-to-all: tokens are dispatched to the GPUs owning their chosen experts and the results are gathered back. Uneven routing makes one GPU the straggler.

open as a page

How do you split one large LLM across two 8-GPU nodes for serving?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Keep tensor parallelism inside each node — TP=8 over the intra-node fabric — and cross the network with pipeline parallelism, PP=2. The node boundary is the slow link, and only pipeline parallelism sends little enough traffic to survive it. First check the model cannot simply fit on one node.

open as a page

Why does speculative decoding stop paying off as an inference server's batch size grows?

level: seniorimportance: should knowfreq 52%

basics

~20 s

At small batch a decode step has idle arithmetic capacity, so scoring extra drafted positions is nearly free. As concurrency rises the GPU's matrix units fill up, and the extra positions — most of the rejected ones pure waste — start costing real time and stealing throughput from other requests.

open as a page

When is scale-to-zero right for a self-hosted LLM endpoint?

level: principalimportance: should knowfreq 38%

basics

~20 s

Scale to zero when idle hours dominate the bill and the first caller after an idle period can tolerate minutes — internal tools, batch pipelines, dev environments, rarely used per-tenant models. Never for interactive traffic with a time-to-first-token target, or where GPU capacity may not be reacquirable.

open as a page

When does renting GPUs to self-host an open model beat paying per token?

level: principalimportance: should knowfreq 44%

basics

~20 s

When sustained volume is high enough to keep GPUs genuinely busy, and when something other than price — data control, a custom checkpoint, predictable latency — is also on the table. Below that, per-token pricing wins because you are not paying for idle hardware.

open as a page

Between chat turns, should the server keep a session's KV blocks resident or recompute?

level: principalimportance: should knowfreq 34%

basics

~20 s

Pinning blocks through a user's think time blocks capacity for every idle session, so it only pays for short gaps and high-value sessions. The usual answer is neither extreme: free the blocks but leave them in the prefix cache, where they are reused if still present and recomputed if not.

open as a page

What belongs in an SLO for a streaming LLM endpoint, and why not just e2e p95?

level: principalimportance: should knowfreq 44%

basics

~20 s

End-to-end latency scales with answer length, which the model and caller choose, so an e2e p95 target is mostly an SLO on how long answers are. Bound time to first token and per-output-token latency instead, each at a stated load level and length regime.

open as a page

When do you add replicas instead of raising the tensor-parallel degree?

level: principalimportance: should knowfreq 52%

basics

~20 s

Default to replicas whenever the model fits on one GPU: independent copies add throughput linearly with no cross-GPU collectives and no shared failure domain. Shard only when the model does not fit, when per-token latency must drop, or when one replica needs a bigger KV-cache pool.

open as a page

Do you standardize one quantization scheme across a serving fleet, or choose per model?

level: principalimportance: should knowfreq 35%

basics

~20 s

Standardize per GPU generation, not across the whole fleet — hardware decides which schemes have fast kernels. Allow per-model exceptions only with evidence, because every model-and-scheme pair carries its own validation, artifact and re-validation-on-upgrade cost.

open as a page

When would you split prefill and decode onto separate LLM server instances?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

When one continuously batched replica can no longer serve two SLOs at once. Disaggregation runs prefill workers and decode workers separately and ships the KV cache between them, so each can be sized, tuned and scaled independently — worth it only at cluster scale.

open as a page

When is speculative decoding the wrong way to buy inference latency for a fleet?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

When the constraint is cost or capacity rather than per-request latency. Speculation raises total GPU work and consumes memory that limits concurrency, so a throughput-bound or budget-bound fleet is usually better served by more replicas, a smaller or quantized model, or better batching.

open as a page

showing 31–45 of 45