LLM Inference Serving
You will learn how open-weight models are actually served in production: the engines (vLLM, TGI, Triton/TensorRT-LLM) and the mechanics that decide whether a GPU costs you $200 or $20,000 a month. Interviewers ask because 'we self-host Llama' is now a normal answer, and it immediately raises batching, KV-cache, quantization, and GPU-sizing questions that most candidates cannot answer.
on this pageshowhide
explore
- Serving Concepts45 questions
- Continuous Batching5 questions
- KV Cache and Paged Attention6 questions
- Quantization for Serving5 questions
- Speculative Decoding6 questions
- TTFT, Throughput and SLOs6 questions
- GPU Memory Sizing and Cost5 questions
- Multi-GPU Parallelism6 questions
- Autoscaling and Model Routing6 questions
- vLLM18 questions
- Engine and PagedAttention6 questions
- OpenAI-Compatible Server6 questions
- Deployment and Tuning6 questions
- Text Generation Inference (TGI)12 questions
- Running the TGI Server6 questions
- Batching and Performance Features6 questions
- NVIDIA Triton and TensorRT-LLM18 questions
- Model Repository and Backends6 questions
- Batching, Instances and Ensembles6 questions
- TensorRT-LLM Engine Build6 questions
- KServeempty
questions
93 · 4 sectionsWhy is GPU utilization a poor autoscaling signal for an LLM inference server?
basics
~20 sGPU utilization only reports that a kernel was running during the sampling window. A server that batches continuously sits near 100% at low load and looks identical when its queue is backing up. Scale on waiting-request count, KV-cache occupancy and measured time-to-first-token instead.
How does continuous batching differ from static batching in an LLM inference server?
basics
~20 sStatic batching forms a batch, runs it to completion, and only then accepts new work, so every slot waits for the longest generation. Continuous batching re-forms the batch before each decode step: finished sequences leave and queued requests join immediately.
Will a 70B model fit on 2xA100-80GB at 32k context? How do you check?
basics
~20 sAdd the four terms up. 70B weights at bf16 are about 140 GB of the 160 GB on the node, so after CUDA context and activations only a few GB are left for KV cache. It loads, but it serves almost no concurrency.
How does a PagedAttention block table let one sequence's KV cache be non-contiguous?
basics
~20 sPagedAttention splits the KV cache into fixed-size blocks, each holding a set number of tokens. Every sequence keeps a block table listing which physical blocks hold its logical positions, and the attention kernel follows that table instead of striding through one range.
Why does a contiguous per-request KV cache waste most of an LLM server's VRAM?
basics
~20 sA contiguous allocator must reserve one slab per request sized for the longest output that request could produce. Most requests never grow into it, so the reserved tail sits idle, and the leftover gaps between slabs are too small for the next request to use.
In vLLM, how do you run offline batch inference with LLM and SamplingParams?
basics
~20 sConstruct an LLM object with the model name, build a SamplingParams, then call llm.generate(prompts, params) with the whole prompt list. vLLM batches the list internally and returns one RequestOutput per prompt; the text sits in output.outputs[0].text.
How do you point an OpenAI SDK client at a vLLM server instead of api.openai.com?
basics
~10 sStart the model with vllm serve, set the SDK's base_url to http://localhost:8000/v1, and send the model name the server lists at /v1/models. api_key can be any placeholder unless the server was started with --api-key.
In vLLM, what does --gpu-memory-utilization control, and what does raising it buy?
basics
~20 s--gpu-memory-utilization is the fraction of each GPU's total memory one vLLM instance may use, default 0.9. Weights and peak activations are subtracted first, and whatever is left becomes KV cache — so raising it buys concurrency and context, not speed.
In vLLM, what does --max-model-len cap, and why can it block startup?
basics
~20 s--max-model-len caps prompt plus generated tokens for a single request, defaulting to the model config's context length. vLLM refuses to start when its KV block pool cannot hold even one sequence that long, since it must be able to honour the limit it advertises.
In vLLM, what does AsyncLLM give you that the offline LLM class does not?
basics
~20 sAsyncLLM is vLLM's asynchronous engine: requests can be added at any moment while others are running, and its generate() is an async generator that yields partial results as tokens appear. The offline LLM class takes a fixed prompt list and returns only when all of it is finished.
What does a minimal TGI docker run command need to serve a Hub model?
basics
~20 sGive the container --gpus all so it sees the GPU, publish its port 80, mount a host volume at /data so downloaded weights survive restarts, and pass --model-id with the Hub repo id. Sharded runs also need --shm-size 1g.
In TGI, what does --max-batch-prefill-tokens control, and what breaks if it is too high?
basics
~20 sTGI's --max-batch-prefill-tokens caps how many prompt tokens the router packs into one prefill step. Raising it improves prompt-processing throughput but raises peak activation memory, so setting it too high produces CUDA out-of-memory crashes under concurrent long prompts.
TGI fails to download a gated Llama repo — how do you fix the launch?
basics
~20 sAccept the model's licence with the Hugging Face account that owns the token, then inject the token into the container at runtime as the HF_TOKEN environment variable. The download happens once at startup, so the failure appears in the launcher log before the server ever binds.
How do TGI's /generate and /v1/chat/completions routes differ?
basics
~20 s/generate is TGI's native route: you send raw prompt text under "inputs" with a "parameters" object and get back "generated_text". /v1/chat/completions is the Messages API: you send a "messages" array, the server applies the model's chat template, and the response is OpenAI-shaped.
How does TGI's grammar parameter force a response to match a JSON schema?
basics
~20 sTGI accepts a grammar object in a /generate request's parameters, either type json with a JSON Schema or type regex with a pattern. At each decoding step the server masks out any token that could not continue a valid match, so the output is guaranteed to parse.
What is a Triton ensemble model, and what cost does it remove from a pipeline?
basics
~20 sA Triton ensemble is a scheduling-only model, declared with platform: "ensemble", that wires several real models into a pipeline inside the server. The client sends one request instead of three, and intermediate tensors never leave the server, removing network round-trips and serialization between stages.
How must a Triton model repository be laid out on disk for a model to load?
basics
~20 sA Triton model repository is a directory of model directories. Each model directory holds an optional config.pbtxt plus one or more numerically named version subdirectories, and the model file itself lives inside a version directory.
What does trtllm-build produce in TensorRT-LLM, and why is that artifact not portable?
basics
~20 strtllm-build compiles a converted checkpoint into a serialized TensorRT engine: fused kernels, hardware-specific tactics and fixed shape limits baked into a binary. That binary is tied to the GPU architecture, the precision and the TensorRT-LLM version that built it.
How do you convert a Hugging Face checkpoint into a TensorRT-LLM engine?
basics
~20 sTwo steps. A per-model convert_checkpoint.py script rewrites the Hugging Face weights into a TensorRT-LLM checkpoint directory — a config.json plus one weight file per rank, sharded by the --tp_size you choose. Then trtllm-build compiles that directory into an engine.
In Triton, how do preferred_batch_size and max_queue_delay_microseconds shape dynamic batching?
basics
~20 sTriton's dynamic batcher holds arriving requests briefly and merges them into one model execution. preferred_batch_size lists batch sizes worth executing immediately; max_queue_delay_microseconds caps how long a partial batch waits for more requests before it runs anyway.