skip to content

LLM Inference Serving

You will learn how open-weight models are actually served in production: the engines (vLLM, TGI, Triton/TensorRT-LLM) and the mechanics that decide whether a GPU costs you $200 or $20,000 a month. Interviewers ask because 'we self-host Llama' is now a normal answer, and it immediately raises batching, KV-cache, quantization, and GPU-sizing questions that most candidates cannot answer.

on this pageshow

explore

questions

93 · 4 sections

Why is GPU utilization a poor autoscaling signal for an LLM inference server?

level: middleimportance: must knowfreq 60%
basics
~20 s

GPU utilization only reports that a kernel was running during the sampling window. A server that batches continuously sits near 100% at low load and looks identical when its queue is backing up. Scale on waiting-request count, KV-cache occupancy and measured time-to-first-token instead.

open as a page

How does continuous batching differ from static batching in an LLM inference server?

level: middleimportance: must knowfreq 85%
basics
~20 s

Static batching forms a batch, runs it to completion, and only then accepts new work, so every slot waits for the longest generation. Continuous batching re-forms the batch before each decode step: finished sequences leave and queued requests join immediately.

open as a page

Will a 70B model fit on 2xA100-80GB at 32k context? How do you check?

level: middleimportance: must knowfreq 72%
basics
~20 s

Add the four terms up. 70B weights at bf16 are about 140 GB of the 160 GB on the node, so after CUDA context and activations only a few GB are left for KV cache. It loads, but it serves almost no concurrency.

open as a page

How does a PagedAttention block table let one sequence's KV cache be non-contiguous?

level: middleimportance: must knowfreq 66%
basics
~20 s

PagedAttention splits the KV cache into fixed-size blocks, each holding a set number of tokens. Every sequence keeps a block table listing which physical blocks hold its logical positions, and the attention kernel follows that table instead of striding through one range.

open as a page

Why does a contiguous per-request KV cache waste most of an LLM server's VRAM?

level: middleimportance: must knowfreq 62%
basics
~20 s

A contiguous allocator must reserve one slab per request sized for the longest output that request could produce. Most requests never grow into it, so the reserved tail sits idle, and the leftover gaps between slabs are too small for the next request to use.

open as a page

In vLLM, how do you run offline batch inference with LLM and SamplingParams?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Construct an LLM object with the model name, build a SamplingParams, then call llm.generate(prompts, params) with the whole prompt list. vLLM batches the list internally and returns one RequestOutput per prompt; the text sits in output.outputs[0].text.

open as a page

How do you point an OpenAI SDK client at a vLLM server instead of api.openai.com?

level: juniorimportance: must knowfreq 75%
basics
~10 s

Start the model with vllm serve, set the SDK's base_url to http://localhost:8000/v1, and send the model name the server lists at /v1/models. api_key can be any placeholder unless the server was started with --api-key.

open as a page

In vLLM, what does --gpu-memory-utilization control, and what does raising it buy?

level: middleimportance: must knowfreq 72%
basics
~20 s

--gpu-memory-utilization is the fraction of each GPU's total memory one vLLM instance may use, default 0.9. Weights and peak activations are subtracted first, and whatever is left becomes KV cache — so raising it buys concurrency and context, not speed.

open as a page

In vLLM, what does --max-model-len cap, and why can it block startup?

level: middleimportance: must knowfreq 66%
basics
~20 s

--max-model-len caps prompt plus generated tokens for a single request, defaulting to the model config's context length. vLLM refuses to start when its KV block pool cannot hold even one sequence that long, since it must be able to honour the limit it advertises.

open as a page

In vLLM, what does AsyncLLM give you that the offline LLM class does not?

level: middleimportance: must knowfreq 52%
basics
~20 s

AsyncLLM is vLLM's asynchronous engine: requests can be added at any moment while others are running, and its generate() is an async generator that yields partial results as tokens appear. The offline LLM class takes a fixed prompt list and returns only when all of it is finished.

open as a page

What does a minimal TGI docker run command need to serve a Hub model?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Give the container --gpus all so it sees the GPU, publish its port 80, mount a host volume at /data so downloaded weights survive restarts, and pass --model-id with the Hub repo id. Sharded runs also need --shm-size 1g.

open as a page

In TGI, what does --max-batch-prefill-tokens control, and what breaks if it is too high?

level: middleimportance: must knowfreq 58%
basics
~20 s

TGI's --max-batch-prefill-tokens caps how many prompt tokens the router packs into one prefill step. Raising it improves prompt-processing throughput but raises peak activation memory, so setting it too high produces CUDA out-of-memory crashes under concurrent long prompts.

open as a page

TGI fails to download a gated Llama repo — how do you fix the launch?

level: middleimportance: must knowfreq 55%
basics
~20 s

Accept the model's licence with the Hugging Face account that owns the token, then inject the token into the container at runtime as the HF_TOKEN environment variable. The download happens once at startup, so the failure appears in the launcher log before the server ever binds.

open as a page

How do TGI's /generate and /v1/chat/completions routes differ?

level: middleimportance: must knowfreq 58%
basics
~20 s

/generate is TGI's native route: you send raw prompt text under "inputs" with a "parameters" object and get back "generated_text". /v1/chat/completions is the Messages API: you send a "messages" array, the server applies the model's chat template, and the response is OpenAI-shaped.

open as a page

How does TGI's grammar parameter force a response to match a JSON schema?

level: middleimportance: should knowfreq 45%
basics
~20 s

TGI accepts a grammar object in a /generate request's parameters, either type json with a JSON Schema or type regex with a pattern. At each decoding step the server masks out any token that could not continue a valid match, so the output is guaranteed to parse.

open as a page

What is a Triton ensemble model, and what cost does it remove from a pipeline?

level: juniorimportance: must knowfreq 52%
basics
~20 s

A Triton ensemble is a scheduling-only model, declared with platform: "ensemble", that wires several real models into a pipeline inside the server. The client sends one request instead of three, and intermediate tensors never leave the server, removing network round-trips and serialization between stages.

open as a page

How must a Triton model repository be laid out on disk for a model to load?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A Triton model repository is a directory of model directories. Each model directory holds an optional config.pbtxt plus one or more numerically named version subdirectories, and the model file itself lives inside a version directory.

open as a page

What does trtllm-build produce in TensorRT-LLM, and why is that artifact not portable?

level: middleimportance: must knowfreq 60%
basics
~20 s

trtllm-build compiles a converted checkpoint into a serialized TensorRT engine: fused kernels, hardware-specific tactics and fixed shape limits baked into a binary. That binary is tied to the GPU architecture, the precision and the TensorRT-LLM version that built it.

open as a page

How do you convert a Hugging Face checkpoint into a TensorRT-LLM engine?

level: middleimportance: must knowfreq 45%
basics
~20 s

Two steps. A per-model convert_checkpoint.py script rewrites the Hugging Face weights into a TensorRT-LLM checkpoint directory — a config.json plus one weight file per rank, sharded by the --tp_size you choose. Then trtllm-build compiles that directory into an engine.

open as a page

In Triton, how do preferred_batch_size and max_queue_delay_microseconds shape dynamic batching?

level: middleimportance: must knowfreq 70%
basics
~20 s

Triton's dynamic batcher holds arriving requests briefly and merges them into one model execution. preferred_batch_size lists batch sizes worth executing immediately; max_queue_delay_microseconds caps how long a partial batch waits for more requests before it runs anyway.

open as a page