skip to content

vLLM

You will learn vLLM, the de-facto open-source inference engine: its PagedAttention memory manager, its OpenAI-compatible server, and the flags that decide throughput. Interviewers name it constantly because it is what most teams reach for when they stop paying per token.

on this pageshow

questions

18

In vLLM, how do you run offline batch inference with LLM and SamplingParams?

level: juniorimportance: must knowfreq 68%

answer

  1. one call, whole prompt list
  2. engine runs in your process
  3. scheduler batches, not you
  4. max_tokens default is small
  5. RequestOutput.outputs[0].text

basics

~20 s

Construct an LLM object with the model name, build a SamplingParams, then call llm.generate(prompts, params) with the whole prompt list. vLLM batches the list internally and returns one RequestOutput per prompt; the text sits in output.outputs[0].text.

solid answer

~50 s

`LLM` is vLLM's offline, in-process entry point: constructing it loads the weights onto the GPU and profiles memory, so it is expensive to create once and cheap to reuse. You hand `generate()` a **list** of prompts plus a `SamplingParams` (or a list of them, one per prompt) and vLLM's scheduler batches them continuously inside that single call — you never loop over prompts yourself. It returns a list of `RequestOutput` objects in input order; each carries `prompt`, `prompt_token_ids` and an `outputs` list of completions with `.text`, `.token_ids` and a finish reason. `SamplingParams` is where decoding settings live, and its `max_tokens` default is only 16, which is the classic reason a first offline run looks truncated. For chat models, `llm.chat(messages)` applies the model's chat template for you instead of making you format the prompt string by hand.

code

python · 10 lines
python
from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct")
params = SamplingParams(temperature=0.8, top_p=0.95, max_tokens=256)

prompts = ["Explain paged attention in one paragraph.", "Write a haiku about GPUs."]
outputs = llm.generate(prompts, params)

for out in outputs:
    print(out.prompt, "->", out.outputs[0].text, out.outputs[0].finish_reason)

go deeper

for a junior

Be able to write the four lines from memory: import LLM and SamplingParams, construct LLM(model=...), build params, call generate(prompts, params). Know that the text is at output.outputs[0].text.

for a middle

Explain why the whole prompt list goes into one call — the scheduler keeps the batch full — and name the max_tokens default as the reason a first run looks truncated. Know chat() applies the model's chat template and generate() does not.

for a senior

Show judgment about the process: one LLM per process, weights loaded once, no streaming, no metrics, and a crash costs the run. Talk about checking finish_reason across a batch job before trusting its output.

for a principal

Own the decision of when embedding the engine in a job beats standing up a shared endpoint, and what that costs the platform: an un-observable GPU consumer, no request-level SLOs, and retry/resume logic that becomes your team's problem.

## What the LLM class is vLLM has two ways in. `vllm serve` starts an HTTP server; the `LLM` class embeds the same engine directly in your Python process. There is no socket, no HTTP, no separate service — you import it, construct it, and call it. That makes `LLM` the right tool for offline work: scoring a dataset, generating synthetic data, running an eval harness, benchmarking a checkpoint. ``` from vllm import LLM, SamplingParams llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct") ``` Construction is the expensive part. It downloads or memory-maps the checkpoint, loads weights onto the GPU, runs a profiling forward pass to work out how much memory is left for the KV cache, and allocates that cache. Expect tens of seconds to minutes for a large model. The consequence for how you write code: create **one** `LLM` and keep it for the life of the process. Creating one per request, or one per worker on the same GPU, is the most common beginner mistake — the second object will fight the first for VRAM and usually fail outright. ## The three things you pass `generate()` takes prompts, sampling parameters, and (optionally) per-request extras such as LoRA requests. Prompts may be a single string or a list of strings; they may also be dicts carrying pre-tokenized ids or multimodal inputs, but a list of strings covers most work. `SamplingParams` is a plain value object holding the decoding configuration for a request: `temperature`, `top_p`, `top_k`, `max_tokens`, `n`, `stop`, `seed`, `logprobs` and so on. Pass one object and every prompt in the batch uses it; pass a list of the same length as the prompt list and each prompt gets its own. The parameter that trips people up is `max_tokens`, whose default is 16 — small enough that a first run returns a sentence fragment and looks like a broken model. What you do **not** control here is how requests are batched. There is no `batch_size` argument, because vLLM does not do fixed batches: the engine's scheduler decides, on every iteration, which of the outstanding sequences run in the next forward pass. ## Why you pass the whole list This is the point of the offline API and the thing interviewers actually probe. Handing all 50,000 prompts to one `generate()` call means the scheduler's queue is never empty: whenever a sequence finishes and frees its KV blocks, a waiting prompt is admitted in the same iteration. GPU occupancy stays high from start to finish. Calling `generate()` once per prompt in a Python `for` loop destroys that. Each call blocks until its single request completes, so the GPU decodes one sequence at a time — memory-bandwidth-bound work with a batch of one, which is the least efficient thing a serving GPU can do. The throughput difference between the two shapes of code is routinely an order of magnitude, and nothing about the API stops you from writing the slow one. `generate()` also has no streaming: it returns when the last request in the batch has finished. If you need incremental output, that is the asynchronous engine's job, or the server's. ## What comes back One `RequestOutput` per prompt, in the order you supplied them (the engine completes them out of order internally and reorders for you). Each has: - `prompt` and `prompt_token_ids` — what went in; - `outputs` — a list of `CompletionOutput`, one per requested sample, so `outputs[0]` unless you set `n > 1`; - per completion: `text`, `token_ids`, `cumulative_logprob`, optional `logprobs`, and `finish_reason` (`"stop"`, `"length"`, …); - `finished` and timing metrics. Always read `finish_reason` on a batch job. A run where most completions say `"length"` means `max_tokens` truncated them, not that the model chose to stop. ## Chat models Instruction-tuned checkpoints expect their own chat template — role markers and special tokens that the tokenizer config carries. `llm.chat(messages)` applies that template before generating, taking OpenAI-style message dicts. Passing raw user text to `generate()` on an instruct model skips the template and usually produces noticeably worse, rambling output. If you want the raw completion behaviour, that is exactly what `generate()` gives you. ## When it is the wrong tool One process owns the GPU. There is no HTTP surface, no metrics endpoint, no way for another team to send a request, no streaming, and if the process dies six hours into a run you lose the run. That is the trade the offline API makes in exchange for maximum throughput and zero deployment.

  • What actually goes wrong if you call generate() inside a Python for-loop, one prompt at a time?
    Each call blocks until that single request finishes, so the engine decodes with a batch of one. Decode is memory-bandwidth-bound, so a batch of one wastes almost all of the card's throughput, and you also pay the scheduler ramp-up repeatedly. Pass the full list instead and let the scheduler keep the batch full.
  • How do you give each prompt in a batch different sampling settings?
    Pass a list of SamplingParams the same length as the prompt list; element i applies to prompt i. That is how you mix, say, greedy extraction requests and higher-temperature creative requests in one run without splitting the job into two calls and losing batch density.
  • Can the offline LLM class stream tokens as they are produced?
    No. generate() returns only when every request in the call has finished, so there is no incremental callback. Streaming lives in the asynchronous engine, AsyncLLM, and in the HTTP server built on it. For an offline job that is rarely a loss — you want the whole batch fast, not the first token fast.
  • Where do sampling defaults come from if you leave SamplingParams fields unset?
    vLLM reads the checkpoint's own generation_config.json by default (the --generation-config option, whose default is auto), so a model that ships an opinionated temperature will use it. Pass --generation-config vllm to ignore the file and take vLLM's built-in defaults instead. Either way, explicit SamplingParams values win.

saying these in an interview costs you the question

  • Loops over prompts calling generate() once each to "batch" them
  • Expects generate() to stream tokens as they are produced
  • Blames the model when output stops at 16 tokens
  • Constructs a new LLM object per request or per worker
  • Thinks LLM() sends HTTP to a separate vllm serve process

context

open as a page

How do you point an OpenAI SDK client at a vLLM server instead of api.openai.com?

level: juniorimportance: must knowfreq 75%

basics

~10 s

Start the model with vllm serve, set the SDK's base_url to http://localhost:8000/v1, and send the model name the server lists at /v1/models. api_key can be any placeholder unless the server was started with --api-key.

open as a page

In vLLM, what does --gpu-memory-utilization control, and what does raising it buy?

level: middleimportance: must knowfreq 72%

basics

~20 s

--gpu-memory-utilization is the fraction of each GPU's total memory one vLLM instance may use, default 0.9. Weights and peak activations are subtracted first, and whatever is left becomes KV cache — so raising it buys concurrency and context, not speed.

open as a page

In vLLM, what does --max-model-len cap, and why can it block startup?

level: middleimportance: must knowfreq 66%

basics

~20 s

--max-model-len caps prompt plus generated tokens for a single request, defaulting to the model config's context length. vLLM refuses to start when its KV block pool cannot hold even one sequence that long, since it must be able to honour the limit it advertises.

open as a page

In vLLM, what does AsyncLLM give you that the offline LLM class does not?

level: middleimportance: must knowfreq 52%

basics

~20 s

AsyncLLM is vLLM's asynchronous engine: requests can be added at any moment while others are running, and its generate() is an async generator that yields partial results as tokens appear. The offline LLM class takes a fixed prompt list and returns only when all of it is finished.

open as a page

In vLLM, what breaks at /v1/chat/completions when a model ships no chat template?

level: middleimportance: must knowfreq 62%

basics

~20 s

Chat requests fail outright. vLLM renders the messages array into a single prompt using the Jinja chat template stored in the model's tokenizer_config.json; a base checkpoint has none, so you must supply one with --chat-template or use /v1/completions instead.

open as a page

In vLLM, which flags turn on OpenAI-style tool calling for a chat model?

level: middleimportance: must knowfreq 55%

basics

~20 s

Two flags together: --enable-auto-tool-choice and --tool-call-parser with the parser matching the served model's family. The parser converts the model's own tool-call text into the tool_calls field of the response; without it, calls arrive as ordinary content.

open as a page

What does a minimal vllm/vllm-openai docker run need to serve a Hub model?

level: juniorimportance: should knowfreq 56%

basics

~20 s

GPU access, a published port, host shared memory, and a mounted Hugging Face cache so weights survive restarts. Gated repos also need a Hub token in the environment. Engine arguments go after the image name, because the image already starts the OpenAI-compatible server.

open as a page

In vLLM, what does automatic prefix caching hash to match a reused prefix?

level: middleimportance: should knowfreq 48%

basics

~20 s

Each full KV block gets a chained hash: the previous block's hash combined with that block's token ids, plus keys such as the LoRA adapter id. A new request reuses a cached block only when the chain matches from token 0, so matching is prefix-only and block-granular.

open as a page

Deploying vLLM with --tensor-parallel-size 4 on Kubernetes: what must the pod provide?

level: seniorimportance: should knowfreq 46%

basics

~20 s

All four GPUs must be visible to one container in one pod, since vLLM's tensor-parallel workers are local processes sharing memory — not pods that find each other. The pod also needs a large /dev/shm, somewhere durable for weights, and probes that tolerate minutes of loading.

open as a page

Which vLLM /metrics series show whether the server is queueing or just slow?

level: seniorimportance: should knowfreq 50%

basics

~10 s

vllm:num_requests_waiting against vllm:num_requests_running separates queueing from execution; vllm:kv_cache_usage_perc and vllm:num_preemptions say whether memory is the cause. vllm:time_to_first_token_seconds and vllm:inter_token_latency_seconds give the latency that clients feel.

open as a page

In vLLM, when do you pass --quantization, and which values does it accept?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Usually you pass nothing: vLLM reads the checkpoint's own quantization config and picks a kernel. You pass --quantization to force a specific kernel path, to quantize an unquantized checkpoint at load time, or to name a method the checkpoint does not declare. Values come from a fixed list.

open as a page

In vLLM, what decides how many KV-cache blocks the engine allocates at startup?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A startup profiling run measures weights plus peak activation memory, subtracts that from the share of VRAM allowed by gpu_memory_utilization, and divides the remainder by the bytes one block costs. The quotient is the block count, fixed for the process; --num-gpu-blocks-override replaces it.

open as a page

What does vLLM's --api-key flag actually protect on the server?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only one thing: it requires callers to present a shared static token as an Authorization Bearer header on the API routes. There are no identities, scopes, quotas, rotation or rate limits, and operational routes outside the API prefix stay open.

open as a page

How do you serve several LoRA adapters from one vLLM server and select one per request?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Start with --enable-lora and register adapters via --lora-modules name=path. Each name appears in /v1/models, and a request selects one by putting that name in its model field. Adapters must all target the base checkpoint the server loaded.

open as a page

In vLLM, when is the offline LLM class the right tool instead of the server?

level: principalimportance: should knowfreq 31%

basics

~20 s

Use the in-process LLM class when one job owns the GPU and all the work is known up front: offline scoring, evals, dataset generation. Use vllm serve when several clients share the model, output must stream, or the deployment has to be operated and scaled independently.

open as a page

Migrating from a hosted OpenAI endpoint to vLLM, what still breaks despite API compatibility?

level: principalimportance: should knowfreq 38%

basics

~20 s

The protocol matches; the platform does not. Token counts shift with a different tokenizer, length limits become hard rejections, provider rate limits and retry semantics disappear, vendor-only parameters and surfaces have no equivalent, and the model itself is a different model.

open as a page

In vLLM V1, why does EngineCore run in its own process from the front end?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Splitting them keeps Python-side CPU work off the GPU's critical path. The front-end process handles HTTP, tokenization and detokenization while EngineCore — scheduler plus model executor — loops on the GPU without interruption, exchanging requests and outputs over local sockets.

open as a page