skip to content

Engine and PagedAttention

You will learn vLLM from the inside: the offline LLM API, the async engine loop that drives continuous batching, and the block manager that makes PagedAttention and prefix caching work. Interviewers ask how vLLM gets its throughput, and this is the honest answer.

on this pageshow

questions

6

In vLLM, how do you run offline batch inference with LLM and SamplingParams?

level: juniorimportance: must knowfreq 68%

answer

  1. one call, whole prompt list
  2. engine runs in your process
  3. scheduler batches, not you
  4. max_tokens default is small
  5. RequestOutput.outputs[0].text

basics

~20 s

Construct an LLM object with the model name, build a SamplingParams, then call llm.generate(prompts, params) with the whole prompt list. vLLM batches the list internally and returns one RequestOutput per prompt; the text sits in output.outputs[0].text.

solid answer

~50 s

`LLM` is vLLM's offline, in-process entry point: constructing it loads the weights onto the GPU and profiles memory, so it is expensive to create once and cheap to reuse. You hand `generate()` a **list** of prompts plus a `SamplingParams` (or a list of them, one per prompt) and vLLM's scheduler batches them continuously inside that single call — you never loop over prompts yourself. It returns a list of `RequestOutput` objects in input order; each carries `prompt`, `prompt_token_ids` and an `outputs` list of completions with `.text`, `.token_ids` and a finish reason. `SamplingParams` is where decoding settings live, and its `max_tokens` default is only 16, which is the classic reason a first offline run looks truncated. For chat models, `llm.chat(messages)` applies the model's chat template for you instead of making you format the prompt string by hand.

code

python · 10 lines
python
from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct")
params = SamplingParams(temperature=0.8, top_p=0.95, max_tokens=256)

prompts = ["Explain paged attention in one paragraph.", "Write a haiku about GPUs."]
outputs = llm.generate(prompts, params)

for out in outputs:
    print(out.prompt, "->", out.outputs[0].text, out.outputs[0].finish_reason)

go deeper

for a junior

Be able to write the four lines from memory: import LLM and SamplingParams, construct LLM(model=...), build params, call generate(prompts, params). Know that the text is at output.outputs[0].text.

for a middle

Explain why the whole prompt list goes into one call — the scheduler keeps the batch full — and name the max_tokens default as the reason a first run looks truncated. Know chat() applies the model's chat template and generate() does not.

for a senior

Show judgment about the process: one LLM per process, weights loaded once, no streaming, no metrics, and a crash costs the run. Talk about checking finish_reason across a batch job before trusting its output.

for a principal

Own the decision of when embedding the engine in a job beats standing up a shared endpoint, and what that costs the platform: an un-observable GPU consumer, no request-level SLOs, and retry/resume logic that becomes your team's problem.

## What the LLM class is vLLM has two ways in. `vllm serve` starts an HTTP server; the `LLM` class embeds the same engine directly in your Python process. There is no socket, no HTTP, no separate service — you import it, construct it, and call it. That makes `LLM` the right tool for offline work: scoring a dataset, generating synthetic data, running an eval harness, benchmarking a checkpoint. ``` from vllm import LLM, SamplingParams llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct") ``` Construction is the expensive part. It downloads or memory-maps the checkpoint, loads weights onto the GPU, runs a profiling forward pass to work out how much memory is left for the KV cache, and allocates that cache. Expect tens of seconds to minutes for a large model. The consequence for how you write code: create **one** `LLM` and keep it for the life of the process. Creating one per request, or one per worker on the same GPU, is the most common beginner mistake — the second object will fight the first for VRAM and usually fail outright. ## The three things you pass `generate()` takes prompts, sampling parameters, and (optionally) per-request extras such as LoRA requests. Prompts may be a single string or a list of strings; they may also be dicts carrying pre-tokenized ids or multimodal inputs, but a list of strings covers most work. `SamplingParams` is a plain value object holding the decoding configuration for a request: `temperature`, `top_p`, `top_k`, `max_tokens`, `n`, `stop`, `seed`, `logprobs` and so on. Pass one object and every prompt in the batch uses it; pass a list of the same length as the prompt list and each prompt gets its own. The parameter that trips people up is `max_tokens`, whose default is 16 — small enough that a first run returns a sentence fragment and looks like a broken model. What you do **not** control here is how requests are batched. There is no `batch_size` argument, because vLLM does not do fixed batches: the engine's scheduler decides, on every iteration, which of the outstanding sequences run in the next forward pass. ## Why you pass the whole list This is the point of the offline API and the thing interviewers actually probe. Handing all 50,000 prompts to one `generate()` call means the scheduler's queue is never empty: whenever a sequence finishes and frees its KV blocks, a waiting prompt is admitted in the same iteration. GPU occupancy stays high from start to finish. Calling `generate()` once per prompt in a Python `for` loop destroys that. Each call blocks until its single request completes, so the GPU decodes one sequence at a time — memory-bandwidth-bound work with a batch of one, which is the least efficient thing a serving GPU can do. The throughput difference between the two shapes of code is routinely an order of magnitude, and nothing about the API stops you from writing the slow one. `generate()` also has no streaming: it returns when the last request in the batch has finished. If you need incremental output, that is the asynchronous engine's job, or the server's. ## What comes back One `RequestOutput` per prompt, in the order you supplied them (the engine completes them out of order internally and reorders for you). Each has: - `prompt` and `prompt_token_ids` — what went in; - `outputs` — a list of `CompletionOutput`, one per requested sample, so `outputs[0]` unless you set `n > 1`; - per completion: `text`, `token_ids`, `cumulative_logprob`, optional `logprobs`, and `finish_reason` (`"stop"`, `"length"`, …); - `finished` and timing metrics. Always read `finish_reason` on a batch job. A run where most completions say `"length"` means `max_tokens` truncated them, not that the model chose to stop. ## Chat models Instruction-tuned checkpoints expect their own chat template — role markers and special tokens that the tokenizer config carries. `llm.chat(messages)` applies that template before generating, taking OpenAI-style message dicts. Passing raw user text to `generate()` on an instruct model skips the template and usually produces noticeably worse, rambling output. If you want the raw completion behaviour, that is exactly what `generate()` gives you. ## When it is the wrong tool One process owns the GPU. There is no HTTP surface, no metrics endpoint, no way for another team to send a request, no streaming, and if the process dies six hours into a run you lose the run. That is the trade the offline API makes in exchange for maximum throughput and zero deployment.

  • What actually goes wrong if you call generate() inside a Python for-loop, one prompt at a time?
    Each call blocks until that single request finishes, so the engine decodes with a batch of one. Decode is memory-bandwidth-bound, so a batch of one wastes almost all of the card's throughput, and you also pay the scheduler ramp-up repeatedly. Pass the full list instead and let the scheduler keep the batch full.
  • How do you give each prompt in a batch different sampling settings?
    Pass a list of SamplingParams the same length as the prompt list; element i applies to prompt i. That is how you mix, say, greedy extraction requests and higher-temperature creative requests in one run without splitting the job into two calls and losing batch density.
  • Can the offline LLM class stream tokens as they are produced?
    No. generate() returns only when every request in the call has finished, so there is no incremental callback. Streaming lives in the asynchronous engine, AsyncLLM, and in the HTTP server built on it. For an offline job that is rarely a loss — you want the whole batch fast, not the first token fast.
  • Where do sampling defaults come from if you leave SamplingParams fields unset?
    vLLM reads the checkpoint's own generation_config.json by default (the --generation-config option, whose default is auto), so a model that ships an opinionated temperature will use it. Pass --generation-config vllm to ignore the file and take vLLM's built-in defaults instead. Either way, explicit SamplingParams values win.

saying these in an interview costs you the question

  • Loops over prompts calling generate() once each to "batch" them
  • Expects generate() to stream tokens as they are produced
  • Blames the model when output stops at 16 tokens
  • Constructs a new LLM object per request or per worker
  • Thinks LLM() sends HTTP to a separate vllm serve process

context

open as a page

In vLLM, what does AsyncLLM give you that the offline LLM class does not?

level: middleimportance: must knowfreq 52%

basics

~20 s

AsyncLLM is vLLM's asynchronous engine: requests can be added at any moment while others are running, and its generate() is an async generator that yields partial results as tokens appear. The offline LLM class takes a fixed prompt list and returns only when all of it is finished.

open as a page

In vLLM, what does automatic prefix caching hash to match a reused prefix?

level: middleimportance: should knowfreq 48%

basics

~20 s

Each full KV block gets a chained hash: the previous block's hash combined with that block's token ids, plus keys such as the LoRA adapter id. A new request reuses a cached block only when the chain matches from token 0, so matching is prefix-only and block-granular.

open as a page

In vLLM, what decides how many KV-cache blocks the engine allocates at startup?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A startup profiling run measures weights plus peak activation memory, subtracts that from the share of VRAM allowed by gpu_memory_utilization, and divides the remainder by the bytes one block costs. The quotient is the block count, fixed for the process; --num-gpu-blocks-override replaces it.

open as a page

In vLLM, when is the offline LLM class the right tool instead of the server?

level: principalimportance: should knowfreq 31%

basics

~20 s

Use the in-process LLM class when one job owns the GPU and all the work is known up front: offline scoring, evals, dataset generation. Use vllm serve when several clients share the model, output must stream, or the deployment has to be operated and scaled independently.

open as a page

In vLLM V1, why does EngineCore run in its own process from the front end?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Splitting them keeps Python-side CPU work off the GPU's critical path. The front-end process handles HTTP, tokenization and detokenization while EngineCore — scheduler plus model executor — loops on the GPU without interruption, exchanging requests and outputs over local sockets.

open as a page