In vLLM, what does AsyncLLM give you that the offline LLM class does not?
answer
- online arrival versus fixed list
- async generator, not a return value
- abort by request id
- V1 class, old name kept
- same scheduler underneath
basics
~20 sAsyncLLM is vLLM's asynchronous engine: requests can be added at any moment while others are running, and its generate() is an async generator that yields partial results as tokens appear. The offline LLM class takes a fixed prompt list and returns only when all of it is finished.
solid answer
~50 s`AsyncLLM` (in `vllm.v1.engine.async_llm`) is the engine the HTTP server is built on. Two things make it different from `LLM`: requests arrive over time rather than as one list, and `generate()` is an `async` generator that yields `RequestOutput` objects as the engine produces them, which is what makes SSE streaming possible. It also supports aborting a request by id, so a client disconnect frees that sequence's KV blocks instead of paying to decode tokens nobody will read. One naming point worth getting right in vLLM 0.27: the old `AsyncLLMEngine` name still imports, but its own docstring says it is an alias of `vllm.v1.engine.async_llm.AsyncLLM` — vLLM is V1-only now, so name `AsyncLLM` and mention the alias. You rarely construct it directly; you get it by running `vllm serve`, or by embedding it when you need your own request front end.
code
python · 6 linesfrom vllm import AsyncLLMEngine
from vllm.v1.engine.async_llm import AsyncLLM
# In vLLM 0.27 the historical name resolves to the V1 engine class.
print(issubclass(AsyncLLMEngine, AsyncLLM))
print(AsyncLLMEngine.__mro__)go deeper
Know there are two entry points: LLM for a fixed batch of prompts, and an asynchronous engine behind the server for streaming. Say plainly that streaming responses come from the async one.
Explain the async generator shape, the per-request id, and that AsyncLLMEngine is now just an alias for AsyncLLM in a V1-only vLLM. Be able to say why both paths batch continuously.
Talk about abort on disconnect and why an unaborted stream burns batch slots and KV blocks. Be ready to say when you would embed AsyncLLM directly rather than run the packaged server, and what you take on by doing it.
Own the boundary question: does your platform expose the engine through vLLM's own server or through a house front end on AsyncLLM? The second buys protocol control and costs you metrics, templating, error mapping and abort semantics you must then maintain.
## Two entry points, one engine vLLM exposes the same core engine through two front ends. `LLM` is synchronous and offline: you give it every prompt up front, it blocks, it returns everything. `AsyncLLM` is asynchronous and online: requests are submitted individually at arbitrary times, each gets its own stream of incremental outputs, and any of them can be cancelled mid-flight. The scheduler underneath is identical. Continuous batching is *not* something `AsyncLLM` adds — the offline path batches too. What `AsyncLLM` adds is a request lifecycle that matches a server: open-ended arrival, per-request streaming, per-request abort. ## The alias, and why it matters If you have read older vLLM code or older interview material, the class you know is `AsyncLLMEngine`. In vLLM 0.27 that symbol still imports and still works, but it is an alias for the V1 class `AsyncLLM`; the docstring says so in as many words. The V0 engine it originally belonged to is gone — 0.27 is V1-only. Saying "vLLM runs on AsyncLLMEngine" is not wrong enough to be a hard error, but it dates you by about a year, and any question about the engine's internals will immediately go somewhere V0 no longer describes. Name `AsyncLLM`, and say the old name is retained as an alias. ## The output stream `AsyncLLM.generate(prompt, sampling_params, request_id)` returns an async generator. You iterate it with `async for`, and each iteration yields a `RequestOutput` reflecting the engine's progress on that request: text produced so far (or the delta, depending on how you configure it), token ids, logprobs if requested, and a `finished` flag with a finish reason on the last one. Note what it does *not* yield: single tokens. An engine step can emit more than one token for a request — speculative decoding accepts several drafted tokens per target forward pass, for instance — so the unit of yield is "whatever the last engine step produced for you", not "one token". Code that assumes one yield equals one token will silently miscount, and any UI that renders per-yield chunks handles this naturally anyway. The `request_id` you pass is the handle the whole system uses: the front end tags outputs with it to route them back to the right client stream, the logs carry it, and abort takes it. ## Aborts and disconnects This is the operationally important capability. In an HTTP server, clients close connections — a user navigates away, a proxy times out, a retry supersedes an in-flight call. Without abort, the engine would keep decoding that sequence to its stop condition, holding KV blocks and occupying batch slots that paying requests need. `AsyncLLM` lets the front end abort by request id when the stream's consumer goes away; the scheduler drops the sequence and frees its blocks on the next step. Under load, this is the difference between a server that degrades gracefully and one where a wave of client timeouts leaves the GPU busy generating garbage for nobody. ## Where you actually meet it Most people never construct `AsyncLLM` themselves. `vllm serve` builds it and wraps it in the OpenAI-compatible HTTP layer; that is the path production uses, and it comes with metrics, an API surface, and all the request handling already written. You reach for `AsyncLLM` directly when you need a *different* front end: a gRPC service, a custom protocol, an agent runtime that wants engine-level control over request admission, or a research harness that submits and cancels requests dynamically. The cost of taking that route is that you now own everything the server was giving you free — tokenization policy, chat templating, error mapping, metrics, and the abort-on-disconnect logic described above. ## Choosing between them The honest rule: if the set of work is known up front and one process owns the GPU, use `LLM` — less machinery, no per-request overhead, maximum throughput. If work arrives over time from multiple callers, or anything needs to see tokens before the response is complete, you need the async path, and almost always you want the server rather than the raw class. A nuance worth being able to state: choosing `AsyncLLM` does not make the GPU faster. Both paths run the same scheduler and reach similar throughput when the queue is equally full. What differs is the shape of the interaction, and the operational surface you get with it.
- Why does the async generator yield RequestOutput objects rather than one token at a time?Because an engine step does not always produce exactly one token for a request — speculative decoding can accept several accepted tokens in a single target pass, and the yield also has to carry logprobs, the finish reason and the finished flag. Yielding an output object per step is honest about that; yielding "a token" would not be.
- What happens to a running request when the HTTP client disconnects mid-stream?The front end aborts it by request id. The scheduler drops the sequence on the next step and returns its KV blocks to the pool, so the freed capacity goes to requests someone is still waiting on. Without that, the engine would decode the abandoned sequence to its stop condition while holding a batch slot.
- Is it a mistake to write AsyncLLMEngine in new vLLM code?It still works — the name is kept as an alias — but write AsyncLLM. The alias exists for compatibility with pre-V1 code, and using it in new work signals you are reasoning about an engine architecture that no longer exists in 0.27. If you inherit code that uses it, treat it as a rename, not a migration.
- Does switching from LLM to AsyncLLM increase throughput?No. Both drive the same scheduler with the same continuous batching, so with an equally full queue they land in the same place. AsyncLLM changes the interaction model — arrival over time, streaming, abort — not the GPU efficiency. If anything the async path pays a little more CPU per request for the front-end work.
saying these in an interview costs you the question
- Thinks AsyncLLMEngine is a different, older engine still in use
- Believes continuous batching only happens on the async path
- Assumes each yielded output equals exactly one token
- Thinks "async" means several models or GPU multi-threading
- Never mentions aborting requests when clients disconnect