skip to content

In vLLM, when is the offline LLM class the right tool instead of the server?

level: principalimportance: should knowfreq 31%

answer

  1. who owns the GPU?
  2. one call keeps the queue full
  3. no metrics, no streaming, no abort
  4. weight load paid per process
  5. batch job as a client instead

basics

~20 s

Use the in-process LLM class when one job owns the GPU and all the work is known up front: offline scoring, evals, dataset generation. Use vllm serve when several clients share the model, output must stream, or the deployment has to be operated and scaled independently.

solid answer

~50 s

The technical argument for the offline path is queue density: handing the scheduler every prompt at once guarantees the batch is never starved, with no HTTP, no per-request serialization and no client-side concurrency tuning to get wrong. For a bounded job on a dedicated GPU that is the cheapest tokens you will buy. The argument against it is everything an operable service has: no streaming, no metrics endpoint, no per-request abort, no way for another team to send work, and a crash six hours into a run loses the run. So the split I draw is by ownership rather than by workload size — if exactly one job owns the card for a bounded period, embed the engine; if the model is a shared asset with an SLO, run the server and let the batch job be one more concurrent client. The batch job then pays a little HTTP overhead and gains observability, retries and fleet-level sharing, which for most platforms is the better trade.

go deeper

for a junior

Know there are two ways to use vLLM — a Python class you embed and an HTTP server you deploy — and that streaming and multi-client use require the server.

for a middle

Explain why handing the whole prompt list to the offline API keeps the scheduler saturated, and list what the server adds: streaming, metrics, abort, and a shared surface for other callers.

for a senior

Weigh the operational side: weight-load cost paid per process, no resumability in an embedded job, and the VRAM contention you create by co-locating an embedded job with a serving replica.

for a principal

Set the policy. Decide whether the model is a shared platform asset or a per-job resource, how batch traffic is isolated from interactive traffic, and who owns the resumability and quality gates that an embedded job must supply itself.

## The two paths vLLM's engine can be embedded in a Python process (`LLM`) or fronted by an HTTP server (`vllm serve`). Both run the same scheduler, the same paged KV cache and the same kernels. The choice is not about speed of the model; it is about who owns the GPU and what the surrounding system needs. ## The case for embedding **Queue density is free.** The offline API takes the entire prompt list in one call, so from the first iteration to the last the scheduler always has waiting work to admit whenever a sequence finishes. You get maximum-throughput behaviour without writing any concurrency control. The equivalent through the server requires a client that keeps hundreds of requests in flight, handles backpressure, and does not accidentally serialize on a connection pool. **No per-request overhead.** No HTTP parsing, no JSON, no socket. At small payload sizes and very high request counts that overhead is measurable. **Nothing to deploy.** One process, one container, one lifetime. For an eval that runs in CI or a nightly scoring job, standing up a service and then tearing it down is machinery you do not need. **Cost.** A bounded job on a dedicated or spot GPU with no idle time is the cheapest way to buy tokens from open weights, because you pay for the card exactly as long as there is work. ## The case for the server **Sharing.** A GPU costs real money per hour. If one team's batch job owns the card exclusively, everyone else waits or you buy another card. The server turns the model into a shared asset. **Observability.** A metrics endpoint, per-request logging, queue depth and preemption counters — none of which the embedded path gives you. An embedded job that slows down is opaque; a server that slows down tells you whether it is queueing or decoding. **Streaming and abort.** Anything user-facing needs incremental output, and anything user-facing needs a way to stop a request whose client left. Neither exists in the offline API. **Operational independence.** Weight loading and KV profiling take minutes. A long-lived server pays that once; a per-job process pays it every run, which for a job that runs every ten minutes is most of the wall clock. ## Where the line falls I draw it at ownership, not at size: - **One job, bounded duration, dedicated GPU** → embed. A 50-million-token backfill on a spot H100 wants nothing between it and the scheduler. - **Model is a shared asset with an SLO** → server, always. Then the batch job is a concurrent client of it, with a concurrency cap that leaves interactive traffic its headroom. - **Mixed** → server, with a separate replica or node pool for batch. Isolation by replica rather than by process is easier to reason about, easier to autoscale, and easier to bill. The common failure I have seen is the middle case handled badly: a batch job embedding its own engine on a node that already runs a serving replica. Two processes now contend for the same VRAM, and vLLM's memory fraction is computed against *total* card memory, not free memory, so the second one either fails at startup or silently starves the first of KV blocks. ## The operational realities of a long embedded run If you do embed for a long job, own these: - **Resumability.** Write outputs incrementally, keyed by input id, and make the job skip what is already written. Without this a preemption at hour five costs five hours. - **Sharding.** Split the prompt set across workers, one GPU each. There is no shared scheduler across processes, so the split has to be static. - **Warm-up amortization.** Weight load plus profiling is paid per process. If the job is short relative to that, the server was the right answer. - **Finish reasons.** A batch job silently truncating at `max_tokens` produces a dataset that looks fine and is not. Aggregate finish reasons as a quality gate. ## The answer that lands Name the real trade — throughput and simplicity against observability and sharing — and then say which failure you are optimizing against. "We embed for the nightly backfill because it owns its own spot node and finishes in ninety minutes; everything else goes through the shared endpoint so we have one place to see latency and one place to enforce quotas" is a complete answer. "Offline is for demos" is not.

  • Your embedded batch job takes six hours and runs on a preemptible node. What changes in the design?
    It has to become resumable: write outputs incrementally keyed by input id and skip completed inputs on restart, shard the prompt set across workers so a lost worker costs a fraction, and checkpoint often enough that a preemption costs minutes rather than hours. The offline API gives you no persistence, so all of that is yours to build.
  • Could you reach the same throughput by pointing many concurrent clients at vllm serve instead?
    Close to it, provided concurrency is high enough to keep the scheduler saturated and the client does not serialize on its connection pool. You pay HTTP and tokenization overhead and you own the backpressure logic, and you gain metrics, sharing and the ability to cap the batch job so interactive traffic keeps its headroom. For most platforms that trade is worth a few percent of throughput.
  • Why is running an embedded batch job on the same GPU as a serving replica a bad idea?
    vLLM sizes its KV cache from a fraction of total card memory measured at startup, not from free memory. The second process either fails to allocate blocks or succeeds while starving the first, and neither one knows about the other's preemption pressure. If they must share hardware, share by node with separate GPUs, not by GPU.
  • How do you decide whether a nightly eval belongs in the offline API or the shared endpoint?
    Ask how long the run is relative to weight-load and profiling time, and whether the eval's results must be comparable across model versions. A short eval pays proportionally huge startup cost, so it belongs on a warm shared endpoint; a long one on its own node is cheaper embedded and gives you a clean, uncontended measurement.

saying these in an interview costs you the question

  • Treats the offline API as demo-only, never production
  • Assumes HTTP adds no meaningful per-request cost
  • Puts an embedded job on a GPU already serving traffic
  • Ignores weight-load time when launching per-job processes
  • Thinks high throughput requires the OpenAI endpoint

context