skip to content

OpenAI-Compatible Server

You will learn to stand up vllm serve as a drop-in OpenAI endpoint — same routes, same SDKs, plus tool calling, chat templates, and hot-swappable LoRA adapters. Interviewers care because API compatibility is what makes migrating off a paid provider a config change instead of a rewrite.

on this pageshow

questions

6

How do you point an OpenAI SDK client at a vLLM server instead of api.openai.com?

level: juniorimportance: must knowfreq 75%

answer

  1. Configuration change, not a code change
  2. Three client settings, nothing else
  3. The SDK appends the route to base_url
  4. Model string must match /v1/models
  5. --served-model-name aliases the provider id

basics

~10 s

Start the model with vllm serve, set the SDK's base_url to http://localhost:8000/v1, and send the model name the server lists at /v1/models. api_key can be any placeholder unless the server was started with --api-key.

solid answer

~50 s

`vllm serve meta-llama/Llama-3.1-8B-Instruct` loads the model and starts an HTTP server on port 8000 that publishes `/v1/chat/completions`, `/v1/completions` and `/v1/models` with the same request and response JSON a hosted OpenAI-style endpoint uses. Only three client settings change. **base_url** must include the `/v1` prefix, because the SDK appends the route to it. **model** must match a name the server reports at `/v1/models` — by default that is the exact model string you launched with, and `--served-model-name` lets you alias it to whatever id the client already hardcodes. **api_key** is a non-empty placeholder that vLLM ignores unless you started it with `--api-key`. Streaming, `usage`, `finish_reason` and tool-call shapes follow the same schemas, so the SDK's `stream=True` path works unchanged. What does not come across is the account layer: no billing, no provider-side rate limits, no files or batch endpoints.

code

bash · 3 lines
bash
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --port 8000 \
  --served-model-name gpt-4o-mini

go deeper

for a junior

Be able to name the three settings that change — base_url ending in /v1, a model name from /v1/models, and a placeholder api_key — and to say that vLLM runs the model locally rather than proxying anywhere.

for a middle

Explain why the SDK works at all: the routes and JSON schemas match, so the client is unchanged. Mention --served-model-name and that /v1/models is the authoritative list of what this process will accept.

for a senior

Show that you know where compatibility stops: no billing, no provider quotas, no managed uptime, one model per process. Talk about health probes, readiness, and how client timeout and retry policy must change when overload becomes queueing.

for a principal

Own the boundary decision: which cross-cutting concerns (auth, quotas, routing, fallback to a hosted provider) stay in a gateway so that swapping between self-hosted and hosted models really is a config flip for every consuming team.

## The server behind the routes `vllm serve <model>` loads the weights onto the GPU and starts an HTTP server (FastAPI on uvicorn) listening on port 8000 by default. The routes it publishes are deliberately the ones a hosted OpenAI-style provider publishes: `POST /v1/chat/completions`, `POST /v1/completions`, `GET /v1/models`, and — when the loaded model is an embedding model — `POST /v1/embeddings`. Beside them sit routes that are vLLM's own and not part of that API surface: `GET /health` for readiness probes, `GET /metrics` for Prometheus scraping, and `POST /tokenize` / `POST /detokenize` for inspecting how text becomes tokens. Compatibility means the JSON shapes match, not just the paths. You send `messages` with `role` and `content`; you get back `choices[0].message.content`, a `finish_reason`, and a `usage` object with prompt and completion token counts. With `"stream": true` you get a `text/event-stream` response whose chunk objects are `chat.completion.chunk` and whose final line is `data: [DONE]`. An SDK written for the hosted API is just an HTTP client for those shapes, so it drives vLLM without a rewrite. That is the whole point of the compatibility layer: migrating off a paid provider becomes a configuration change rather than a code change. ## The three client settings that change **base_url.** The OpenAI SDKs append the route path to whatever base URL you configure, so the value must end at `/v1` — `http://localhost:8000/v1`. Point it at `http://localhost:8000` and the client will request `/chat/completions`, which the server does not serve, and the 404 will look like a broken server rather than a misconfigured client. **model.** The hosted-provider habit of passing any model id and getting a model back does not apply. A vLLM process serves exactly one model, and the `model` field is validated against the names it advertises. Query `GET /v1/models` to see them; the default name is the exact string you passed to `vllm serve` (a Hugging Face repo id or a local directory path). Send anything else and the request is rejected. **api_key.** The SDKs refuse to construct a client with an empty key, so pass a placeholder such as `"not-used"`. vLLM only checks the value if the server was started with `--api-key` (or the `VLLM_API_KEY` environment variable), in which case the client's key must match that shared token exactly. ## The served model name is a real contract `--served-model-name` decouples the public name from the checkpoint on disk. If your application already hardcodes a provider model string, launching with `--served-model-name` set to that string means literally nothing in the application changes except the base URL. It is also how you keep a stable public name while swapping the underlying checkpoint — clients keep calling the same name across a quantized rebuild or a version bump. When LoRA adapters are enabled, their names show up in `/v1/models` too, and are addressed through the same `model` field. ## What does not come along The protocol is compatible; the platform is not. There is no billing, no per-key quota, no organization-level rate limiting, no files, batch or moderation surface, and no managed uptime. You now own GPU capacity, model load time on startup, version upgrades and failure handling. Client code that relied on a provider returning `429` when you exceeded a quota gets a different experience here: an overloaded vLLM queues requests, so the symptom is a growing time-to-first-token rather than a rejection, and client timeouts must be set with that in mind. ## One process, one model A `vllm serve` process holds one model in GPU memory. Serving a second model means a second process (and, usually, a second GPU) behind a router. The one exception is LoRA adapters, which share a base model in the same process and are selected through the `model` field. Embeddings are a related trap: `/v1/embeddings` is only usable when the served model runs as an embedding (pooling) model, so a chat server cannot also answer your retrieval pipeline's embedding calls — run a second server for the embedding model. ## Verifying in thirty seconds `curl http://localhost:8000/v1/models` confirms the server is up and tells you the exact name to send. Then post a one-line chat completion and check that `choices[0].message.content` comes back. Use `/health` for your orchestrator's readiness probe rather than a completion request, so probes do not consume GPU time.

  • Can that same server also answer /v1/embeddings for your retrieval pipeline?
    No. A `vllm serve` process holds one model, and `/v1/embeddings` is only usable when that model runs as an embedding (pooling) model. A chat model's server will reject embedding requests. Run a second `vllm serve` for the embedding model and point the retrieval client at it, or use a separate embedding service entirely.
  • Your application hardcodes a provider model string in a hundred places. How do you avoid touching them?
    Launch with `--served-model-name` set to that exact string. The server then advertises it at `/v1/models` and accepts it in the `model` field, while loading whatever checkpoint you actually pass to `vllm serve`. The only change left in the application is the base URL, which is normally a single environment variable.
  • The hosted client had retry-on-429 logic. What changes under vLLM?
    vLLM has no provider-side quota, so it rarely answers with a rate-limit status; overload shows up as queueing and a long time-to-first-token instead. Naive retries on timeout therefore add load to an already-saturated server. Set client timeouts that account for queue wait, cap concurrency at the client, and put quota enforcement in a gateway if you need it.

saying these in an interview costs you the question

  • Thinks vLLM forwards requests to OpenAI's servers
  • Omits the /v1 suffix from base_url
  • Assumes any model string works regardless of what is loaded
  • Expects one server process to host several different models
  • Believes an OpenAI platform key is validated by vLLM

context

open as a page

In vLLM, what breaks at /v1/chat/completions when a model ships no chat template?

level: middleimportance: must knowfreq 62%

basics

~20 s

Chat requests fail outright. vLLM renders the messages array into a single prompt using the Jinja chat template stored in the model's tokenizer_config.json; a base checkpoint has none, so you must supply one with --chat-template or use /v1/completions instead.

open as a page

In vLLM, which flags turn on OpenAI-style tool calling for a chat model?

level: middleimportance: must knowfreq 55%

basics

~20 s

Two flags together: --enable-auto-tool-choice and --tool-call-parser with the parser matching the served model's family. The parser converts the model's own tool-call text into the tool_calls field of the response; without it, calls arrive as ordinary content.

open as a page

What does vLLM's --api-key flag actually protect on the server?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only one thing: it requires callers to present a shared static token as an Authorization Bearer header on the API routes. There are no identities, scopes, quotas, rotation or rate limits, and operational routes outside the API prefix stay open.

open as a page

How do you serve several LoRA adapters from one vLLM server and select one per request?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Start with --enable-lora and register adapters via --lora-modules name=path. Each name appears in /v1/models, and a request selects one by putting that name in its model field. Adapters must all target the base checkpoint the server loaded.

open as a page

Migrating from a hosted OpenAI endpoint to vLLM, what still breaks despite API compatibility?

level: principalimportance: should knowfreq 38%

basics

~20 s

The protocol matches; the platform does not. Token counts shift with a different tokenizer, length limits become hard rejections, provider rate limits and retry semantics disappear, vendor-only parameters and surfaces have no equivalent, and the model itself is a different model.

open as a page