skip to content

Why do self-hosted Llama servers expose an OpenAI-compatible endpoint?

level: middleimportance: must knowfreq 58%

answer

  1. one schema everything already speaks
  2. swap the base URL, keep the client
  3. schema portability, not capability parity
  4. model string must match what is served
  5. tool-call parsing needs server flags

basics

~20 s

Because the OpenAI chat-completions shape became the de facto wire protocol: exposing /v1/chat/completions lets existing SDKs, gateways and frameworks talk to a local model by changing only the base URL. Compatibility covers the common fields, not every parameter or feature.

solid answer

~60 s

vLLM, TGI, Ollama and llama.cpp's `llama-server` all serve an OpenAI-shaped `/v1/chat/completions` (and usually `/v1/completions` and `/v1/models`). The reason is ecosystem gravity: client SDKs, LangChain-style frameworks, proxies and observability tools already speak that schema, so adopting it means self-hosted Llama drops into existing code with a `base_url` change and a dummy API key. What you must not assume is parity. Compatibility is at the level of `messages`, `model`, `temperature`, `max_tokens` and SSE streaming deltas. Beyond that it frays: `model` must match the name the server actually serves (the Hugging Face repo path, `--served-model-name`, or the Ollama tag), unsupported parameters may be ignored silently rather than rejected, `/v1/embeddings` only exists if you serve an embedding model, and tool calling in vLLM is off until you start it with `--enable-auto-tool-choice` and a matching `--tool-call-parser`. Ollama's OpenAI route also has no field for context length — that is set natively or in the Modelfile. Treat the schema as portable and the behaviour as server-specific: re-test streaming, tool calls and parameter handling after every swap.

code

python · 16 lines
python
from openai import OpenAI

# Same client class, three different servers
local_vllm = OpenAI(base_url="http://localhost:8000/v1", api_key="not-used")
local_ollama = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

# The model string must match what the server actually serves
print([m.id for m in local_vllm.models.list().data])

resp = local_vllm.chat.completions.create(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[{"role": "user", "content": "Say hi."}],
    stream=True,
)
for chunk in resp:
    print(chunk.choices[0].delta.content or "", end="")

go deeper

for a junior

Know that self-hosted servers speak the OpenAI chat-completions shape, so you point the same client at a different base URL and use the model name that server reports.

for a middle

Explain the boundary: the schema and SSE streaming are portable, but the model identifier, honoured parameters, embeddings availability and tool-call parsing are server-specific.

for a senior

Show you have been burned — silently ignored parameters, tool calls arriving as plain text, mismatched token counts — and that you keep a compatibility test suite that runs on every server version bump.

for a principal

Own the abstraction as a strategy: standardising on one wire contract keeps providers swappable and preserves negotiating leverage, provided you also own the gateway that supplies the auth, rate limiting and cost attribution the contract does not.

## How one vendor's schema became the protocol There is no standards body for LLM HTTP APIs. What happened instead is that the OpenAI chat-completions request and response shape — a `messages` array of `{role, content}`, a `model` string, sampling parameters, and a `choices` array with `finish_reason` and a `usage` object — arrived first and was implemented by every client library, framework and gateway. Serving that shape is therefore the cheapest possible integration decision for a self-hosted stack: nobody has to write a client for you. The practical consequence is that an application can be written once against an OpenAI-style client and pointed at a hosted API in production, a vLLM cluster in staging, and Ollama on a laptop, differing only by `base_url`, the model string, and a placeholder key that local servers ignore. ## What each stack actually exposes - **vLLM**: starts an HTTP server (port 8000 by default) with `/v1/chat/completions`, `/v1/completions`, `/v1/models`, plus `/health` and a Prometheus `/metrics` endpoint. `/v1/embeddings` is served when the loaded model is an embedding model. - **TGI**: has its own native `/generate` and `/generate_stream` routes and additionally serves an OpenAI-compatible `/v1/chat/completions`. - **Ollama**: native API at `/api/generate`, `/api/chat` and `/api/embed`, with an OpenAI-compatible surface under `/v1/`. - **llama.cpp `llama-server`**: native `/completion` alongside an OpenAI-compatible `/v1/chat/completions`. So most servers have two APIs: a native one that exposes everything the engine can do, and a compatibility one that exposes the subset expressible in OpenAI's schema. ## Where compatibility ends **The model identifier.** Hosted APIs accept a published model name. A self-hosted server accepts whatever name it loaded the weights under — usually the Hugging Face repo id such as `meta-llama/Llama-3.1-8B-Instruct`, overridable with `--served-model-name`, or an Ollama tag like `llama3.1:8b`. Call `/v1/models` and use exactly what it returns; a mismatch is a 404-class error, not a fallback. **Parameters.** The common sampling fields are honoured. Beyond them, behaviour varies by server and version: some parameters are unimplemented and rejected, some are accepted and ignored, and some engine-specific knobs have no OpenAI field at all. Ollama's context length is the classic example — there is no OpenAI parameter for it, so you set it through the native `/api/chat` options or a `Modelfile` `PARAMETER num_ctx`, and a client that only speaks OpenAI simply cannot change it. Never infer from a 200 response that a parameter took effect; verify behaviourally. **Tool calling.** The request schema accepts `tools` and `tool_choice`, but the server has to parse the model's raw output back into structured `tool_calls`, and that parsing is model-format-specific. In vLLM you must launch with `--enable-auto-tool-choice` and select a `--tool-call-parser` appropriate to the model family; without those flags the tool call arrives as plain text in `content` and your client silently sees no tool call. This is the single most common "it worked against the hosted API and broke locally" bug. **Everything operational.** Hosted APIs give you organisation-level rate limits, quota headers, billing and usage dashboards. A self-hosted server gives you a queue and a `/metrics` endpoint. Rate limiting, authentication, multi-tenancy and cost attribution are yours to build in front of it — commonly with a gateway or ingress that terminates auth and enforces per-tenant limits before traffic reaches the model server. **Response details.** `usage` token counts come from the local tokenizer and will not match another provider's counts for identical text. `finish_reason` values map onto the engine's stop conditions. Streaming is SSE with incremental deltas terminated by `[DONE]`, which is the part with the highest fidelity across stacks. ## The right mental model OpenAI compatibility is a *transport and schema* agreement, not a *capability* agreement. It guarantees your HTTP client, your streaming parser and your message-shaping code keep working. It guarantees nothing about which parameters are honoured, whether tool calls are parsed, what the context limit is, or how the server behaves under load. Write an integration test that exercises streaming, a tool call and a long prompt against every server you deploy against, and run it on every version bump — servers change their compatibility surface far more often than the hosted APIs change theirs.

  • Your client sends `tools` to a vLLM server and gets prose back instead of `tool_calls`. Why?
    The server accepted the schema but was not started with tool-call parsing enabled, so the model's structured output was returned verbatim as assistant text. Launch vLLM with `--enable-auto-tool-choice` and a `--tool-call-parser` matching the model family, then confirm that `choices[0].message.tool_calls` is populated. Nothing in the HTTP status code warns you about this — the request succeeds either way.
  • How do you pick the right value for the `model` field against a self-hosted server?
    Query `/v1/models` and use the id it returns. vLLM defaults to the Hugging Face repo path it loaded unless you pass `--served-model-name`; Ollama uses its tag, such as `llama3.1:8b`. Hard-coding a hosted provider's model name is a guaranteed failure, and it is worth asserting the id at startup so a misconfigured deployment fails loudly rather than at first request.
  • Why do token counts in `usage` differ from another provider's for the same text?
    Counts come from the tokenizer of the model actually loaded. Llama's tokenizer segments text differently from other families, so the same string yields different token totals. Use the local counts for capacity planning against your own context limit, and never port a token budget computed against one provider to another without recounting.
  • What does OpenAI compatibility not give you that a hosted API does?
    Everything operational: authentication, per-tenant rate limits and quota headers, billing, usage dashboards, and global availability. A self-hosted server exposes a queue and a metrics endpoint. You put a gateway in front for auth, rate limiting and cost attribution, and you own capacity, upgrades and incident response.

saying these in an interview costs you the question

  • Assuming every OpenAI parameter is honoured because the request returns 200
  • Passing a hosted provider's model name to a local server
  • Expecting tool calls to be parsed without enabling the server-side parser
  • Believing token counts are comparable across providers
  • Thinking compatibility includes rate limits, quotas or billing

context