What does vLLM's --api-key flag actually protect on the server?
answer
- One shared secret, nothing more
- Authenticates a token, not a caller
- No quotas, no rotation, no transport security
- Operational routes answer without it
- Real controls belong at the gateway
basics
~20 sOnly one thing: it requires callers to present a shared static token as an Authorization Bearer header on the API routes. There are no identities, scopes, quotas, rotation or rate limits, and operational routes outside the API prefix stay open.
solid answer
~50 s`--api-key` (or the `VLLM_API_KEY` environment variable) installs a middleware that compares the `Authorization: Bearer <token>` header against a fixed secret and rejects mismatches with a 401. That is the entire feature. It gives you no notion of *who* is calling, so there are no per-tenant quotas, no scopes, no revocation of one consumer, no usage attribution, and no rate limiting — a single leaked token is total access. The check also applies to the API routes; operational endpoints such as `/health` and `/metrics` sit outside that prefix and answer unauthenticated, which leaks model identity and traffic shape to anyone who can reach the port. And any authorized caller can request the full context length, monopolizing the batch. Treat the flag as a tripwire against accidental access on a trusted network: bind the server to a private interface, and put a gateway in front for TLS, real authentication, per-tenant keys, quotas and request limits.
code
bash · 2 linesexport VLLM_API_KEY="$(cat /run/secrets/vllm-key)"
vllm serve meta-llama/Llama-3.1-8B-Instruct --host 127.0.0.1 --port 8000go deeper
Know that --api-key makes the server require an Authorization Bearer header matching one configured secret, and that without the flag any placeholder key is accepted.
Explain that it is a single shared token with no identity behind it, so there is no attribution, no per-caller revocation, and no rate limiting — and that it says nothing about transport security.
Demonstrate the deployment stance: private binding, a gateway owning TLS, real authentication, quotas and body-size limits, the key kept on as defence in depth, and awareness that the metrics route discloses model and traffic shape to anyone who can reach the port.
Own the cost-abuse framing: every authorized request buys GPU seconds with a three-order-of-magnitude spread, so tenancy, quota and chargeback design matter more than the token, and rotation must be possible without reloading tens of gigabytes of weights.
## What the flag does, precisely Starting the server with `--api-key <token>` (equivalently, setting `VLLM_API_KEY`) adds an HTTP middleware. For requests to the API routes, it reads the `Authorization` header, expects the literal form `Bearer <token>`, compares it against the configured value, and returns 401 if it does not match. That is the whole mechanism, and its simplicity is the point: it exists so that an OpenAI-compatible client, which already sends its key in exactly that header, works unchanged against a server that wants one. Because the SDKs always send *something* in that header, an unauthenticated vLLM ignores whatever placeholder arrives. Turning the flag on turns that placeholder into a real requirement without changing a line of client code. ## Everything it does not do **No identity.** The token authenticates a secret, not a caller. Two teams sharing one server share one token, so you cannot attribute usage, cannot revoke one consumer without breaking the other, and cannot grant one team a smaller allowance. Any per-tenant story you want has to live somewhere else. **No quotas or rate limits.** vLLM will accept whatever arrives and queue it. A misbehaving client with a valid token can saturate the batch, and the visible symptom for everyone else is climbing time-to-first-token rather than a rejection. **No request-size policy.** An authorized caller can request generations up to the server's configured context length. A handful of maximum-length requests occupy KV cache that dozens of short requests could have used, so "authorized" is nowhere near "harmless". **No rotation.** Changing the token is a process restart, which on a large model means minutes of downtime while weights reload. Rotation therefore needs a fronting layer that can swap credentials without touching the GPU process. **No TLS.** The flag is about a header, not about transport. Sent over plain HTTP, the token is on the wire in clear text on every request. ## The routes that stay open The key check covers the API paths. Operational routes — the health endpoint used by orchestrator probes and the Prometheus metrics endpoint — sit outside that prefix and answer without a token. This is convenient (your probe and your scraper do not need the secret) and it is a leak: metrics disclose which model is loaded, how much traffic it takes, queue depth and latency distributions. Anyone who can reach the port can profile your deployment even without the key. The correct conclusion is not "vLLM should protect them" but "the port should not be reachable by anyone you would not tell". ## What production actually looks like Bind the server to a private interface or a cluster-internal service, never a public address, and enforce that with network policy or security groups rather than trusting a configuration flag. Put a gateway, ingress or reverse proxy in front, and give it the jobs vLLM does not do: terminate TLS, authenticate real callers (per-tenant keys, mTLS, OIDC — whatever your organization already uses), enforce per-tenant rate limits and concurrency caps, cap request body size and requested output length, and emit usage records for chargeback. Keep `--api-key` on as well, as defence in depth: it means a pod that is accidentally exposed, or a misconfigured route that bypasses the gateway, still refuses anonymous traffic. One more habit worth adopting: pass the token through the environment rather than the command line. Process arguments are readable from the process table and land in orchestrator manifests, shell history and crash dumps, so `VLLM_API_KEY` sourced from a secret store is strictly better than an inline `--api-key` in a Kubernetes command array. ## How to answer this in an interview The question is really testing whether you treat an inference server as an application server. A model endpoint is a compute-expensive, unusually easy-to-abuse resource: every request costs GPU seconds, and the cost per request varies by three orders of magnitude depending on prompt and output length. A single shared bearer token is the least interesting part of protecting that. Say plainly what the flag is (a tripwire), name the missing controls, and put them at the layer that already owns them in your stack.
- Two internal teams share one vLLM deployment and you need per-team quotas. Where does that live?Not in vLLM. Put a gateway in front that issues a distinct credential per team, enforces per-credential rate and concurrency limits, and records usage for chargeback; it forwards to vLLM with the single shared token. vLLM keeps --api-key on as defence in depth, so traffic that somehow bypasses the gateway is still refused.
- Why prefer VLLM_API_KEY over passing --api-key on the command line?Command-line arguments are visible in the process table and get copied into orchestrator manifests, shell history and crash dumps. An environment variable sourced from a secret store keeps the value out of those surfaces and lets the platform's normal secret handling — mounting, auditing, rotation on redeploy — apply to it.
- With --api-key set, is it safe to expose the port to the internet?No. There is no TLS, so the token travels in clear text unless something else terminates HTTPS; there are no rate limits, so a valid token means unbounded GPU spend; and the health and metrics routes answer unauthenticated, disclosing the model and traffic shape. Bind to a private interface and front it with a gateway.
saying these in an interview costs you the question
- Treats --api-key as production-grade authentication
- Assumes it enforces per-user rate limits
- Thinks it protects the metrics and health endpoints
- Believes the key is encrypted in transit by itself
- Says rotating the key needs no restart