skip to content

What does a minimal vllm/vllm-openai docker run need to serve a Hub model?

level: juniorimportance: should knowfreq 56%

answer

  1. it is not a stateless web app
  2. the entrypoint is already the server
  3. 64 MiB of shared memory is not enough
  4. weights need somewhere to live
  5. gated repos need a token in env

basics

~20 s

GPU access, a published port, host shared memory, and a mounted Hugging Face cache so weights survive restarts. Gated repos also need a Hub token in the environment. Engine arguments go after the image name, because the image already starts the OpenAI-compatible server.

solid answer

~40 s

The documented invocation is `docker run --runtime nvidia --gpus all -v ~/.cache/huggingface:/root/.cache/huggingface --env "HUGGING_FACE_HUB_TOKEN=$HF_TOKEN" -p 8000:8000 --ipc=host vllm/vllm-openai:latest --model <repo-id>`. Four things matter. GPU access via the NVIDIA container runtime, or vLLM sees no device and fails. `--ipc=host` (or an explicit `--shm-size`) because PyTorch uses shared memory between processes and Docker's 64 MiB default is far too small once tensor parallelism spawns workers. A volume for the Hub cache, or every restart re-downloads tens of gigabytes into the container's writable layer. And a token in the environment for gated repos. Arguments after the image name are passed straight to the API server, so `--max-model-len`, `--tensor-parallel-size` and friends go there.

code

bash · 11 lines
bash
docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --env "HUGGING_FACE_HUB_TOKEN=$HF_TOKEN" \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model mistralai/Mistral-7B-Instruct-v0.2 \
  --max-model-len 8192

# once it answers, confirm what it is serving
curl -s localhost:8000/v1/models

go deeper

for a junior

Be able to recite the four essentials — GPU flags, published port, shared memory, cache volume — and know that engine arguments follow the image name because the image already runs the API server.

for a middle

Explain why each flag exists rather than listing them: the container toolkit for devices, /dev/shm for PyTorch's inter-process tensors, the volume for cold starts, the token for gated repos. Mention /health and /v1/models for verification.

for a senior

Talk about running this as a real workload: pinned version tags, secrets rather than env literals in scripts, weights pre-staged on the node or baked in, and health checks whose timeouts tolerate minutes of model load.

for a principal

Decide the platform contract — one blessed base image and version cadence, where weights live cluster-wide, and how teams get GPUs — so every service does not reinvent this command line with a different set of mistakes.

## The image's contract `vllm/vllm-openai` is the official image whose entrypoint is the OpenAI-compatible API server. That single fact shapes the command line: you do not pass a shell command, you pass **engine arguments**, and they land on the server directly. The image listens on port 8000 by default, so `-p 8000:8000` makes it reachable. ## The four things a minimal run must supply **GPU access.** `--runtime nvidia --gpus all` (or `--gpus '"device=0,1"'` to select) exposes devices through the NVIDIA container toolkit. Without it the container starts and then fails during engine init because no CUDA device is visible. This is the single most common first failure and it is worth being able to say what the toolkit does: it injects the driver libraries and device nodes into the container. **Shared memory.** vLLM runs on PyTorch, which passes tensors between processes through shared memory — most visibly when tensor parallelism spawns one worker per GPU. Docker gives a container 64 MiB of `/dev/shm` by default, which is not enough; the run hangs or dies with a bus error. `--ipc=host` shares the host's IPC namespace and is what vLLM's own documentation shows; `--shm-size=16g` is the more contained alternative when you would rather not share the namespace. **A weights cache.** Mounting `~/.cache/huggingface` to `/root/.cache/huggingface` means the model is downloaded once per host rather than once per container start. Skipping it is not a correctness bug, it is a cold-start and disk bug: a 70B checkpoint is well over a hundred gigabytes, it lands in the container's writable layer, it is thrown away on every restart, and it turns every crash-loop into a bandwidth incident. In production this is a persistent volume, a pre-warmed node-local cache, or weights baked into an image. **Credentials for gated repos.** `--env "HUGGING_FACE_HUB_TOKEN=..."` lets the download authenticate. Without it, gated families fail at download with an authorization error that reads confusingly like a missing-model error. The token is a secret: it belongs in a secret store or an orchestrator secret, never in the image or in shell history. ## Engine arguments Everything after the image name is server configuration, and it is where the deployment's real decisions live: `--model` or the positional model id, `--tensor-parallel-size`, `--max-model-len`, `--gpu-memory-utilization`, `--quantization`, `--kv-cache-dtype`, `--served-model-name` to control the name clients pass in the `model` field. Because they are ordinary process arguments, they can come from an orchestrator's `args` list and be templated per environment. ## Verifying it came up The server exposes `/health` for liveness-style checks and `/v1/models` to confirm which model name it is serving; `/metrics` returns Prometheus text. A server that answers `/health` has finished loading — and loading is slow, because it means reading tens of gigabytes of weights from disk into VRAM and capturing CUDA graphs. Anything watching the container needs to tolerate minutes of startup, which is why a naive health check with a short timeout will kill a perfectly good server in a loop. ## Things people get wrong Treating the image like a stateless web app is the theme of every mistake here. It is not stateless: it holds a large model in VRAM, it wants a warm on-disk cache, and it takes minutes to become useful. It does not autoscale in seconds, it does not survive an aggressive liveness probe during load, and running two of them on one card without lowering each one's memory budget will simply OOM. The image is a convenient package for a heavyweight process, not a lightweight service. ## Variants The published tags are versioned as well as `latest`; pinning an exact version is the right default for a deployment, because engine argument names and defaults do move between releases. Building your own image on top is common when you want to bake in weights, add a startup wrapper, or pin CUDA and driver expectations to your fleet.

  • Why do vLLM's own docs suggest --ipc=host rather than leaving the default?
    Because PyTorch shares tensors between processes through /dev/shm, and Docker sizes that at 64 MiB by default. Once vLLM spawns worker processes — which it does for tensor parallelism — that is exhausted immediately and the run hangs or aborts. --ipc=host shares the host's IPC namespace so the limit is the host's; --shm-size=16g is the narrower alternative when sharing the namespace is not acceptable.
  • Is pinning to :latest acceptable for a production deployment?
    No. Engine argument names, defaults and even accepted values change between vLLM releases, so :latest means a node replacement can silently change your server's behaviour or reject a flag that worked yesterday. Pin an exact version tag, upgrade deliberately, and re-run your latency and quality checks on the new version before rolling it out.
  • How would you avoid downloading weights at container start altogether?
    Two options. Bake the weights into an image layer, which makes the image enormous but makes start-up a pure image pull that node-level caching already optimises. Or keep weights on a shared read-only volume mounted into the container, so the first pod on a node pays nothing and the download happens once per cluster. Both trade image or storage management for a much shorter and more predictable cold start.

saying these in an interview costs you the question

  • Forgets GPU access flags and blames the model
  • Leaves the default 64 MiB /dev/shm
  • Re-downloads weights on every container start
  • Puts a Hub token into the image
  • Expects it to start in seconds like a stateless service

context