skip to content

Deploying vLLM with --tensor-parallel-size 4 on Kubernetes: what must the pod provide?

level: seniorimportance: should knowfreq 46%

answer

  1. shards are processes, not pods
  2. one container holds every GPU
  3. 64 MiB /dev/shm is a trap
  4. loading is measured in minutes
  5. in-flight streams need a grace period

basics

~20 s

All four GPUs must be visible to one container in one pod, since vLLM's tensor-parallel workers are local processes sharing memory — not pods that find each other. The pod also needs a large /dev/shm, somewhere durable for weights, and probes that tolerate minutes of loading.

solid answer

~50 s

Tensor parallelism in vLLM is intra-process: the engine spawns one worker per shard on the **same host**, and those workers exchange activations through shared memory and NCCL over the local interconnect. So the pod requests `nvidia.com/gpu: 4` on a single container — you cannot spread a tensor-parallel group across four one-GPU pods, and the node must actually have four GPUs with a fast link between them. Beyond the device request: mount an `emptyDir` with `medium: Memory` at `/dev/shm` and size it in GiB, because the 64 MiB default breaks the workers' shared-memory transport. Give weights a durable home — a PVC, a pre-warmed host cache, or a baked image — or every restart re-downloads them. And use a `startupProbe` against `/health` with a generous failure budget, since loading weights and capturing CUDA graphs takes minutes; a liveness probe tuned for a web app will crash-loop a healthy server.

code

yaml · 32 lines
yaml
containers:
  - name: vllm
    image: vllm/vllm-openai:v0.27.1
    args:
      - --model
      - meta-llama/Llama-3.1-70B-Instruct
      - --tensor-parallel-size
      - "4"
      - --max-model-len
      - "8192"
    ports:
      - containerPort: 8000
    resources:
      limits:
        nvidia.com/gpu: 4
    volumeMounts:
      - name: dshm
        mountPath: /dev/shm
      - name: hf-cache
        mountPath: /root/.cache/huggingface
    startupProbe:
      httpGet: { path: /health, port: 8000 }
      periodSeconds: 10
      failureThreshold: 90
volumes:
  - name: dshm
    emptyDir:
      medium: Memory
      sizeLimit: 16Gi
  - name: hf-cache
    persistentVolumeClaim:
      claimName: hf-models

go deeper

for a junior

Know that the container must request as many GPUs as the tensor-parallel size, and that the model has to load before the server answers anything, so start-up is slow.

for a middle

Explain why the shards must share a container: they are worker processes exchanging activations locally through shared memory and the GPU interconnect. Name the /dev/shm emptyDir and the weights volume as required, not optional.

for a senior

Show the operational reflexes: startupProbe with a long budget gating liveness, a grace period longer than a generation, pre-staged weights, and pinned image versions. Be ready to diagnose a crash-loop as a probe misconfiguration rather than a model problem.

for a principal

Own the fleet shape — which node SKUs you buy, whether tensor parallelism or more replicas is the standard answer at your model sizes, how weights are distributed cluster-wide, and what rollout policy keeps capacity flat while each replacement takes minutes.

## The shape of the workload A vLLM replica is one process group pinned to one node's GPUs. That single sentence determines almost everything about the pod spec, and it is what interviewers are checking: candidates who model an LLM server as a horizontally sliceable stateless service get this wrong immediately. ## GPU count and placement `--tensor-parallel-size N` shards each layer's weights across N devices and requires all N to be addressable from one process. In Kubernetes that means one container requesting `nvidia.com/gpu: N` in its resource limits. GPU resources are integral and non-overcommittable, so requests equal limits by construction, and the scheduler will only place the pod on a node that can satisfy the whole request at once. If you also use `--pipeline-parallel-size`, total devices needed is tensor size times pipeline size. The corollary matters more than the syntax: **four pods with one GPU each is not a tensor-parallel deployment**, it is four independent replicas each trying to load the whole model. If the model does not fit on one GPU, they all fail. And because tensor parallelism exchanges activations on every layer, the node's interconnect quality is part of the deployment decision — the same flag on a machine whose GPUs talk over PCIe behaves very differently from one with a high-bandwidth link. ## Shared memory vLLM's workers are separate processes, and PyTorch moves data between them through `/dev/shm`. A container's default 64 MiB is far below what those transports need, and the symptom is ugly: a hang during initialisation, or a bus error, with no clear message pointing at shared memory. The fix in a pod spec is an `emptyDir` volume with `medium: Memory` mounted at `/dev/shm`, with a `sizeLimit` in the gigabytes. Note that memory-backed `emptyDir` counts against the pod's memory limit, so the container's memory request must cover it plus the process's own host-RAM use — which is itself substantial, because weights are read through host memory on the way to the device. ## Weights Model weights are the dominant fact of the deployment's lifecycle. Tens to hundreds of gigabytes must reach the node before the server is useful. The options, in increasing order of operational effort and decreasing order of cold-start time: download into ephemeral storage on every start (simplest, worst), mount a shared read-write-many volume or a node-local pre-warmed cache (good middle ground), or bake the weights into the container image (largest images, but pulls are cached and deduplicated by the node's runtime). Ephemeral downloads also interact badly with crash-loops, turning one bad config into sustained bandwidth cost. ## Probes vLLM exposes `/health` on the API port, plus `/v1/models` and `/metrics`. The trap is timing. Between container start and a healthy server sit: image pull, weight download or read, host-to-device transfer, worker startup for each shard, profiling and CUDA graph capture. Minutes, routinely. A liveness probe with a short `initialDelaySeconds` restarts the container mid-load, and the restart repeats the same slow path forever. The correct pattern is a `startupProbe` on `/health` with a long failure budget that gates the liveness and readiness probes, so slow-start is tolerated exactly once while a genuinely wedged server is still restarted later. Readiness deserves its own thought: a vLLM server accepts requests as soon as it is up and queues what it cannot run immediately, so readiness reflects "loaded", not "has spare capacity". Anything that needs the second thing has to look at the queue and cache metrics rather than at the probe. ## Shutdown In-flight generations can run for tens of seconds. A short `terminationGracePeriodSeconds` cuts streaming responses off mid-sentence during every rollout, which users perceive as an outage. Give the pod a grace period longer than your longest expected generation, and roll one replica at a time given how expensive each replacement is. ## What stays out of this How the scheduler chooses the node, how taints keep non-GPU work off expensive hardware, and how probes work in general are Kubernetes concerns with their own answers. What is specific to vLLM is the shape it imposes: all shards in one container, big shared memory, big slow-loading state, and long-lived streaming connections.

  • Why can't you run four one-GPU pods and have them form a tensor-parallel group?
    Because vLLM's tensor-parallel workers are processes launched by one engine on one host, coordinating through shared memory and the local GPU interconnect on every layer. There is no discovery mechanism for pods to join a shard group, and even if there were, per-layer collectives over pod networking would be catastrophically slow. Multi-node serving exists, but it is pipeline parallelism across nodes on top of tensor parallelism within each node.
  • Your rollout cuts off users mid-response. What in the pod spec is wrong?
    terminationGracePeriodSeconds is shorter than a generation. On termination the pod gets SIGTERM, and when the grace period expires it is killed regardless of in-flight streaming responses. Set the grace period above your longest expected completion time, and roll replicas one at a time with maxUnavailable at zero so capacity does not dip while a several-minute replacement loads.
  • Should readiness reflect that the replica has spare capacity?
    No — keep readiness as a liveness-of-the-model check. vLLM queues what it cannot run immediately, so a saturated replica is still correctly serving; flipping it out of the endpoint list would shift its queue onto peers and cascade. Capacity belongs in autoscaling signals drawn from the queue and cache metrics, not in a probe that controls traffic membership.

saying these in an interview costs you the question

  • Splits a tensor-parallel group across several pods
  • Leaves /dev/shm at the container default
  • Uses a short liveness probe and crash-loops on load
  • Downloads weights into ephemeral storage every start
  • Treats a GPU replica as horizontally interchangeable in seconds

context