skip to content

Why does a new GPU replica take minutes to serve LLM traffic?

level: seniorimportance: must knowfreq 55%

answer

  1. weights are tens of gigabytes
  2. node, image, weights, load, warmup
  3. the correction arrives minutes later
  4. ready before warmup is a lie
  5. cache weights on the node

basics

~20 s

A new replica must wait for a GPU node, pull a multi-gigabyte container image, fetch tens of gigabytes of weights, load and shard them into VRAM, and run a warmup pass before it is useful. Each stage is minutes, and none of them is code you wrote.

solid answer

~60 s

Cold start on a GPU fleet is a chain of slow stages. The cluster autoscaler must obtain a GPU node, which can take minutes and can fail outright when the instance type is scarce. The serving image carries CUDA, PyTorch and compiled kernels and runs to many gigabytes. The weights are the real bill: bf16 costs roughly 2 bytes per parameter, so a 70B checkpoint is about 140 GB pulled from object storage or a model hub, then loaded and sharded across tensor-parallel ranks. Finally the engine warms up — vLLM compiles and captures CUDA graphs unless you pass `--enforce-eager`, and TGI runs a warmup pass to size its batches. You shorten this by attacking each stage: pre-pull or bake the image onto nodes, cache weights on node-local NVMe or a shared read-many volume, keep a warm node pool so provisioning is out of the path, and make sure the readiness probe (vLLM's `/health`, Triton's `/v2/health/ready`) only passes after warmup so the load balancer does not send traffic into a loading process.

code

yaml · 10 lines
yaml
containers:
  - name: vllm
    image: vllm/vllm-openai:latest
    startupProbe:
      httpGet: { path: /health, port: 8000 }
      periodSeconds: 10
      failureThreshold: 60   # tolerate ~10 min of load + warmup
    readinessProbe:
      httpGet: { path: /health, port: 8000 }
      periodSeconds: 5

go deeper

for a junior

Be able to say that model weights are tens of gigabytes and must be downloaded and loaded into GPU memory before the server answers anything, so a new replica is not instant.

for a middle

Walk the stages in order — node, image, weights, load, warmup — and give rough magnitudes, including the bytes-per-parameter arithmetic that makes the weight fetch dominate.

for a senior

Show the fixes you have actually applied: pre-pulled images, node-local or shared weight cache, warm node pool, and a readiness probe pointed at the engine health route with a startup probe sized for real load time.

for a principal

Own the tradeoff between permanent warm capacity and cold-start risk across a fleet, including whether GPU capacity is reliably obtainable in your region and what that means for scale-in policy and for any scale-to-zero promise.

## The chain, stage by stage Scaling a stateless web pod is fast because the only slow step is pulling a small image. An LLM replica has five slow steps, and they are serial. **1. Getting a GPU node.** If the cluster has no spare GPU node, the node autoscaler must ask the cloud provider for one. That is typically one to five minutes, and unlike CPU capacity it can simply fail: GPU instance types are frequently unavailable in a given zone, so the pod stays Pending. Any plan that assumes a node is always obtainable on demand is a plan with an unbounded worst case. **2. Pulling the image.** Serving images ship CUDA runtime, PyTorch, and precompiled attention and quantization kernels. They are large — several gigabytes and up — and a cold node pulls them over the network into a fresh layer store. **3. Fetching the weights.** This dominates. Bytes per parameter times parameter count is the arithmetic: bf16 is roughly 2 bytes/param, so an 8B model is about 16 GB and a 70B model about 140 GB. Pulling that from object storage or a model hub at, say, 1 GB/s is 20 seconds for the small one and over two minutes for the large one — and that is with excellent bandwidth and no throttling. Quantized checkpoints shrink this proportionally, which is a cold-start argument for quantization independent of the memory argument. **4. Loading and sharding.** Weights are read into host memory and copied to the device, and under tensor parallelism each rank takes its slice while the ranks establish their collective communicator. vLLM's `--load-format` selects how weights are read, and the choice matters when the storage path is the bottleneck. **5. Warmup.** vLLM's V1 engine compiles the model and captures CUDA graphs at startup so steady-state decode avoids per-step launch overhead; `--enforce-eager` skips that at the cost of slower steady-state execution. TGI runs a warmup pass to determine safe batch sizes. TensorRT-LLM avoids compilation at startup because the engine was built ahead of time, but it still has to load that engine. Warmup is tens of seconds, and it is real work — skipping it moves the cost into the first user's request. ## Why this breaks web-tier scaling instincts An HPA-style control loop assumes the correction arrives soon after the signal. Here the correction arrives three to ten minutes later. Two consequences follow. First, you must trigger scale-out on a leading indicator while the current replicas still have real slack — a threshold tuned to "we are full now" guarantees a multi-minute outage window. Second, scale-in must be far more conservative than scale-out, because a replica you drop is expensive to get back; asymmetric stabilization windows are the norm. ## Shortening each stage - **Node:** keep a warm pool. Either hold a floor of idle GPU nodes, or run low-priority placeholder workloads that a real replica can evict, so provisioning happens before demand rather than during it. - **Image:** pre-pull it onto every GPU node (a pre-puller DaemonSet is the usual mechanism) or bake it into the node image. Once cached, this stage rounds to zero. - **Weights:** cache them on node-local NVMe, or mount a shared read-many volume that all replicas on that model share. Fetching from the public hub on every scale-out is both the slowest option and the one most likely to rate-limit you. Fetching from same-region object storage is far better than cross-region. - **Load:** choose a fast load path and prefer formats that memory-map rather than deserialize; where the network is the constraint, streaming loaders that overlap fetch and load help. - **Warmup:** accept it, and make sure it happens before traffic arrives. ## The readiness-probe trap The most common self-inflicted wound is a probe that passes too early. If readiness checks a port that binds before weights are loaded, the service endpoint goes ready while the process is still loading, the load balancer routes real traffic to it, and every request there times out — which looks like a scaling failure and is actually a probe bug. Point readiness at the engine's own health route (vLLM `/health`, TGI `/health`, Triton `/v2/health/ready`) and give the container a startup probe with a failure threshold generous enough to cover the worst observed load time, so a slow start is not mistaken for a crash loop and restarted into a loop. ## What to measure Instrument the stages separately: pod scheduled, image pulled, weights loaded, first successful health check. A single "time to ready" number tells you it is slow; the breakdown tells you which stage to fix, and they have completely different fixes. And publish the number — every autoscaling threshold, every scale-to-zero decision and every incident review downstream depends on knowing that a replica takes, say, four minutes rather than forty seconds.

  • How does cold-start time change what threshold you set on the autoscaler?
    It forces headroom. If a replica needs four minutes, the trigger has to fire while the fleet still has roughly four minutes of absorbable growth left, which means a lower target and a shorter scale-out window than a web tier would use. Scale-in gets the opposite treatment — a long stabilization window — because reacquiring a replica is expensive and may fail if GPU capacity is scarce.
  • Does a quantized checkpoint help cold start, and by how much?
    Yes, roughly in proportion to the byte reduction, because the weight fetch dominates. A 70B model at 4-bit is about 35 GB instead of 140 GB, cutting the download and load stages to a quarter. That is a separate benefit from the VRAM saving, and it is often the more visible one on an autoscaling fleet.
  • Your pods go Ready in 40 seconds but the first requests all time out. What happened?
    The readiness probe is passing before the engine can serve. Something bound the port — the HTTP frontend, or a probe pointed at a TCP check — while weights were still loading, so the endpoint was added to the service and traffic arrived into a loading process. Point readiness at the engine's health route and add a startup probe sized for the real load time.
  • Where do prebuilt TensorRT-LLM engines sit on this curve?
    They remove compilation from startup because the engine was built ahead of time, so warmup is mostly load rather than compile. The tradeoff moves elsewhere: the engine is pinned to a GPU architecture and precision, so your warm pool and your image must match that SKU, and a fleet spanning two GPU generations needs two artifacts.

saying these in an interview costs you the question

  • Assumes GPU pods scale like stateless web pods
  • Ignores that the GPU node itself may not exist yet
  • Readiness probe passes before weights are loaded
  • Downloads weights from the public hub on every scale-out
  • Treats scale-in as symmetric with scale-out

context