skip to content

When is scale-to-zero right for a self-hosted LLM endpoint?

level: principalimportance: should knowfreq 38%

answer

  1. idle GPU hours versus first-request pain
  2. the node, not just the pod
  3. frequent bursts mean constant thrash
  4. capacity may not be there on return
  5. warm floor beats true zero for users

basics

~20 s

Scale to zero when idle hours dominate the bill and the first caller after an idle period can tolerate minutes — internal tools, batch pipelines, dev environments, rarely used per-tenant models. Never for interactive traffic with a time-to-first-token target, or where GPU capacity may not be reacquirable.

solid answer

~60 s

The decision is a trade of idle GPU cost against first-request latency and availability risk. Scale-to-zero pays when utilization is genuinely bursty — a few hours of use a day — and the workload is asynchronous or internally facing, because an idle H100 bills the same as a busy one. It fails for user-facing chat, because the cold start is minutes: node acquisition, a multi-gigabyte image, tens of gigabytes of weights, then warmup. Two second-order risks decide most real cases. First, releasing the node is what actually saves money; if the pod scales to zero but the GPU node stays, you have saved nothing. Second, scarce GPU SKUs may not be available when you want to come back, so scale-to-zero silently converts a cost optimization into an availability risk. Mechanically, stock HPA will not go below one replica, so scale-to-zero comes from KEDA or a serverless layer that holds the first request while a replica starts. Intermediate positions usually win: a warm floor of one small replica for interactive traffic bursting to larger ones, scheduled scaling to a known daily curve, or a warm node pool with cold pods.

go deeper

for a junior

Know that an idle GPU still costs full price, but that starting a replica from nothing takes minutes, so turning an endpoint off is only sensible when nobody is waiting.

for a middle

Compare idle cost against measured cold-start time, and name the workloads that fit — batch, internal tools, dev — versus interactive traffic with a first-token target.

for a senior

Show the mechanics: releasing the node rather than the pod, a component that holds the first request while a replica starts, hysteresis against thrash, and timeouts and retry policy that survive a minutes-long activation.

for a principal

Own the whole trade — idle spend against latency and availability risk, whether the SKU is reliably reacquirable, the intermediate positions (warm floor, scheduled scaling, consolidation, cheaper idle), and the explicit contract you publish to callers.

## Frame it as a trade, not a feature Scale-to-zero is not free efficiency; it exchanges money for latency and for availability. To decide, you need three numbers: the hourly cost of an idle replica, the measured cold-start time, and what the first request after an idle period is worth. A dev-environment endpoint used twenty minutes a day at several dollars an hour is an obvious yes. A customer-facing assistant with a two-second TTFT target is an obvious no. Most real cases sit between, and the analysis is what distinguishes a senior answer from a principal one. ## When it clearly pays - **Internal tools and dev/staging.** Nobody is paying for latency, and idle hours are the overwhelming majority. - **Batch and asynchronous pipelines.** The caller submits work and collects results later, so a four-minute spin-up is amortized over an hour of processing. This is the strongest case: run the fleet hot while the queue is deep, and hold nothing when it is empty. - **Long-tail per-tenant models.** With hundreds of fine-tuned models of which a handful are used daily, keeping each warm is untenable. Note the alternative first: if they share a base model, multi-LoRA multiplexing keeps them all warm on one replica and beats scale-to-zero outright. - **Expensive, scarce-use experiments.** A large model evaluated occasionally is precisely the workload nobody should keep resident. ## When it clearly does not - **Interactive traffic with a TTFT SLO.** A user waiting minutes for the first token has left. No amount of clever queueing hides a cold GPU. - **Traffic that is bursty but frequent.** Arrivals every few minutes cause continuous thrash: the endpoint is always either starting up or shutting down, and you pay the start cost repeatedly while barely saving idle time. Hysteresis — a long idle timeout before scale-in — is the fix, and once the timeout is long enough to avoid thrash, the savings often evaporate. - **Scarce GPU capacity.** If your SKU is intermittently unavailable in your region, releasing a node is a bet that you can get it back. Losing that bet turns a cost optimization into an outage, and the failure is silent until the day it matters. ## The mechanics you must not get wrong Three things trip teams up. **The pod is not the cost.** Deleting a pod frees nothing if its GPU node stays in the pool. Real savings require the node to be released and the instance stopped, which adds node provisioning to the return path — the slowest stage of cold start. Conversely, keeping a warm node pool with no pods running is a middle position: pods start fast because the image and possibly the weight cache are local, but you keep paying for the hardware. **Something must hold the first request.** Stock HPA does not scale below one replica (a scale-to-zero feature gate exists but has long been alpha), so zero comes from KEDA or a serverless/gateway layer that can accept a request with no backend, trigger activation and hold the connection while the replica starts. Whatever holds it must survive minutes, which usually means generous proxy timeouts and a client that will not retry midway — an impatient retry policy against a cold endpoint multiplies the very work you are waiting on. **Cold start must be measured, not assumed.** Everything above depends on that number, and it is the one teams quote from memory and get wrong by a factor of three. ## The intermediate positions that usually win - **A warm floor with burst.** Keep one replica — possibly a smaller model or a smaller GPU — always on to absorb interactive traffic instantly, and scale the expensive tier up on demand. Users never see a cold start; you pay for one replica, not the peak. - **Scheduled scaling.** Most internal traffic follows a known daily and weekly curve. Scaling to zero overnight and back before the workday costs nothing in user-visible latency and captures most of the savings, without betting on a reactive control loop. - **Consolidation.** Before choosing zero, ask whether several sparsely used endpoints could be one endpoint: same base model with adapters, or a shared multi-model server. A shared replica at 40% utilization beats four idle replicas and four cold starts. - **Cheaper idle.** A smaller GPU SKU, a quantized checkpoint or a spot instance can cut idle cost enough that keeping the endpoint warm becomes affordable — often a better answer than eliminating idle entirely. ## The organizational half Whoever chooses scale-to-zero must own the consequence. State it as an explicit contract: this endpoint's first request after idle takes up to N minutes, it is not covered by the interactive latency SLO, and here is the failure mode if GPU capacity is unavailable. That sentence is worth more than the configuration, because it stops the endpoint from being quietly adopted by a latency-sensitive caller six months later — which is how most scale-to-zero incidents actually begin.

  • Scaling to zero pods saved nothing on the bill. Why?
    Because the GPU node stayed in the pool. Cloud billing is per instance-hour, not per pod, so the saving only lands when the node is released and the instance stopped. That also lengthens the return path, since node provisioning rejoins the cold-start chain — which is exactly the trade a warm-node/cold-pod setup makes in the other direction.
  • Traffic arrives every few minutes. What does scale-to-zero do?
    It thrashes — the endpoint spends its life starting up and shutting down, paying repeated cold starts while saving little idle time, and every gap-crossing request eats the full spin-up. Lengthen the idle timeout until the pattern stabilizes, and then check whether the remaining savings still justify it. Usually they do not, and a warm floor is the better answer.
  • How do you present a scale-to-zero endpoint to its callers?
    As an explicit contract: first request after idle may take up to N minutes, it is excluded from the interactive latency SLO, and capacity acquisition can fail. Ideally make the shape asynchronous — accept the job, return a handle — so the wait is structural rather than a mysteriously slow synchronous call that a caller will eventually wrap in a retry loop.
  • When is consolidation a better answer than scaling to zero?
    When the sparse endpoints share a base model. Several fine-tunes as LoRA adapters on one always-warm replica, or several models on one multi-model server, gives instant response at a fraction of the cost of separate replicas — and removes the cold-start question entirely rather than managing it.

saying these in an interview costs you the question

  • Applies scale-to-zero to interactive chat traffic
  • Assumes deleting the pod stops the GPU bill
  • Ignores that the GPU SKU may be unavailable on return
  • Expects HPA alone to scale a Deployment to zero
  • Sets a short idle timeout on bursty traffic

context