skip to content

When a workload is allowed to scale to zero copies, what does the first request after an idle period wait for?

level: middleimportance: should knowfreq 42%

answer

  1. no copies, no per-copy signal
  2. the trigger must come from outside
  3. something holds the request meanwhile
  4. pull, start, ready, warm
  5. the floor decides idle cost and first latency

basics

~20 s

It waits for the whole cold path: something in the request path must notice it, raise the count from zero, and hold the request while a host is chosen, the image is pulled if absent, the process starts and the readiness check passes. That is the cold start.

solid answer

~40 s

At zero copies the workload emits no per-copy signal, so the usual measurement cannot restart it - the trigger has to come from outside: a component in the request path, or a queue whose depth is watched independently. That component accepts the request and holds it, or enqueues the work, while the count goes from zero to one. The request then pays placement, an image pull if the chosen host has no cached copy, process start-up, readiness and any warm-up - seconds to tens of seconds. Scale to zero is therefore a fit for spiky, latency-tolerant or asynchronous work. For a latency-sensitive request path, the usual alternative is a floor of one warm copy: you pay for it continuously and you no longer scale to zero.

go deeper

for a junior

Remember that a workload at zero copies has nothing running at all, so the first request afterwards waits for a copy to be started from scratch before it can be answered.

for a middle

Explain why the trigger has to come from outside the workload at zero, and list the cold path in order: placement, image pull, start, readiness, warm-up.

for a senior

Judge the fit rather than the feature: identify which workloads tolerate the wait, attack the dominant term in the cold path, and name the availability risk a floor of zero concentrates on the return.

for a principal

Decide where in the estate the trade is acceptable, since scale to zero converts a steady predictable cost into occasional user-visible latency and a new dependency on the registry being reachable.

## At zero, the normal signal does not exist Almost every scaling signal is denominated per copy - utilisation per copy, in-flight requests per copy, queued items per copy. At a count of zero there are no copies to measure, and a per-copy average over an empty set is undefined. So scale to zero is not simply "the loop kept subtracting until it reached zero"; it is a distinct mechanism with a **separate trigger for the trip back up**, and that trigger has to live outside the workload: - **In the request path.** A component that is always running accepts the connection, sees that the workload has no copies, asks for one, and holds the request until a copy is ready. The client experiences it as one very slow request. - **In the work source.** For queue-fed work, the depth of the queue is watched by something independent of the workload. A non-zero depth raises the count from zero, and the work simply waits in the queue meanwhile. The second shape is much more forgiving, because nothing is holding a connection open and there is no user watching a spinner. ## What the first request actually pays for The cold path is the same chain any new copy walks, with none of it hidden behind already-running capacity: 1. **Trigger and decision** - the arriving request or queued item is noticed and the desired count goes from 0 to 1. 2. **Placement** - a host with room for the copy's reservation is chosen. 3. **Image fetch** - if that host has no cached copy of the image, its layers are pulled from a registry. This is frequently the largest single term, and it is the one most sensitive to image size. 4. **Process start** - the container starts and the process initialises. 5. **Readiness** - the copy reports it can serve, and only then is it routed to. 6. **Warm-up** - empty caches, cold connection pools and lazily loaded configuration mean the first requests are slower even after routing begins. From a cold host that is tens of seconds; from a host with the image already cached and a fast-starting process it can be a couple of seconds. The spread is enormous, which is why "how long is your cold start" is unanswerable in general and very answerable for a specific workload. ## The trade this is actually making | Choice | Idle cost | First-request latency | Fits | |---|---|---|---| | Floor of zero | Nothing | Full cold start | Spiky, rare, asynchronous or latency-tolerant work | | Floor of one | One copy, always | Warm | Steady user-facing paths where a cold start is not acceptable | | Floor of two or more | Two copies, always | Warm, and survives losing one | Anything where availability matters at idle as much as at peak | The floor is one number and it decides all three columns. The common mistake is to treat scale to zero as free efficiency: it moves a continuous, predictable cost into an occasional, user-visible latency cost, and whether that is a good trade depends entirely on who is waiting. ## Making the cold start smaller If the shape fits but the wait is too long, attack the chain in order of size: - **Shrink the image.** Less to pull is directly less to wait for, and a minimal base image with no shell or package manager can be a fraction of a general-purpose one. - **Keep the image cached.** A host that already holds the layers skips the largest term entirely; some platforms can keep images warm on hosts for exactly this reason. - **Start faster.** Defer optional work, open connections in parallel, and avoid reading large state at start-up. - **Make the wait asynchronous.** If the caller can accept an acknowledgement and collect the result later, the cold start stops being a latency problem and becomes a throughput one. ## The failure worth naming Scaling from zero concentrates risk on the moment of return. If the image can no longer be pulled - the registry is unreachable, credentials expired, a rate limit applies - a workload at zero copies is *completely* down, whereas a workload with a warm floor would have kept serving on copies whose images were pulled long ago. The same is true of a host pool that has no room: at a floor of one, you are degraded; at a floor of zero, you are off. That asymmetry, not the arithmetic, is usually the reason a user-facing path keeps a floor.

  • Why can the usual per-copy signal not bring a workload back from zero?
    Because there is nothing left to measure. A per-copy average over zero copies is undefined, and a utilisation reading cannot be produced by a workload that is not running. The trip back up therefore needs an external trigger - a component in the request path that sees the arriving request, or an independent watcher of the queue that feeds the work.
  • What makes queue-fed work a much better fit for scale to zero than a synchronous request path?
    Nothing is waiting on an open connection. The item sits in the queue while the first copy starts, and once it is serving, the backlog drains - the cold start shows up as end-to-end latency on a few items rather than as a user staring at a slow page or a client timing out mid-request.
  • Beyond latency, what risk does a floor of zero add that a floor of one does not?
    It makes the return path a single point of failure. A workload at zero must pull its image and find a host in order to exist at all, so an unreachable registry, an expired credential, a registry rate limit or a full host pool means total unavailability. A warm copy is already running on an image it pulled long ago and keeps serving through all of those.

saying these in an interview costs you the question

  • Thinks the per-copy signal can still drive scaling at a count of zero
  • Treats scale to zero as free, ignoring the latency it moves onto the first caller
  • Assumes the cold start is just process start-up, forgetting the image pull
  • Applies scale to zero to a latency-sensitive user-facing request path
  • Forgets that a workload at zero is fully down if its image cannot be pulled
  • Believes a copy serves traffic as soon as its process is running