You push a CDS update to a running Envoy and traffic does not move to the new cluster for several seconds. What is cluster warming, what blocks on it, and what happens if warming never completes?
answer
- a cluster is not usable until it has endpoints
- the old version keeps serving meanwhile
- the symptom is delay, not error
- first EDS response, or first DNS resolution
- initial_fetch_timeout is the only bound
basics
~20 sA cluster added or changed via CDS is not usable until it warms: Envoy waits for its first endpoint data — an initial EDS response or initial DNS resolution — before it can take traffic. During warming the previous version keeps serving, so an update looks delayed rather than broken.
solid answer
~50 sEnvoy never puts a half-built cluster into service. When CDS adds or modifies a cluster, Envoy constructs it and then *warms* it: for an EDS cluster it waits for the first `ClusterLoadAssignment`, for `STRICT_DNS`/`LOGICAL_DNS` it waits for the first resolution to return. Only then is the cluster swapped into the cluster manager. For an update, the old version of that cluster keeps serving throughout, so the visible effect is latency in the change, not an error — which is exactly the several-second gap you saw. Listeners warm too: one whose HCM uses RDS will not start accepting connections until its route configuration arrives. If warming never completes — the control plane never sends endpoints, or the DNS name does not resolve — the cluster stays stuck warming indefinitely unless the `ConfigSource` sets `initial_fetch_timeout`, which bounds the wait and lets Envoy proceed with what it has. Watch the `warming_clusters` gauge; a non-zero value that never returns to zero is the signature.
go deeper
Know the term: a new or updated cluster is not put into service until Envoy has its endpoints, so a config push can take effect a moment after it is accepted.
Say what warming waits on per discovery type — first EDS response, first DNS resolution, nothing for static — and that the previous version of an updated cluster keeps serving throughout.
Diagnose the stuck case: a cluster or listener warming forever because the resource it needs never arrives, visible as a warming gauge that never returns to zero, and bounded only by initial_fetch_timeout.
Own the startup-ordering policy for the fleet: whether proxies should block until fully configured or start degraded, what that means for orchestrator health checks and restart loops, and how a slow control plane at scale-out time turns into a capacity incident.
## Why warming exists Envoy's contract with itself is that a cluster in the cluster manager is usable. If a cluster could appear before it had any endpoints, every request routed to it in that window would fail immediately with no upstream to try — the proxy would be manufacturing errors out of its own update process. So cluster construction has two phases. Envoy builds the cluster object from the CDS resource, then holds it *warming* until it has enough to serve: - **EDS clusters** — until the first `ClusterLoadAssignment` for that cluster's EDS service name arrives. - **STRICT_DNS / LOGICAL_DNS clusters** — until the first DNS resolution completes. - **STATIC clusters** — nothing to wait for; endpoints came with the definition. Only when warming finishes is the cluster made available to workers. ## Add versus update The two cases feel different from outside. **Adding** a cluster: it simply does not exist until warm. A route pointing at it fails until it does — which is precisely why the ADS ordering rule sends clusters before the routes referencing them, giving warming a head start. **Updating** a cluster: the old version stays in service for the entire warming period and is swapped out atomically when the new one is ready. Nothing is dropped; existing connections to the old cluster's endpoints are not torn down mid-request. The observable symptom is purely *delay* — you pushed a change and, for a few seconds, the proxy behaved as though you had not. That is the answer to the interview scenario. "Traffic did not move" is not evidence the push failed. Check whether the cluster is warming before you go looking for a rejection. ## Listener warming, the sibling case The same idea applies one level up. A listener whose `http_connection_manager` gets its routes over RDS does not begin accepting connections until that route configuration arrives; a listener that references secrets over SDS waits for them too. This is the mechanism behind a very common startup complaint: a freshly started Envoy that is up, connected, and answering nothing on its port, because its listener is still warming for a route configuration the control plane has not sent. ## When warming never ends Warming has no built-in deadline. If the control plane never sends the endpoints — a resource-name mismatch between what the cluster asks for and what the server is keyed on is the classic cause — the cluster warms forever. The proxy is healthy, the stream is up, the CDS update was accepted, and the cluster is simply not in service. On a cold start with a listener warming on RDS, the proxy accepts nothing at all and health checks against it fail, so an orchestrator may keep restarting a proxy that is doing exactly what it was told. The bound is `initial_fetch_timeout` on the `ConfigSource`. It caps how long Envoy waits for the *first* response for a subscription; when it fires, Envoy proceeds with an empty resource rather than blocking, and the listener or cluster completes warming. Setting it is close to mandatory for anything that must start in a degraded environment: without it, "the control plane was slow" and "the control plane never answered" have the same symptom, forever. Be honest about the trade. `initial_fetch_timeout` firing means proceeding with *no* endpoints or *no* routes — the proxy starts and then fails requests, rather than never starting. Which failure you prefer depends on whether something upstream of the proxy interprets "not accepting connections" as "do not send me traffic". If it does, blocking is often the safer state. ## Confirming it Envoy exposes `cluster_manager.warming_clusters` as a gauge and `listener_manager.total_listeners_warming` for the listener case. In steady state both should be zero; a value that rises with each push and settles back is normal, and one that never settles is the fault. Dumping the proxy's effective configuration shows which clusters made it into service and which did not, and its per-cluster view shows whether any endpoints were ever learned. ## The related trap: warming is not health checking A warm cluster is one that has endpoints, not one whose endpoints are healthy. Envoy will finish warming a cluster whose endpoints all fail active health checks — the load assignment arrived, so warming's condition is met. Warming is about configuration completeness; health is about reachability. Conflating them leads people to expect warming to protect them from a bad deploy, which it will not do.
- A newly started Envoy is connected to its control plane but answers nothing on its listener port. What is the likely cause?Listener warming. A listener whose http_connection_manager fetches routes over RDS — or whose TLS context comes over SDS — does not begin accepting connections until that resource arrives. The proxy is healthy and the stream is up; it is simply waiting. Setting `initial_fetch_timeout` on the config source bounds that wait so the listener opens (with no routes) instead of blocking forever.
- What does setting initial_fetch_timeout actually trade away?It converts an indefinite block into a degraded start. When the timeout fires, Envoy stops waiting and proceeds with an empty resource — so the listener opens with no routes, or the cluster comes up with no endpoints, and requests fail instead of queueing behind a closed socket. That is better when something upstream needs the proxy to be reachable, and worse when a closed port correctly keeps traffic away.
- Does a cluster finish warming if all of its endpoints are failing health checks?Yes. Warming's condition is that endpoint data has been received, not that any endpoint is usable. A ClusterLoadAssignment listing ten hosts completes warming even if active health checking then marks all ten unhealthy. Warming is a configuration-completeness gate; health checking and outlier ejection are what keep traffic off bad hosts afterwards.
saying these in an interview costs you the question
- Thinks the old cluster is removed while the new one warms
- Assumes a cluster that finished warming has healthy endpoints
- Believes warming has a built-in timeout by default
- Reads a delayed traffic shift as evidence the push was rejected
- Confuses cluster warming with upstream connection pre-warming