A host running three device-telemetry ingester replicas stops reporting to the control plane - why does the platform wait before replacing them?
answer
- silence has several causes
- the control plane only hears reports
- suspect, then declared gone
- timeout, then the count drops
- fast recovery against false replacement
basics
~20 sSilence is ambiguous: a host that stops reporting may be dead, or its agent may have crashed, or only the network between it and the control plane is broken. Platforms wait out an unreachable-node timeout before declaring its workloads gone and replacing them.
solid answer
~50 sThe control plane learns about a host only through periodic reports from the agent on it. When those stop, the control plane knows one thing - it stopped hearing - and that has several causes: the machine is gone, the agent died while the containers keep running, or the network path broke while everything on the host is fine. Because those look identical from the deciding side, platforms apply an **unreachable-node timeout**: after some period without a report, the host is marked unreachable and its replicas are declared gone, which is what finally creates the count gap the loop fills. Until the timeout expires the replicas are still counted as running, so nothing is replaced and the workload simply runs short. The timeout is a tuning knob between recovering fast and replacing copies that were never actually dead.
code
pseudocode · 13 linesfor each host in cluster:
silentFor = now - host.lastReportReceived
if silentFor > unreachableTimeout:
host.state = UNREACHABLE
for each replica placed on host:
declare replica gone # bookkeeping only; nothing is stopped
# observed count now drops below declared, so the replica
# loop will ask for the missing copies on its next pass
else:
# still inside the wait: the replicas are counted as running,
# so the loop sees no gap and starts nothing
keep replica accounting unchangedgo deeper
Remember that the deciding half of a cluster only knows what each host reports, and that a host going quiet does not immediately mean its work is gone.
Explain the ambiguity of silence, name the timeout that resolves it, and describe what is true of the workloads during the wait: still counted, possibly still running, not yet replaced.
Reason about the trade-off with numbers - how long the workload runs short, what a shortened timeout costs in churn and overlap, and why sizing for a host loss often beats tuning the wait down.
Treat the timeout as a cluster-wide default that different workloads want set differently, and be explicit about which workloads you would give a different tolerance and what you are trading for it.
## What the control plane can actually see The deciding half of a cluster does not observe hosts directly. An agent on each host reports in on a period - here is this machine, here is what it has capacity for, here is what is running on it. Everything the control plane believes about a host is a recent copy of one of those reports. So when the reports stop, the only fact available is **"I have not heard from that host since T"**. That single fact is produced by at least four different worlds: - the machine lost power, or its disk, and everything on it is genuinely gone; - the agent crashed or is wedged, while the containers it started keep running and serving; - the network between that host and the control plane is broken, while the host and its clients are fine; - the control plane's own side is degraded and is not processing reports it did receive. In three of those four, the workloads are still alive. Replacing them immediately would mean running more copies than declared, and the extra copies would compete with the originals for anything that assumes a bounded number of writers or consumers. ## The waiting period Every platform therefore has some **unreachable-node timeout**: a period after the last report during which the host is suspect but its workloads are still counted. When it expires, the control plane marks the host unreachable and declares the workloads on it gone. That is a bookkeeping decision, not a stop - nothing on a silent host can be stopped by a control plane that cannot reach it. But it is the decision that makes the observed count drop, and the count gap is what the replica loop fills. The sequence, with the three ingester copies: 1. Last report arrives at T. Nine replicas of the ingester are running, three of them on this host. 2. Between T and T plus the timeout, the control plane still counts nine. No replacement is started. Whether the three still receive traffic is decided by the separate serving signal, not by this timeout. 3. At T plus the timeout, the host is marked unreachable and its three copies are declared gone. The observed count drops to six against a declared nine. 4. The loop asks for three more copies, which are placed on hosts that are still reporting. 5. If the original host comes back later, its agent reconciles with the control plane and stops the copies it is no longer supposed to be running. ## Why not set the timeout to zero This is the trade-off the question is really about, and it has a bad end at both extremes. | Shorter timeout | Longer timeout | |---|---| | Capacity is restored sooner after a real host loss | The workload runs below its declared count for longer | | A brief network blip triggers replacement of copies that were never dead | Transient blips resolve with nothing replaced | | Widens the window where old and new copies run at once | Narrows that window, at the cost of slower recovery | | Churn costs real work: image fetches, cold starts, refilled caches | Recovery from a genuine failure is delayed by exactly the timeout | Platforms differ in the default they ship and in how finely it can be set - some expose it per cluster, some allow a per-workload tolerance for how long a copy may sit on an unreachable host. What is common to all of them is *that* there is a wait, and why. ## What the wait does not mean Three misreadings are worth naming: - **It is not a health check.** A check that judges whether a copy is serving correctly is a different mechanism with a different consequence; this timeout is about whether the host is talking to the control plane at all. - **It is not an in-place restart.** Nothing is retried on the silent host. When the timeout expires, the copies are written off and new ones are started elsewhere. - **It is not a promise the old copies stopped.** During the whole window, and beyond it until the host is reachable again, those three copies may be running and doing their work. The control plane merely stopped counting them. ## The operational shape of it The user-visible symptom of this design is a service that is degraded for a fixed period and then recovers on its own. If the ingester was sized so that losing a third of its copies hurts, the outage lasts roughly the timeout plus the start-up time of the replacements - and no amount of watching the replacement logs explains the first part of it, because during that part the platform believes nothing is wrong. Sizing for a host loss, rather than tuning the timeout down, is usually the better answer.
- During the wait, does traffic keep going to the three copies on the silent host?That is decided by a different signal. Routing follows whether each copy reports itself able to serve, and where that reporting flows through the same broken path it will also stop, pulling the copies out of routing well before the unreachable timeout expires. So traffic often stops long before replacement starts, which is why the two mechanisms are worth naming separately.
- What happens to the three original copies when the host finally comes back?Its agent re-establishes contact, compares what it is running against what the control plane says it should be running, and stops the copies that are no longer assigned to it. That reconciliation is the only thing that actually ends the overlap, and it can only happen once the host is reachable again.
- A team lowers the timeout sharply after a slow recovery. What should they expect to see next?More replacements triggered by transient network trouble, each paying an image fetch and a cold start, and a wider window in which a partitioned host's copies run alongside their replacements. If the original complaint was the length of the degradation, sizing the workload to survive a host loss usually helps more than shortening the wait.
A ship that misses one radio check-in is not declared lost, and one that misses them all week is not still under way. The timeout is where a fleet operator draws that line.
saying these in an interview costs you the question
- Says the control plane detects a dead host instantly
- Assumes silence proves the machine is down
- Thinks the replicas are stopped the moment the host goes quiet
- Wants the timeout at zero for faster recovery
- Believes a shorter timeout has no downside