skip to content

A cluster's control loop stops for an hour, then resumes: why do workloads converge instead of staying stuck on missed changes?

level: middleimportance: must knowfreq 52%

answer

  1. what the pass reads, not what happened
  2. state is the input, the event is a hint
  3. a lost notification costs latency only
  4. a periodic full re-read sits behind it
  5. level-triggered, not edge-triggered

basics

~20 s

Each pass re-derives everything from the current declared state and the current live state, so no notification needs to have been seen. A loop resuming after an hour reads the world as it is now and closes whatever difference it finds.

solid answer

~40 s

Cluster control loops are **level-triggered**: the input to a pass is the current state of the world, not the event that changed it. Change notifications are used as a hint about *when* to run, and they are backed by a periodic full re-read, so a dropped, duplicated or delayed notification costs latency rather than correctness. That is why an hour-long gap strands nothing: on its first pass back, the loop reads the declared state and the live state fresh, computes one difference, and acts on it. An **edge-triggered** design does the opposite — it acts on the transition, so a lost notification loses the change permanently and the system needs a separate repair path to find what it missed.

go deeper

for a junior

Recall that the loop compares the whole current state each time it runs, so it does not need to have witnessed a change in order to correct it.

for a middle

Explain level-triggered against edge-triggered: the input is state rather than a transition, which is why a lost or repeated notification changes only how quickly the loop converges.

for a senior

Show the operational consequence: you diagnose from state — what did the last pass read and compute — rather than hunting a missing message, and you alert on the age of a difference instead of on delivery errors.

for a principal

Weigh what the design costs: full re-reads scale with the estate, so you are trading read load and eventual convergence against never needing a message bus to be reliable.

## Two ways to build a loop Any automation that keeps a system in shape has to decide what its input is. There are two answers, and they behave completely differently under failure. - **Edge-triggered.** The input is a *transition*: something changed, here is the change, react to it. The work is proportional to the number of changes, and each change must be delivered exactly once or the system is wrong. - **Level-triggered.** The input is the *current state*: here is what is declared, here is what exists, compute the difference. Work is proportional to the size of the state, and delivery of change notifications is an optimisation rather than a correctness requirement. Cluster control loops are built the second way, and the hour-long gap is the clearest demonstration of why. ## What the first pass back actually reads When the loop resumes, it does not look for a backlog. It performs the same three steps it performs every time: read the declared state, read the live state, act on the difference. Everything that happened during the hour is already baked into those two readings. | Event during the gap | What an edge-triggered design needs | What the level-triggered loop sees on resume | |---|---|---| | Declared count changed 6 → 8 | the change notification, delivered once | declared 8 against live 6: start two | | Declared count changed 6 → 9 → 8 | three notifications, in order | declared 8 against live 6: start two | | A copy was removed by hand | a deletion notification | live 5 against declared 6: start one | | Nothing changed | no notification | no difference: do nothing | Notice the third row: an intermediate value of 9 was never acted on and never will be. That is not a bug — a level-triggered loop makes no promise to have passed through every state, only to arrive at the declared one. Anything that needs each intermediate value observed needs a different mechanism entirely. ## So why watch for changes at all? Because a loop that only re-read everything on a slow timer would be *correct* and *slow*. Platforms therefore combine the two: 1. **Change notifications** collapse latency — a declaration edit is usually acted on in well under a second instead of at the next sweep. 2. **A periodic full re-read** (often called a resync) is the correctness guarantee — it re-derives every difference from scratch, regardless of what was or was not delivered. 3. **The pass itself ignores the notification's contents.** The notification says *this object may be interesting*; the pass then reads the object. That is what makes duplicates free: two notifications for the same object produce two passes, and the second finds no difference. This also explains why rapid successive edits do not produce a burst of contradictory actions. Platforms collapse pending work per object, so five edits in ten seconds typically produce one or two passes against whatever the state says at that moment. ## What this buys and what it costs **Buys:** - No message bus needs to be reliable for the platform to be correct. - Restarting a loop is routine — there is no state to recover and no backlog to replay. - A difference introduced outside the platform's knowledge, by any route, is still found. **Costs:** - Full re-reads scale with the size of the estate, so read load grows with object count, and very large installations tune the sweep interval carefully. Designs differ in how aggressively they cache the live state to make this cheap. - Convergence is *eventual*. There is a window between a change and the pass that acts on it, and no ordering promise across objects. - Because nothing fails when a notification is lost, a slow loop is invisible unless you measure it — the useful signal is the age of the last successful pass per object, not an error count. ## Debugging in this model The practical consequence for an on-call engineer is a changed first question. You do not hunt for a missing event. You ask **what the last pass read and what it computed**: if the declared state says one thing and the live state still says another, either no pass has run recently, or a pass ran and could not act. Those are two different failures with two different fixes, and both are visible from state rather than from a message trail.

  • If current state is the input, why do platforms bother watching for changes?
    For latency. A loop that only re-read everything on a timer would be correct but slow, so notifications shorten the gap between an edit and the pass that acts on it, often to well under a second. The periodic full re-read stays as the correctness guarantee, which is why losing a notification costs speed rather than convergence.
  • How do you tell that a level-triggered loop is falling behind?
    Not from errors, because a delayed pass raises none. You measure the age of the last successful pass per object and the depth of the pending work queue, and you alert on a declared-against-live difference that persists longer than a few pass intervals. Without those, a loop that has quietly stopped looks exactly like a cluster with nothing to do.

A thermostat keeps reading the room temperature and heats whenever it is below the target; it does not depend on having noticed the moment someone opened a window. A switch you flip once is the other design — miss the moment and nothing ever corrects it.

saying these in an interview costs you the question

  • Says the loop replays the events it missed while it was down
  • Assumes a change must be observed as it happens or it is lost
  • Thinks the loop applies every intermediate value a rapid series of edits passed through
  • Believes change notifications are the source of truth rather than a timing hint
  • Claims a duplicate notification makes the corrective action happen twice