Should leadership restoration to home nodes run continuously on its own, or only when an operator decides to run it?
answer
- cheap to fix, invisible to notice
- automation acts while nobody watches
- threshold, health gate, cooldown
- alert on every firing, publish the per-node view
basics
~20 sNeither extreme survives contact. Continuous restoration prevents skew accumulating but hands traffic back unattended, including to a node that just returned unwell. Operator-run restoration is predictable but relies on someone noticing an invisible drift. The workable policy is automatic above a threshold, with health preconditions and an alert.
solid answer
~50 sThe two ends trade a standing tax against an unattended action. Left to an operator, drift accumulates silently — nothing on a cluster-wide dashboard shows it — and the first symptom is a saturated node during a peak. Left running on its own, skew never builds, but handovers happen with nobody watching: a flapping node that returns is handed its traffic back and then loses it again, and a large correction can fire during an incident when the cluster can least absorb a wave of retries. What usually survives is the middle: restoration that fires only above a skew threshold, only when the home node is present and its copy is caught up, with a cooldown so a flapping node cannot cause repeated handovers, an alert when it fires, and a per-node view of units led that a human is expected to read after every planned change.
go deeper
Recall that leadership does not return home by itself, so something has to put it back — either scheduled automation or a person. Knowing that the choice exists is enough at this level.
Explain the trade in mechanism terms: continuous correction keeps each handover small but acts unattended, while operator-run correction is predictable but depends on someone reading a view that no alert points at.
Argue from the failure modes you have seen: the flapping node handed its traffic back and losing it again, the correction that fired mid-incident, the skew that went unnoticed until a peak. Name the guardrails that address each.
Take a position and define the standard: the threshold, the health precondition, the cooldown, the suppression rule, the alert, and the per-node view the estate must publish so the policy can be audited at all.
## The question behind the question Leadership drift is a fault with two unusual properties: it is invisible on every aggregate view, and correcting it is cheap. That combination is what makes the policy call interesting. An expensive, visible fault gets a runbook and an owner by itself. A cheap, invisible one either gets automated or gets forgotten, and both outcomes have a failure mode. ## What continuous restoration buys, and what it risks Buys: - **Skew never accumulates.** Drift is corrected minutes after the event that caused it, so headroom planned per node stays real. - **No dependence on noticing.** The one thing humans are worst at here — spotting a distribution problem on a totals dashboard — is taken out of the loop. - **Smaller corrections.** Continuous means each correction covers a handful of units, so the interruption burst stays small by construction. Risks: - **Handing traffic to a node that just returned unwell.** A process that is answering is not necessarily a host that is healthy. Restoration that only checks whether the copy is caught up can route most of a unit's work back to a machine about to fail again. - **Repeated handovers on a flapping node.** Every return triggers a restoration, every failure triggers a reassignment, and each one costs a brief interruption for the affected units. The automation amplifies the instability it is reacting to. - **Firing during an incident.** A wave of handovers adds retries and owner-map refreshes to a cluster that is already erroring, which is exactly when it can least absorb them. - **Invisible operations.** If restoration is not recorded and alerted, the latency bump it causes arrives with no explanation attached, and someone spends an hour on it. ## What operator-run restoration buys, and what it risks Buys **predictability**: the correction happens at a chosen moment, with a person watching, after the cluster's state has been looked at. For a small, closely operated estate that is a reasonable answer. Risks the thing the fault is built to exploit: **nobody looks**. Drift produces no alert, no error and no storage symptom. The realistic outcome is that the cluster runs skewed for months, absorbing it in spare capacity, until a peak or the loss of the busiest node converts it into an incident — at which point the skew is large and the correction is the big, bursty kind. ## A policy that survives contact 1. **Automate with a threshold**, expressed as the share of units one node leads relative to its fair share. Below it, do nothing; a perfectly even assignment is not worth an interruption. 2. **Gate on health, not on liveness.** Require the home node to be present, its copy caught up, and the host itself to look sound — not merely that a process answered. 3. **Add a cooldown per node**, so a node that keeps failing cannot trigger a restoration on every return. 4. **Correct in batches**, so the interruptions spread rather than coincide. 5. **Suppress during a declared incident or change freeze**, and make that suppression explicit rather than a side effect. 6. **Alert and record every firing**, so the latency bump has a cause attached and the frequency of firings becomes the signal it should be — frequent restorations mean something upstream is unstable. ## What has to be published either way Whatever the policy, the estate needs a per-node count of units led beside units held, and the expectation that it is read after every planned change to a cluster. Automation without that view is unfalsifiable: nobody can tell whether it is working. Manual correction without it is impossible. ## Where the question changes shape - **Designs with detached storage**, where a unit's records live on shared or remote storage, make ownership cheap to move because a node can serve without holding local records. Skew still happens — serving still concentrates — but the correction is cheaper still and the case for automating it is stronger. - **Designs where consumers compete for work** from a shared queue have no single serving node per unit, so the policy question does not arise for them in this form. - **Rented clusters** may run the restoration for you and not expose the choice. The policy question then becomes whether you can *observe* what it did, and what your provider's behaviour is during your incidents. ## The honest summary The cost of a correction is small and bounded; the cost of a missed correction is a machine running at several times its share with the largest blast radius in the cluster. That asymmetry argues for automating, and the risks argue for constraining the automation rather than withholding it. A principal is expected to say which way they lean and to name the guardrail that makes it safe, not to recite both columns.
- Why is a cooldown per node worth more than simply lowering the threshold?Because the expensive case is instability, not imbalance. A flapping node produces a restoration on every return and a reassignment on every failure, so each cycle costs interruptions and none of them lasts. A cooldown breaks that loop; a lower threshold would only make it fire sooner and more often.
- If restoration is automated, does anyone still need the per-node view of units led?Yes, more than before. Automation that nobody can observe is unfalsifiable: a stuck or suppressed restoration looks exactly like a working one on every aggregate dashboard. The per-node view is how you confirm the automation is doing its job, and how often it has had to.
- Should restoration be suppressed during an incident?By default yes, with a deliberate override. During a degradation the cluster is already producing retries, and a wave of handovers adds more at the worst moment. The exception is when the skew itself is the incident — then correcting it is the treatment, and it should be done in small batches, deliberately.
saying these in an interview costs you the question
- Automatic restoration is always safer than leaving it to an operator
- Manual is fine because someone will notice the imbalance
- Automatic restoration cannot make an ongoing incident worse
- A node is healthy enough to serve because its process answered
- Skew resolves itself once traffic returns to normal
- If restoration is automated, nobody needs the per-node view