skip to content

A healthy primary in a volatile tier stalls for four seconds and the tier promotes a replica; what did that cost, and how do you choose the detection window?

level: seniorimportance: should knowfreq 46%

answer

  1. silence is silence, whatever caused it
  2. the window is a policy, not a detector
  3. a needless promotion is a real outage
  4. the stalled node comes back still believing
  5. measure the stall tail, then set above it

basics

~20 s

A stall longer than the detection window is indistinguishable from death, so the tier pays a full outage for nothing and must then demote a node that returns believing it leads. Size the window above the measured stall tail.

solid answer

~50 s

From the outside, a four-second stop-the-world pause, a stalled write to disk, a starved host or a saturated link all look exactly like a dead node: silence. Any detection window shorter than the stall promotes. The bill is a real outage — the detection, promotion and rediscovery windows in full — for a node that was about to serve, plus every caller reconnecting, plus the recent writes the old primary accepted and had not propagated, which do not survive the switch. Then the stalled node returns still believing it is primary and has to be reconfigured as a replica of the new one; whether that demotion is automatic or needs an operator varies by store and deployment. So the window is chosen against measured stalls, not a round number: set it above the tail you are unwilling to fail over on, and accept that the outage floor rises with it.

go deeper

for a junior

The key fact: a node that pauses and a node that died both look like silence from outside. That is why a failover mechanism can act when nothing was actually broken.

for a middle

Explain the trade in both directions — a short window promotes on ordinary stalls, a long one adds to every real outage — and say that the window is a decision about stalls rather than an accuracy setting.

for a senior

Price the needless promotion completely: an outage for nothing, a reconnection storm, unpropagated writes gone, and a returned node to demote. Then give the measurement method for choosing the window.

for a principal

Decide what the organisation promises and where the effort goes. If the required window is long because your stalls are long, the standing fix is the stalls, the host sharing and the entry sizes — not a timeout that papers over them.

## Why a stall and a death are the same event A deciding party has exactly one signal: whether the primary answers. Every one of these produces silence for the same few seconds and is not a failure at all: - a stop-the-world pause in the node's runtime, or a long allocation stall; - a synchronous write to disk that blocked, on a deployment that persists anything; - a host whose processor time was taken by a noisy neighbour or a backup job; - a saturated or briefly reconfigured network path; - the node itself busy on one very expensive operation that it will not interleave. There is no evidence that distinguishes these from death while they are happening; the distinction only exists afterwards, when the node answers again. So the detection window is not a detection accuracy setting. It is a declaration of how long a stall you are willing to wait out. ## What a needless promotion actually costs 1. **A full outage, for a node that was about to answer.** The caller pays the sum of all three windows — the silence, the reconfiguration, and the time before it stops addressing the old node — and gets nothing for it. 2. **Every caller reconnects.** Connection pools to the old primary are dead; the new primary takes the full connection load at once, from every instance, in the same second. 3. **Recent work vaporises.** Writes the old primary acknowledged but had not propagated do not survive the switch, and on this tier they are simply gone rather than replayable from a durable record. 4. **A node returns believing it leads.** The stalled primary wakes up with no idea that anything happened. It must be demoted — reconfigured as a replica of the new primary, discarding its divergent recent state. Whether that happens automatically on rejoin or waits for an operator varies, and a node that comes back undemoted will happily answer any caller that still holds its address. That last point is what makes an over-eager detection window worse than it looks: it does not just cost an outage, it creates a second node that has to be actively put back in its place. ## Choosing the window The method is measurement, not a round number: - **Measure the stalls you actually have.** Pause durations, disk stall durations, and the host's scheduling behaviour under your real load, at the tail rather than the average. - **Set the window above the tail you are unwilling to fail over on**, and be explicit about which tail that is — a window above the 99.9th percentile stall is a different promise from one above the worst stall ever seen. - **Accept the floor.** Whatever you choose is added to the outage of every *real* failure. There is no setting that is short for deaths and long for stalls. - **Fix the stalls instead, where you can.** If the tail that forces a long window is a pause, an oversized entry blocking the server, or a host you share badly, that is the cheaper thing to change. - **Require more than one voice.** An arrangement where several independent parties must agree already filters out a cut affecting one path, which is a different problem from a stall but removes a whole class of needless promotions. | Window too short | Window too long | |---|---| | Promotes on ordinary stalls | Extends every real outage by the difference | | Recent unpropagated writes lost for nothing | Callers stay pointed at a genuinely dead node | | A returning node must be demoted, sometimes by hand | A human may beat the machinery to the decision | | Connection storms on the new primary | The workload sees a longer unwritable period | ## What varies between stores Stores in this class all expose some silence threshold, and what such a default is *for* is to be comfortably longer than an ordinary hiccup — but the value, the granularity and whether it can differ per role are not the same anywhere, so quoting a number from one product's documentation is how a team ends up with a window that does not match its own stalls. Deployments also differ in whether a candidate's currency is checked before promotion, and in whether a returning node is demoted automatically. Ask those three questions of whatever you are running rather than assuming the arrangement you last used. ## The interview answer Say that the stall is indistinguishable from death by construction, so the window is a policy choice about stalls rather than a tuning knob for accuracy. Then price the mistake — an outage for nothing, a connection storm, unpropagated writes gone, and a node to demote — and give the method: measure the tail, set above it, and reduce the stalls if the required window is embarrassing.

  • The stalled primary comes back. What has to happen to it?
    It has to be demoted: reconfigured as a replica of the new primary and made to discard the divergent recent state it accepted before stalling. Some deployments do that automatically when the node rejoins; others leave it for an operator. Until it happens, the returned node is still answering any caller that kept its address.
  • Is there a detection window that is short for real failures and long for stalls?
    No — not from silence alone, which is all the deciding party has while it is happening. What you can do is reduce the stalls themselves, require agreement from several independent voices so a single cut path does not count as evidence, and accept the window you choose as a floor under every real outage.
  • Why does a needless promotion often hurt more than the outage length suggests?
    Because the recovery is not just time. Every caller's connections to the old primary are dead, so the new primary absorbs the whole reconnection load at once; the writes the old node had not propagated are gone; and a node now exists that believes it is primary and must be put back in its place.

saying these in an interview costs you the question

  • Thinks a stall can be told apart from a crash while it is happening
  • Sets the detection window from a round number rather than measured stalls
  • Treats a needless promotion as free because the data was already copied
  • Forgets the stalled node returns still believing it is primary
  • Assumes the returning node is always demoted automatically