skip to content

An automatic failover of a volatile tier's primary is often reported as one number; which separate windows does that total outage actually consist of?

level: middleimportance: must knowfreq 58%

answer

  1. one reported number, three real clocks
  2. notice, promote, rediscover
  3. the outage is their sum, not the largest
  4. the caller-side clock is the forgotten one

basics

~10 s

A failover outage is the sum of three windows: detection, before anything concludes the primary is gone; promotion, while a replica is reconfigured; and rediscovery, until the last caller stops addressing the dead node.

solid answer

~50 s

Teams quote one failover number, and it is usually the middle clock. Three run in sequence. The **detection window** is how long the deciding party waits out silence before concluding the primary is gone — a timeout somebody chose. The **promotion window** is how long it takes, once decided, to reconfigure a chosen replica as primary and point the surviving copies at it. The **rediscovery window** is how long callers keep addressing the dead node, bounded by connection pools, cached address resolution, and how a caller learns of a change at all. The outage a caller reports is the sum. Which window dominates depends on the arrangement: behind a stable address that a proxy or controller re-points, rediscovery can be near zero; where each caller resolves the address itself, it routinely dwarfs the other two.

go deeper

for a junior

Learn the three names and their order: something notices, something promotes, callers find the new address. The outage is all three added together, not just the promotion.

for a middle

Be able to say which of the three a given timeout controls, and why the promotion is usually the shortest of them. Name what bounds the caller-side window: connection pools and cached address resolution.

for a senior

Bring measurement. Say that the only way to know the rediscovery window is a deliberate failover drill with every caller watched, and give the reason a team that halved its detection timeout saw no improvement.

for a principal

Set what the organisation promises. A stated failover objective is meaningless unless it covers all three windows and names who owns each — the store team owns two of them, and the calling teams own the third.

## Why one number is the wrong unit Ask how long a failover takes and you will usually be told the time it takes to promote a replica. That is the interval an operator can see in one place, and it is almost never what a caller experienced. Three intervals run back to back, each controlled by a different thing, and only their sum is the outage. ## The detection window Something has to conclude, from silence, that the primary is gone. Whatever the deciding party is — the surviving peers, a set of dedicated observer processes, or a controller outside the data plane — it waits some configured period of unanswered probes before it will act. Stores in this class expose such a timeout precisely so you can set it; what a default is *for* is to be long enough that an ordinary hiccup does not trigger a promotion. This window is bounded below by the longest stall you are prepared to survive without failing over, and the fact that the node is already not answering during it means it is pure outage. It is also the only one of the three that a false positive can start from nothing. ## The promotion window Once the decision is made, a candidate is chosen, reconfigured as primary, and the surviving copies are told to follow it. Some deployments add a check here — the candidate must be current enough, or an operator preference decides between two candidates — and that check costs time. This is the interval most people mean when they quote a failover number, and it is frequently the smallest of the three. ## The rediscovery window The new primary can be accepting writes while every caller is still sending to the dead one. How long that lasts depends entirely on how the caller learns the write address: - a caller holding its own copy of the address refreshes it when it refreshes, which may be on error, on a timer, or never; - a caller resolving a name is bounded by how long its runtime caches the resolution, which is often longer than anyone assumes; - a caller behind an intervening proxy or a stable managed endpoint does not rediscover at all — the address never changed, and the re-pointing happened out of its sight. This is the window operators forget, and the one that most often turns a clean ten-second promotion into a four-minute incident. ## Putting a number on it | Window | Who controls it | What shortens it | What shortening it costs | |---|---|---|---| | Detection | The deciding party's configured timeout | A shorter silence threshold | Failing over on ordinary stalls that would have recovered | | Promotion | The store and the deciding party | Fewer eligibility checks on the candidate | Promoting a copy that is further behind than you would like | | Rediscovery | The caller, or whatever fronts the tier | Re-resolving on failure; a stable address something else re-points | Caller-side work, or a component in the data path to operate | The useful habit in an interview, and in a post-incident review, is to say which of the three your timeout actually controls. A team that halves its detection window and still reports the same outage has usually been paying for rediscovery all along. ## The question to ask of any deployment For each window, ask: what sets it, who can change it, and have we ever measured it? The third question is the one that fails. Detection and promotion can be read out of configuration; rediscovery can only be measured by deliberately failing the primary over and watching how long each caller takes to succeed again. A team that has never run that drill does not know its own failover time, whatever the configuration says. And note what all three have in common on this tier specifically: none of them is a replayable gap in a durable record. The tier is unwritable for the whole sum, and the writes callers attempted during it are simply not there afterwards unless the callers themselves retried. ## One failover, three separately controlled windows. The promotion took two seconds; the outage a caller reports is twelve. The numbers are illustrative — the shape is the point ``` t0 primary stops answering; callers' writes begin failing t0 .. +5s detection window - the deciding party waits out the silence, then concludes the primary is gone t0+5s a surviving replica is chosen as the candidate +5s .. +7s promotion window - candidate reconfigured as primary, remaining copies pointed at it; it now accepts writes +7s .. +12s rediscovery window - callers still holding the old address keep failing, then re-resolve and reconnect t0+12s the last caller is writing to the new primary ```

  • Which of the three windows can you not shorten from the store side?
    Rediscovery, in the arrangements where the caller holds the address itself — the store has no way to make a caller ask again. You shorten it either in the callers, by re-resolving on failure instead of at start-up, or by putting something in front that keeps one stable address and re-points it. That second choice moves the problem into a component you then have to operate.
  • Why does halving the detection window often not halve the measured outage?
    Because the detection window is only one of three terms. If rediscovery is the largest, the caller-visible outage barely moves, and you have bought a much higher chance of promoting on an ordinary stall. Measure all three before tuning one, and drill a real failover rather than reading the configuration.
  • Do the three windows still exist when a managed controller performs the failover?
    Yes, but you may not be able to see or set them. The controller has its own silence threshold and its own reconfiguration time, and the rediscovery window is usually collapsed by keeping one stable address that is re-pointed for you. The trade is that the numbers are the provider's, not yours, so you measure them rather than configure them.

A shop with one open till. The queue stops moving; it takes a while before a manager notices rather than assuming the customer is slow; then a second till has to be opened and staffed; and then the people standing in the dead queue have to realise and walk over. A customer's wait is all three added together, and shouting at the manager to notice faster does nothing about the last part.

saying these in an interview costs you the question

  • Quotes the promotion time as the whole failover time
  • Assumes callers reach the new primary the moment it is promoted
  • Thinks lowering one timeout shortens the caller-visible outage
  • Believes the tier is writable during detection because replicas are up
  • Never measured a failover, only read the configured timeouts