skip to content

Automatic Failover

Something has to decide a node is gone and promote another: the window before anything happens, why one observer's opinion is never enough alone, and the client that must find the new address.

on this pageshow

questions

5

In a volatile tier with a primary and two replicas of the whole keyspace, why must a replica be promoted before writes resume?

level: juniorimportance: must knowfreq 62%

answer

  1. copies are not availability
  2. replication and failover are different features
  3. a replica is not the write address
  4. someone outside the dead node must act

basics

~20 s

A replica holds a copy of the keyspace but is not the address that accepts writes, so writes resume only after some deciding party promotes one of them to primary. That promotion is a deliberate step, and it takes time.

solid answer

~40 s

Replication gives you copies, not availability. The replicas already hold the keyspace, so no data has to be moved at failure time — but while a node is a replica it is not the write address. Most stores of this class refuse writes on a replica outright; a few allow a writable replica, and there the writes you make land in a copy nobody else will ever see. So something outside the dead node has to notice it is gone, choose a replica that is current enough, reconfigure it as primary and point the other copies at it. That machinery — surviving peers, dedicated observer processes, or a controller outside the data plane — is what "automatic failover" names. Until it acts, the tier is unwritable.

go deeper

for a junior

Remember the two-part fact: the copy already exists, and the copy is not where writes go. Writes come back only after a replica has been promoted to primary, and that promotion is something a separate piece of machinery does.

for a middle

Be able to list what promotion actually changes — which node accepts writes, which nodes copy from it, and how callers end up addressing it — and to say that replication alone provides none of that.

for a senior

Say which deciding party your deployment uses and what happens if nothing is deployed to decide: the copies are fine, the tier is unwritable, and a human is the failover mechanism. Also know whether your store refuses writes on a replica or quietly accepts them.

for a principal

The call is whether this tier deserves automatic promotion at all. Automatic failover adds a component that can be wrong; manual promotion adds a page and a long outage. Decide per workload, on what the keys hold and who is woken up.

## What a replica is, and what it is not A **replica** in a volatile tier is a node holding a copy of the whole keyspace, kept current from the **primary**. Two consequences get conflated. The first is genuinely good news: the copy is already there, so nothing has to be shipped across the network at the moment of failure. The second is the one candidates miss: holding the copy does not make the node the place callers write to. While it is a replica it is following someone else's stream of changes, and accepting an independent write would make its copy diverge from the thing it exists to mirror. Stores in this class enforce that differently, and the difference matters when you are reasoning about a specific deployment: - most refuse writes on a node while it is configured as a replica, so a caller that tries gets an error rather than a surprise; - some allow a **writable replica**, and there a write succeeds locally and is invisible to the primary and to every other copy — which is worse than an error, because nothing tells you; - whether a replica answers *reads* at all is a separate deployment choice, and plenty of deployments send every read to the primary. So: a replica is a copy, not a hot standby write address. ## Why promotion has to be a decision For the tier to keep serving writes without a human, four things must happen, in order, and each of them needs an actor that is not the failed node: 1. **Decide the primary is gone.** Something has to conclude, from silence, that the node is not coming back soon enough to wait for. 2. **Choose a candidate.** More than one replica may be eligible, and they are not equally current. 3. **Reconfigure.** The chosen node becomes primary and the surviving copies are told to follow it instead of the dead one. 4. **Get the callers there.** The callers must end up addressing the new node, whether by re-resolving a name, by asking the deciding party, or because a stable address in front of the tier was re-pointed. Who performs steps 1 to 3 is one of three arrangements: the **surviving peer nodes** themselves, a set of **dedicated observer processes** deployed alongside the tier, or a **controller outside the data plane** in a managed offering. Replication gives you none of these. A tier with three copies and nothing watching them is a tier that pages a human — which is a perfectly legitimate design, as long as the team knows that is what they bought. ## What the callers observe in the meantime - Writes fail or time out for the whole interval between the failure and the moment the new primary is both serving and being addressed. - Reads may still be answered from the copies where the deployment routes reads to replicas; where reads go to the primary, reads fail too. - Callers that are still holding the old address keep failing *after* a textbook-clean promotion, until they re-resolve. - Recent writes the old primary accepted but had not yet propagated do not survive the switch — how large that window is, and how to size it, is a question of its own. The promotion itself is usually the quickest part. The interval before anything concludes the node is gone, and the interval before the last caller stops addressing it, are both typically longer. ## What varies between stores and deployments | Question | One common arrangement | A real alternative | |---|---|---| | Can a replica accept a write? | It refuses while configured as a replica | A writable replica is allowed, and its writes stay isolated | | Who decides and promotes? | The surviving peer nodes, among themselves | Dedicated observer processes, or a controller outside the data plane | | How current must a candidate be? | The most current surviving copy is preferred, and a far-behind one may be skipped or promotion refused | Any surviving copy is promoted, and currency is not checked | | What happens when the old node returns? | It is reconfigured automatically as a replica of the new primary | It comes back believing it is primary until an operator intervenes | Because all four rows vary, the honest interview answer names the mechanism and then says which arrangement you are describing. ## How to answer this in an interview Say the three things in order: the copies exist, so this is not a data-movement problem; a replica is not a write address, so something must promote one; and the promotion is performed by a named deciding party, which is a component you have to deploy and size deliberately. Then note that the outage a caller experiences is longer than the promotion, because it also contains the time before anything noticed and the time before the caller stopped addressing the dead node.

  • Does adding a third replica make the tier fail over faster?
    No. More copies add promotion candidates and, where reads are routed to them, read capacity — but the intervals are set by how long the deciding party waits out silence, how long reconfiguration takes, and how long callers keep the old address. What does have to be sized deliberately is the deciding party itself, not the number of copies.
  • What makes one surviving replica a better promotion candidate than another?
    How much of the primary's work it had actually received, plus any operator-set preference — some deployments rank candidates by how current they are and will skip, or refuse to promote, a copy that is too far behind. Others promote any survivor. If your store does not check currency, the choice of candidate is effectively arbitrary, which is worth knowing before you need it.

saying these in an interview costs you the question

  • Thinks replicas start taking writes the instant the primary stops answering
  • Says the application can simply write to any replica instead
  • Believes promotion is instant because the copy is already there
  • Treats replication and automatic failover as one feature you get together
  • Assumes every surviving replica holds exactly what the primary held
open as a page

An automatic failover of a volatile tier's primary is often reported as one number; which separate windows does that total outage actually consist of?

level: middleimportance: must knowfreq 58%

basics

~10 s

A failover outage is the sum of three windows: detection, before anything concludes the primary is gone; promotion, while a replica is reconfigured; and rediscovery, until the last caller stops addressing the dead node.

open as a page

A volatile tier's promotion finished in seconds, yet one service kept addressing the dead primary for an hour — what did it get wrong?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Promotion only moves which node accepts writes; it does not reach into callers. A service that resolved the write address once at start-up and never again keeps talking to the dead node, however clean the failover was.

open as a page

A healthy primary in a volatile tier stalls for four seconds and the tier promotes a replica; what did that cost, and how do you choose the detection window?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A stall longer than the detection window is indistinguishable from death, so the tier pays a full outage for nothing and must then demote a node that returns believing it leads. Size the window above the measured stall tail.

open as a page

In a volatile tier that fails over without a human, which parties can be entitled to declare the primary dead, and why is one observer never enough?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Three parties can own the decision: the surviving peer nodes, a set of dedicated observer processes, or a controller outside the data plane. A single observer is never enough — it cannot tell a dead node from its own severed link.

open as a page