In a volatile tier with a primary and two replicas of the whole keyspace, why must a replica be promoted before writes resume?
answer
- copies are not availability
- replication and failover are different features
- a replica is not the write address
- someone outside the dead node must act
basics
~20 sA replica holds a copy of the keyspace but is not the address that accepts writes, so writes resume only after some deciding party promotes one of them to primary. That promotion is a deliberate step, and it takes time.
solid answer
~40 sReplication gives you copies, not availability. The replicas already hold the keyspace, so no data has to be moved at failure time — but while a node is a replica it is not the write address. Most stores of this class refuse writes on a replica outright; a few allow a writable replica, and there the writes you make land in a copy nobody else will ever see. So something outside the dead node has to notice it is gone, choose a replica that is current enough, reconfigure it as primary and point the other copies at it. That machinery — surviving peers, dedicated observer processes, or a controller outside the data plane — is what "automatic failover" names. Until it acts, the tier is unwritable.
go deeper
Remember the two-part fact: the copy already exists, and the copy is not where writes go. Writes come back only after a replica has been promoted to primary, and that promotion is something a separate piece of machinery does.
Be able to list what promotion actually changes — which node accepts writes, which nodes copy from it, and how callers end up addressing it — and to say that replication alone provides none of that.
Say which deciding party your deployment uses and what happens if nothing is deployed to decide: the copies are fine, the tier is unwritable, and a human is the failover mechanism. Also know whether your store refuses writes on a replica or quietly accepts them.
The call is whether this tier deserves automatic promotion at all. Automatic failover adds a component that can be wrong; manual promotion adds a page and a long outage. Decide per workload, on what the keys hold and who is woken up.
## What a replica is, and what it is not A **replica** in a volatile tier is a node holding a copy of the whole keyspace, kept current from the **primary**. Two consequences get conflated. The first is genuinely good news: the copy is already there, so nothing has to be shipped across the network at the moment of failure. The second is the one candidates miss: holding the copy does not make the node the place callers write to. While it is a replica it is following someone else's stream of changes, and accepting an independent write would make its copy diverge from the thing it exists to mirror. Stores in this class enforce that differently, and the difference matters when you are reasoning about a specific deployment: - most refuse writes on a node while it is configured as a replica, so a caller that tries gets an error rather than a surprise; - some allow a **writable replica**, and there a write succeeds locally and is invisible to the primary and to every other copy — which is worse than an error, because nothing tells you; - whether a replica answers *reads* at all is a separate deployment choice, and plenty of deployments send every read to the primary. So: a replica is a copy, not a hot standby write address. ## Why promotion has to be a decision For the tier to keep serving writes without a human, four things must happen, in order, and each of them needs an actor that is not the failed node: 1. **Decide the primary is gone.** Something has to conclude, from silence, that the node is not coming back soon enough to wait for. 2. **Choose a candidate.** More than one replica may be eligible, and they are not equally current. 3. **Reconfigure.** The chosen node becomes primary and the surviving copies are told to follow it instead of the dead one. 4. **Get the callers there.** The callers must end up addressing the new node, whether by re-resolving a name, by asking the deciding party, or because a stable address in front of the tier was re-pointed. Who performs steps 1 to 3 is one of three arrangements: the **surviving peer nodes** themselves, a set of **dedicated observer processes** deployed alongside the tier, or a **controller outside the data plane** in a managed offering. Replication gives you none of these. A tier with three copies and nothing watching them is a tier that pages a human — which is a perfectly legitimate design, as long as the team knows that is what they bought. ## What the callers observe in the meantime - Writes fail or time out for the whole interval between the failure and the moment the new primary is both serving and being addressed. - Reads may still be answered from the copies where the deployment routes reads to replicas; where reads go to the primary, reads fail too. - Callers that are still holding the old address keep failing *after* a textbook-clean promotion, until they re-resolve. - Recent writes the old primary accepted but had not yet propagated do not survive the switch — how large that window is, and how to size it, is a question of its own. The promotion itself is usually the quickest part. The interval before anything concludes the node is gone, and the interval before the last caller stops addressing it, are both typically longer. ## What varies between stores and deployments | Question | One common arrangement | A real alternative | |---|---|---| | Can a replica accept a write? | It refuses while configured as a replica | A writable replica is allowed, and its writes stay isolated | | Who decides and promotes? | The surviving peer nodes, among themselves | Dedicated observer processes, or a controller outside the data plane | | How current must a candidate be? | The most current surviving copy is preferred, and a far-behind one may be skipped or promotion refused | Any surviving copy is promoted, and currency is not checked | | What happens when the old node returns? | It is reconfigured automatically as a replica of the new primary | It comes back believing it is primary until an operator intervenes | Because all four rows vary, the honest interview answer names the mechanism and then says which arrangement you are describing. ## How to answer this in an interview Say the three things in order: the copies exist, so this is not a data-movement problem; a replica is not a write address, so something must promote one; and the promotion is performed by a named deciding party, which is a component you have to deploy and size deliberately. Then note that the outage a caller experiences is longer than the promotion, because it also contains the time before anything noticed and the time before the caller stopped addressing the dead node.
- Does adding a third replica make the tier fail over faster?No. More copies add promotion candidates and, where reads are routed to them, read capacity — but the intervals are set by how long the deciding party waits out silence, how long reconfiguration takes, and how long callers keep the old address. What does have to be sized deliberately is the deciding party itself, not the number of copies.
- What makes one surviving replica a better promotion candidate than another?How much of the primary's work it had actually received, plus any operator-set preference — some deployments rank candidates by how current they are and will skip, or refuse to promote, a copy that is too far behind. Others promote any survivor. If your store does not check currency, the choice of candidate is effectively arbitrary, which is worth knowing before you need it.
saying these in an interview costs you the question
- Thinks replicas start taking writes the instant the primary stops answering
- Says the application can simply write to any replica instead
- Believes promotion is instant because the copy is already there
- Treats replication and automatic failover as one feature you get together
- Assumes every surviving replica holds exactly what the primary held