When a server fails, how does availability differ between a leaderless wide-column store and one that serves each key range from a single server?
answer
- no failover vs reassignment
- enough replicas still answer
- detect, reassign, replay the log
- shared storage shortens recovery
- stale answers vs brief gaps
basics
~20 sA leaderless store keeps serving from the surviving replicas, as long as enough of them answer. A range-served store makes the failed server's key ranges unavailable until they are reassigned and any unflushed writes are replayed from the log.
solid answer
~50 sIn a **leaderless** store there is no failover: every replica of a partition is equal, so when one node dies the others keep taking reads and writes, and a request succeeds as long as enough replicas answer for what it asked. The costs are **staleness** on the recovered node and background work to bring it up to date. In a **range-served** store each range has one server, so its failure makes that server's ranges **briefly unavailable**. A master detects the loss (usually when the server's lease or session expires), reassigns the ranges to healthy servers, and the new owners **replay the commit log** for writes that had not been flushed. Recovery time is detection timeout plus reassignment plus log replay. Stores that keep all data on shared storage move only metadata and recover quickly; some also offer read-only secondary copies to answer reads, possibly stale, meanwhile.
go deeper
Know that leaderless stores keep serving from other replicas while range-served stores reassign the failed server's ranges.
Walk through the range-served recovery steps, detection, fencing, reassignment and log replay, and what a leaderless request needs to succeed.
Estimate recovery time from its factors, set client retry behaviour, and match the model to the application's tolerance for pauses versus stale reads.
Be ready to argue availability targets for a system in terms of these two failure behaviours and the operational investment each requires.
## The question behind the question "What happens when a node dies?" tests whether a candidate knows **which replication model** a store uses, because the two models fail in opposite ways: one keeps answering with possibly stale data, the other briefly stops answering for some keys and then answers correctly. ## Leaderless: no failover to perform Every partition lives on several equal replicas. When a node fails: - **Nothing needs to be promoted.** Coordinators simply stop sending to the dead node. - A read or write succeeds if **enough replicas respond** for the level the request asked for. With three replicas and one down, requests that need one or two answers still succeed; a request that needs all three fails. - Writes the dead node missed are held or repaired in the background, and reads can repair stale copies, so the node catches up after it returns. - The visible effect is **latency and staleness risk**, not an outage for the partition. Longer outages matter differently: a node away for too long needs a full repair before rejoining, and deletes it missed must not be allowed to come back — the deletes-versus-repair subject. ## Range-served: detect, reassign, replay Each key range has **exactly one serving server**. When it fails: 1. **Detection.** A coordination service notices the server's lease or session expired. Timeouts are tuned to avoid false positives from pauses, so detection itself takes seconds. 2. **Fencing.** The master ensures the old server can no longer serve, so two servers never own a range at once. 3. **Reassignment.** The failed server's ranges are assigned to healthy servers. 4. **Log replay.** Writes acknowledged but not yet flushed to data files exist only in the commit log; the new owners replay the relevant log entries before serving writes, and in some designs before serving at all. Until step 4 completes, the affected ranges **cannot be read or written**. Everything else keeps working. ## What makes recovery fast or slow | factor | effect | |---|---| | failure-detection timeout | lower means faster recovery but more false failovers | | amount of unflushed data in the log | more to split and replay | | number of ranges on the failed server | more reassignments | | data on shared storage vs local disks | shared storage moves only metadata | | optional secondary read copies | reads continue, possibly stale, during recovery | ## Choosing with failure in mind - If a few seconds of unavailability for a slice of keys is worse than occasionally stale reads, the leaderless model degrades more gracefully. - If stale or conflicting data is worse than a brief pause, the range-owner model's single-writer semantics are easier to reason about. - Either way, **clients must retry**: leaderless clients retry against other replicas, range-served clients retry until the range's new location is published. ## Interview angle Strong answers describe the failure timeline for the range-owner model step by step, explain why the leaderless model has no failover at all, and connect each to the application's tolerance for pauses versus staleness.
- Why not make failure detection in a range-served store near-instant?A short timeout mistakes garbage-collection pauses or network blips for failures, causing needless reassignments and log replays that hurt more than the pause. Teams pick a timeout that balances recovery time against false failovers.
- What does a leaderless client see while a replica is down?Requests needing fewer answers than the surviving replicas succeed, perhaps with more latency; requests needing every replica fail. Reads may be served from replicas that missed recent writes if the chosen level allows it.
saying these in an interview costs you the question
- Claiming a leaderless store fails over to a new leader for each partition
- Assuming a range-served store's ranges stay writable during reassignment
- Ignoring log replay of unflushed writes as part of recovery time
- Tuning failure detection as short as possible without considering false failovers