When Couchbase auto-fails-over a data node, what happens to its vBuckets and what must you do next?
answer
- Replicas already exist elsewhere
- The map changes, the data does not move
- The unreplicated tail is gone
- Redundancy is not restored automatically
- A timeout and an event cap guard it
basics
~20 sFailover promotes the failed node's replica vBuckets to active on surviving nodes, restoring access immediately but leaving the cluster with fewer replicas than configured. A rebalance afterwards recreates the missing replicas and rebuilds the map.
solid answer
~50 sWhen a data node is failed over, Couchbase removes it from the cluster map and promotes the replica copies of its vBuckets — which live on other nodes — to active. Clients pick up the new map (or get a "not my vBucket" retry) and service resumes in seconds. Two things are then true and both matter. First, any mutation that was in the failed node's memory but not yet replicated is **lost**, unless it was written with a durability level that required replica acknowledgement. Second, the cluster is now **under-protected**: the promoted vBuckets have one fewer replica than the bucket asks for, and only a **rebalance** recreates them. Automatic failover is guarded — a configurable timeout before it fires, and a cap on how many automatic events it will perform before requiring operator intervention — precisely because failing over aggressively in a flaky network is how you turn a blip into data loss.
code
bash · 6 lines# Configure automatic failover, then re-add a recovered node
couchbase-cli setting-autofailover -c host:8091 -u Administrator -p "$PASS" \
--enable-auto-failover 1 --auto-failover-timeout 120
couchbase-cli recovery -c host:8091 -u Administrator -p "$PASS" \
--server-recovery returned-node:8091 --recovery-type deltago deeper
Know that Couchbase keeps replica copies of every vBucket, and that failing a node over promotes those replicas so the data stays reachable.
Explain that failover only rewrites the cluster map — it moves no data — so it is fast but loses whatever was in the failed node's memory and not yet replicated.
Show the post-incident sequence you would actually run: confirm the failover, recognise the cluster is under-replicated, rebalance to restore replicas, and choose delta versus full recovery for the returning node.
Own the policy: auto-failover timeout and event budget, minimum node counts and server-group layout, replica count driven by the durable-write requirement, and the spare capacity needed for the recovery rebalance to succeed.
## Failover versus rebalance These are different operations for different situations and conflating them is the classic Couchbase interview error. **Rebalance** is for planned change on a healthy cluster: it moves data before changing ownership, and loses nothing. **Failover** is for a node that is already gone or is being removed abruptly. It changes the cluster map immediately, without moving data first, by promoting replicas that already exist elsewhere. It is fast by design and it can lose the tail of writes that had not yet replicated. ## The three flavours **Hard failover** takes the node out immediately. Use it when the node is dead or unreachable. Data that lived only in that node's memory is lost. **Graceful failover** is for a node that is still healthy and responsive — you are taking it out for maintenance. It drains the active vBuckets to their replicas first, so no data is lost, then promotes and removes the node. It takes longer than hard failover because it waits for that sync. **Automatic failover** is the cluster performing a hard failover on its own when a node stops responding for longer than the configured auto-failover timeout (default 120 seconds; the minimum configurable value is much lower). It is what keeps a cluster available at 3 a.m. without a human. ## What happens at the moment of failover The orchestrator rewrites the cluster map: for every vBucket whose active copy lived on the failed node, a replica copy on a surviving node becomes active. Clients receive the new configuration by push, and any request still aimed at the old owner gets "not my vBucket" with the current map attached and retries. From the application's perspective there is a short window of errors and elevated latency, then normal service. What is lost is exactly what the failed node had that nobody else had: mutations acknowledged from its memory but not yet replicated. This is where durability levels earn their cost — a write acknowledged under `majority` was on a replica before the client saw success, so it survives the promotion. This is also why running with `num_replicas = 0` and expecting failover to be safe is nonsense: there is nothing to promote, and the vBuckets are simply gone. ## After failover: the cluster is not healthy yet This is the part candidates skip. Post-failover the promoted vBuckets have one fewer copy than the bucket's configured replica count. In a one-replica bucket they now have **no** replica at all — a second node failure loses data outright, and durable writes to those vBuckets become impossible. Nothing recreates replicas automatically. You must run a **rebalance**, which streams fresh replica copies onto the remaining nodes (or onto a replacement node you add in the same rebalance). If the failed node comes back and its data files are intact, you can add it back with a **recovery type**: *delta recovery* reuses the node's on-disk data and only catches it up on what it missed, which is far cheaper; *full recovery* discards its files and re-streams everything. Delta recovery is only viable when the node's storage genuinely survived and the divergence is small. ## Why auto-failover has guardrails Automatic failover is deliberately conservative because a failover is unrecoverable in the direction of data: - **A timeout** (default 120 seconds) before it fires, so a GC pause, a brief network partition, or a node restart does not trigger it. - **A cap on automatic events** before an operator must intervene. This prevents a cascade where a flapping network fails over node after node until the cluster cannot serve at all. The budget is restored by a successful rebalance. - **Safety refusals.** The cluster declines to auto-fail-over when doing so would clearly cause unavailability or guaranteed data loss — most notably in very small data-service deployments, where auto-failover requires at least three data nodes so a partition cannot leave two halves each failing the other over. ## Server groups When server-group (rack/zone) awareness is configured, replicas are placed in a different group from their actives, so an entire group can be lost and every vBucket still has a copy. Recovering from a whole-group failure is still a failover of all its nodes followed by a rebalance, and you need the spare capacity in the surviving groups for that rebalance to succeed — capacity planning, not a runtime feature. ## Interview traps Saying "failover restores full redundancy" is wrong. Saying "use failover to remove a healthy node" is wrong — that is graceful failover at best, a rebalance-out at correct. Claiming auto-failover fires instantly ignores the timeout that exists to prevent flapping. And attributing zero data loss to hard failover ignores the unreplicated memory tail unless durability was requested.
- Why does Couchbase cap the number of automatic failovers before requiring an operator?Because each auto-failover permanently removes a node and can lose its unreplicated writes, and a flapping network can produce a cascade — node after node judged dead until the cluster has too few members to serve. The cap turns an unbounded automatic cascade into one or a small number of events plus an alert. The budget is restored after a successful rebalance, which is the operator's confirmation that the cluster is healthy again.
- When is delta recovery preferable to full recovery for re-adding a failed-over node?When the node's data files survived intact and it was out for a short time. Delta recovery reuses the on-disk vBucket data and streams only the mutations it missed, so the rebalance moves a fraction of the data and finishes far sooner. Choose full recovery when the storage was replaced or is suspect, or when the node was out long enough that the delta approaches a full copy anyway.
- Why does Couchbase require at least three data nodes for automatic failover?With only two data nodes, a network partition between them leaves each side seeing the other as dead. Allowing automatic failover there would let both sides fail the other over and diverge — a split brain. Requiring three means a surviving majority can be identified before any node is removed automatically. In smaller clusters the failover decision is left to an operator who can see the whole picture.
saying these in an interview costs you the question
- Claims failover restores the configured replica count
- Uses hard failover to remove a healthy node for maintenance
- Says failover never loses data regardless of durability settings
- Thinks auto-failover fires the instant a node stops responding
- Runs zero replicas and still expects failover to protect data