For a two-region Couchbase deployment, when would you choose bidirectional XDCR over one stretched cluster?
answer
- Cluster members expect a LAN between them
- Two failure domains beat one big one
- Asynchronous means a backlog, and a backlog is RPO
- Give each document one home region
- Some data should never take two writers
basics
~20 sChoose bidirectional XDCR when the regions are separated by WAN latency and must survive independently: a single Couchbase cluster assumes LAN-latency links. Accept asynchronous replication, a non-zero RPO, and per-document conflict resolution that discards a loser.
solid answer
~50 sA single Couchbase cluster is designed for LAN-latency links. Its replication, durable writes, failure detection and rebalance traffic all assume nodes are close together, so stretching one cluster across regions inflates write latency, makes auto-failover flap on transient WAN blips, and turns a partition into a cluster-wide availability problem. **Two clusters joined by XDCR** is the supported multi-region shape: each region is independently available, sized and rebalanced on its own, and failures do not cross the boundary. The price is that XDCR is asynchronous — there is a replication backlog, so the RPO on losing a region is non-zero and no durability level covers it — and that writing the same document in both regions produces conflicts resolved by discarding one whole version. So my default is bidirectional XDCR *with single-writer routing*: each tenant or key range has a home region, the other region is a warm standby, and conflict resolution is the failover-window safety net rather than a routine code path.
go deeper
Know that Couchbase does multi-region by running a separate cluster per region and connecting them with XDCR, rather than by spreading one cluster across regions.
Explain why a single cluster assumes LAN latency — heartbeats, durable-write acknowledgements and rebalance traffic all cross those links — and that XDCR is asynchronous, so the target can lag.
Show the operational picture: monitoring the replication backlog as an RPO number, rehearsing regional failover and failback, and knowing that index definitions and users do not travel over XDCR.
Own the whole trade: active-passive versus active-active, a documented single-writer routing rule, the named data classes that stay single-region, and the bucket-creation-time conflict-resolution choice you cannot revisit later.
## Frame the decision correctly This is not "which is faster". It is a question about failure domains, write latency, and who is allowed to write what. Answer it in that order. ## Why not stretch a single cluster A Couchbase cluster is a tightly coupled unit. Nodes exchange heartbeats and a shared cluster map; replicas are fed by DCP from actives; durable writes wait on cross-node acknowledgements; rebalance streams large volumes between nodes. Every one of those assumes low, stable latency between members. Put half the nodes in another region and: durable writes pay a WAN round trip on the critical path; the failure detector sees WAN jitter and either flaps toward auto-failover or must be tuned so slow that real failures go unhandled for minutes; a regional partition splits the cluster's membership rather than isolating a self-sufficient half; and rebalances push large data volumes over expensive, slow links. Couchbase's own guidance is unambiguous here: multi-region means multiple clusters connected by XDCR, not one cluster with distant nodes. ## What XDCR actually gives you Two clusters, each complete: its own vBucket map, its own replicas, its own failover behaviour, its own rebalances. XDCR streams mutations between them continuously, checkpointed so it survives pauses and restarts. A region can lose nodes, rebalance, or fall over entirely without dragging the other with it. What it does **not** give you is a cross-cluster commit. Durability levels are intra-cluster: `persistToMajority` says nothing about the other region. The exposure is the replication backlog, and it should be an explicit, monitored number — that backlog *is* your RPO when a region is lost abruptly. ## Active-passive or active-active? **Active-passive** — all writes in one region, the other read-only or dormant — has no conflicts at all. It is the right default whenever the business can accept a failover procedure instead of continuous dual writes. The design work is in the promotion runbook and the DNS/routing switch, not in the data model. **Active-active** — both regions accept writes — buys local write latency for users in both regions and instant capacity on the surviving side. It costs you conflicts, and conflicts in Couchbase are resolved by discarding a whole document version, never by merging fields. So active-active is only safe for data whose contention you have designed away. ## The rule that makes active-active work: single writer per document The practical pattern is to keep both regions active while ensuring each *document* has one home. Route a tenant, a user, or a key range to a region and write it only there. The other region serves reads and stands ready to take over. Conflicts then only occur in the seam — the failover window, or a misrouted request — where the deterministic resolution is a safety net rather than the mechanism your correctness depends on. When state is genuinely shared, restructure it: per-region documents aggregated at read time, append-style keys instead of mutating one hot document, or keeping a small strongly-contended slice single-region while everything else is local. ## What I refuse to put behind active-active Anything where losing one of two concurrent updates is a business error and the merge is not expressible as "pick a version": monetary balances and ledgers, inventory decrements against a hard limit, uniqueness claims such as username registration, and sequence or counter allocation. Those either stay single-writer or move to a design that makes the operation idempotent and commutative. Choosing timestamp (LWW) conflict resolution for such data compounds the problem, since it makes correctness depend on NTP discipline across both regions. ## Choose the conflict-resolution mode before the first document The mode is fixed at bucket creation and both ends must match. Sequence-number mode needs no clocks and is the safer default; timestamp mode gives the intuitive "latest wins" semantic but makes clock skew a silent data-loss mechanism. Changing your mind later means a new bucket and a migration, so this belongs in the design review, not in an operations ticket. ## What you must be able to operate Monitor the XDCR backlog per replication and alert on it as an RPO breach, not as a performance metric. Rehearse regional failover including the routing switch and the return of the failed region — the return is where duplicate writers and conflict storms actually appear. Remember that XDCR carries documents only: index definitions, users, and bucket settings must exist on the target before it is useful, and a standby you never queried is a standby you cannot trust. ## How I would answer in one line Two clusters with XDCR, always, for two regions; active-passive unless local write latency in both regions is a stated requirement; active-active only with a documented single-writer routing rule and a named list of data that stays single-region.
- How would you set the RPO expectation for a Couchbase deployment using XDCR?Measure it rather than assert it. The replication backlog per link is the amount of data that would be lost if the source region vanished right now, so instrument it, chart it against peak write rate, and alert when it exceeds the RPO the business signed up to. Then state the RPO as that number under stated load — and be explicit that durability levels do not shrink it, because they are intra-cluster guarantees only.
- A region fails over and later returns. What is the risk when you bring it back?Duplicate writers and a conflict burst. If routing still points some clients at the returning region while others were moved, both regions accept writes to the same documents and every one of them becomes a conflict resolved by discarding a version. Bring it back deliberately: let XDCR drain in the surviving-to-returning direction first, confirm the backlog is near zero, then move routing in one controlled step rather than letting DNS drift back.
- Which data would you keep single-writer even in an active-active Couchbase deployment?Anything where losing one of two concurrent updates is a business error that a pick-a-version rule cannot express: balances and ledgers, inventory decremented against a hard limit, uniqueness claims like username or seat reservation, and sequence allocation. Those stay in one region, or are redesigned into commutative per-region documents aggregated on read. Everything else — profiles, catalogue, session state, content — is usually fine under single-writer routing with conflict resolution as the safety net.
saying these in an interview costs you the question
- Proposes one Couchbase cluster with nodes in two regions
- Treats bidirectional XDCR as giving synchronous multi-region writes
- Assumes a durability level covers replication to the other cluster
- Puts balances or uniqueness claims behind active-active writes
- Plans no routing rule and expects conflict resolution to sort it out