skip to content

Clustering, Durability & XDCR

How a cluster spreads data over vBuckets, what durability you can demand per write, and how two datacenters stay in sync. XDCR is the headline feature in multi-region discussions.

on this pageshow

questions

6

What do Couchbase durability levels majority, majorityAndPersistActive and persistToMajority guarantee?

level: middleimportance: must knowfreq 65%

answer

  1. Default write is memory on one node
  2. Ask for more per operation
  3. Memory versus disk, active versus majority
  4. Majority means replicas + 1, halved
  5. One replica plus one dead node equals impossible

basics

~20 s

In Couchbase, majority means the mutation is in memory on a majority of the active-plus-replica copies; majorityAndPersistActive adds a disk write on the active; persistToMajority requires it on disk on a majority. Each step up costs latency.

solid answer

~50 s

Couchbase writes are memory-first, so an ordinary `upsert` returns as soon as the **active** node has it in memory — a crash of that node before replication can lose it. Durability levels (synchronous replication, Couchbase 6.5+) let you demand more per operation. **`majority`** holds the write as pending until a majority of the copies — active plus replicas — have it in memory. **`majorityAndPersistActive`** adds the requirement that the active node has written it to disk. **`persistToMajority`** requires it on disk on a majority of the copies, which is the only level that survives simultaneous power loss across the majority. The write is a *sync write*: the active holds it pending, the client sees success only after the requirement is met, and a concurrent write to the same key is rejected while one is pending. Do the majority arithmetic before promising anything: with one replica the majority of two copies is two, so any single node failure makes durable writes impossible.

code

java · 4 lines
java
collection.upsert("order::1001", order,
    UpsertOptions.upsertOptions()
        .durability(DurabilityLevel.PERSIST_TO_MAJORITY)
        .timeout(Duration.ofSeconds(10)));

go deeper

for a junior

Recall that a normal Couchbase write is acknowledged from memory on one node, and that you can ask for a durability level on individual writes to get a stronger guarantee.

for a middle

Distinguish the three levels precisely — memory on a majority, plus disk on the active, versus disk on a majority — and be able to compute the majority as replicas plus one.

for a senior

Show the operational judgment: at least two replicas if durable writes must survive a node failure, ambiguous timeouts require idempotent writes, and durable writes on a hot key serialise because a pending sync write blocks the next one.

for a principal

Own where the line is drawn — which classes of data pay the extra round trips, what RPO that buys inside the cluster, and the explicit statement that no durability level says anything about the XDCR target.

## Why the default is not durable Couchbase's data service is memory-first. A plain `upsert` is acknowledged when the active node holds the mutation in its managed cache; replication to replicas and persistence to disk both happen asynchronously afterwards. This is what gives Couchbase its latency profile, and it is also a real durability hole: if the active node's process or machine dies in that window, the acknowledged mutation is gone. Durability levels, introduced with synchronous replication in Couchbase 6.5 and standard in 7.x, let a caller pay for a stronger guarantee on the operations that need it, per operation, without changing the bucket's behaviour for everything else. ## The three levels **`majority`** — the mutation must be in memory on a majority of the copies of its vBucket, counting the active and its replicas. It survives the loss of the active node, because at least one surviving replica already holds it and will be promoted. **`majorityAndPersistActive`** — majority in memory, *and* written to disk on the active node. It additionally survives a scenario where the active node's process restarts and its cache is empty, since the value is on the active's disk. **`persistToMajority`** — the mutation must be on disk on a majority of the copies. This is the level that survives a correlated power loss of the majority of nodes holding the vBucket, and it is the slowest because it waits on disk I/O on multiple machines. ## How a sync write actually executes A durable write is a **sync write**. The active node accepts the mutation and holds it in a *pending* state: the value is not yet visible as committed, and a second write to the same key while one is pending is refused rather than queued behind it silently. The active replicates the mutation over DCP and waits for acknowledgements matching the requested level — in-memory acks for `majority`, disk acks where the level demands persistence. When the requirement is met, the active commits and only then returns success to the client, and the commit is propagated to the replicas. ## Majority arithmetic — the part people get wrong "Majority" is a majority of the copies of that vBucket, which is `replicas + 1`: - `num_replicas = 1` → 2 copies → majority is **2** → both the active and its single replica must ack. One node down means durable writes to the affected vBuckets are impossible. - `num_replicas = 2` → 3 copies → majority is **2** → the active plus either replica. Tolerates one node down. - `num_replicas = 3` → 4 copies → majority is **3**. The consequence is blunt: **if you want durable writes to keep working through a single node failure, you need at least two replicas.** A common production surprise is a bucket with one replica where every durable write starts failing the moment a node goes down. ## What the client sees when it cannot be satisfied The SDKs distinguish two failure shapes, and the distinction matters for retry logic. When the cluster *cannot* meet the level — not enough replicas are online for a majority — the operation is rejected up front with a durability-impossible error. Nothing was written; retrying immediately will fail the same way, so the correct response is to shed, queue, or degrade, not to hammer. When the operation times out or the connection breaks while the sync write is pending, the SDK reports a durability-*ambiguous* condition. The write may or may not commit afterwards. This is the honest answer, and applications that care must make their writes idempotent (write a full document keyed deterministically, use CAS, or check-and-retry) rather than assuming a timeout means failure. ## Costs and mixing Every level above the default adds at least one extra network round trip inside the cluster, and the persist levels add disk-flush latency on top. Throughput for durable writes on the same key is also bounded by the pending-write rule: you cannot pipeline many mutations of one hot key. Durability is per operation, so the usual design is to demand it on the writes that represent money or state you cannot reconstruct, and leave counters, caches, and derived data on the default. One more scoping point: these guarantees are **intra-cluster**. XDCR to another cluster is asynchronous and is not covered by any durability level; a `persistToMajority` write can still be absent from the remote cluster. ## Legacy PersistTo / ReplicateTo Older Couchbase SDKs offered `PersistTo` and `ReplicateTo` observe-based options, where the client polled until enough nodes reported the item. That is client-driven and weaker — it does not make the write atomic or hold it pending server-side. On 6.5+ prefer durability levels; recognising the old options and why they were replaced is a good senior signal.

  • A bucket has one replica and one node goes down. What happens to durable writes?
    They fail. With one replica there are two copies of each vBucket, so a majority is two — the active and its replica must both acknowledge. Losing either node makes the requirement unsatisfiable for the affected vBuckets and the SDK reports a durability-impossible error immediately rather than blocking. Surviving a single node failure with durable writes intact requires at least two replicas.
  • How should an application handle a durability-ambiguous timeout?
    Treat it as "unknown", never as "failed". The sync write may still commit after the client gave up. Make the write idempotent — write a whole document under a deterministic key, or use CAS so a retry cannot double-apply — and then retry, or read back and reconcile. Blindly retrying a non-idempotent mutation such as an in-place increment is how duplicates get created.
  • Does a persistToMajority write guarantee the data reached the XDCR target cluster?
    No. Durability levels are entirely intra-cluster: they describe how many copies within the source cluster hold the mutation in memory or on disk. XDCR is asynchronous, source-driven replication with its own queue, so a fully durable local write can still be missing from the remote cluster at the moment the source region is lost. Cross-cluster exposure is measured as replication lag and RPO, not as a durability level.

saying these in an interview costs you the question

  • Assumes every Couchbase write is already replicated before acknowledgement
  • Reads majority as a majority of all cluster nodes
  • Promises durable writes through a node failure with one replica
  • Treats a durability timeout as a definite failure and retries blindly
  • Thinks a durability level also covers XDCR to another cluster

context

open as a page

How does Couchbase map a document key to a node using vBuckets?

level: middleimportance: must knowfreq 70%

basics

~20 s

Couchbase hashes the document key into one of a bucket's 1024 vBuckets, then looks that vBucket up in the cluster map to find the node holding its active copy. The SDK does this locally, so requests go straight to the owning node.

open as a page

What happens to vBuckets and connected clients during a Couchbase rebalance?

level: middleimportance: should knowfreq 55%

basics

~20 s

A Couchbase rebalance streams vBuckets to their new nodes while the cluster stays online, flipping ownership only after each transfer completes. Clients hitting the old owner get a "not my vBucket" reply and retry with the refreshed map.

open as a page

How does Couchbase XDCR resolve conflicts between two clusters that both accepted writes?

level: seniorimportance: should knowfreq 60%

basics

~10 s

XDCR resolves conflicts deterministically per document using the bucket's conflict-resolution mode, fixed at bucket creation: sequence-number (revision metadata) or timestamp last-write-wins. Both clusters pick the same winner and converge; the losing version is discarded.

open as a page

When Couchbase auto-fails-over a data node, what happens to its vBuckets and what must you do next?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Failover promotes the failed node's replica vBuckets to active on surviving nodes, restoring access immediately but leaving the cluster with fewer replicas than configured. A rebalance afterwards recreates the missing replicas and rebuilds the map.

open as a page

For a two-region Couchbase deployment, when would you choose bidirectional XDCR over one stretched cluster?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Choose bidirectional XDCR when the regions are separated by WAN latency and must survive independently: a single Couchbase cluster assumes LAN-latency links. Accept asynchronous replication, a non-zero RPO, and per-document conflict resolution that discards a loser.

open as a page