skip to content

How would you distribute Atlas cluster nodes across regions to survive a full region outage?

level: principalimportance: should knowfreq 40%

answer

  1. Availability follows the votes
  2. Two regions is the classic trap
  3. Three vote-bearing regions survive losing one
  4. Some nodes hold data but never vote
  5. Majority acknowledgement crosses a region

basics

~20 s

Place electable nodes in at least three regions so a surviving majority can still elect a primary; two regions cannot survive losing the larger half. Use region priority to steer the primary, and add read-only or analytics nodes elsewhere without giving them votes.

solid answer

~50 s

Survival depends on where the **votes** live, not where the data lives. A 2+1 split across two regions dies when the two-node region goes: the survivor cannot form a majority and the cluster has no primary. Spreading electable members across **three** regions means any single region loss still leaves a voting majority. Atlas expresses this per region config: electable nodes with a **priority** (highest priority region holds the primary), plus optional **read-only** nodes that serve local reads but never vote or become primary, and **analytics** nodes that isolate reporting traffic and are likewise never electable. The cost is latency: majority-acknowledged writes must reach a member in a second region, so your write latency floor becomes the round trip to the nearest majority-forming node. That is the real design tension — spread the votes far enough to survive a region, but not so far that every write pays an ocean crossing. Multi-cloud adds provider-failure tolerance on the same terms.

code

json · 13 lines
json
{
  "replicationSpecs": [{
    "regionConfigs": [
      { "providerName": "AWS", "regionName": "US_EAST_1", "priority": 7,
        "electableSpecs": { "instanceSize": "M30", "nodeCount": 2 } },
      { "providerName": "AWS", "regionName": "US_WEST_2", "priority": 6,
        "electableSpecs": { "instanceSize": "M30", "nodeCount": 2 } },
      { "providerName": "AWS", "regionName": "EU_WEST_1", "priority": 5,
        "electableSpecs": { "instanceSize": "M30", "nodeCount": 1 },
        "readOnlySpecs": { "instanceSize": "M30", "nodeCount": 1 } }
    ]
  }]
}

go deeper

for a junior

Know that an Atlas cluster can place its replica set members in several regions, and that some of those members can serve reads without ever becoming the primary.

for a middle

Explain why a majority of voting members must survive, why a two-region split cannot elect after losing the larger side, and what region priority controls.

for a senior

Show that you weigh the latency cost: majority writes must cross a region, failover can move the primary away from the application tier, and read routing is a separate per-workload decision.

for a principal

Own the framing — name the failure domain you are buying protection against, the write latency budget you will accept for it, and whether multi-cloud is bought for resilience, residency or procurement.

## The rule is about votes, not copies A MongoDB replica set stays writable only while a **majority of voting members** can reach each other. Everything about multi-region design follows from that single sentence. Copies of your data in another region are worth nothing for availability if the members holding them cannot participate in electing a primary. The classic mistake is the two-region deployment. Put three electable members in two regions and you get a 2+1 split. Lose the region with two members and the remaining member is one vote out of three — no majority, no primary, no writes. You have paid for cross-region replication and bought yourself a read-only cluster at the moment you needed it most. The only fix is a **third region** holding at least one electable member, so any single region failure still leaves two of three votes alive. ## The node types Atlas gives you Atlas lets each region in a cluster contribute different kinds of members: - **Electable nodes** vote in elections and can become primary. Each region config carries a **priority**; the region with the highest priority is where Atlas steers the primary, and lower-priority regions take over only when it is unavailable. - **Read-only nodes** hold a full copy and serve reads, but they have no vote and can never be elected. They are how you put a local read replica in a distant region without dragging that region into your election quorum or your majority write path. - **Analytics nodes** are also non-electable read replicas, intended to isolate long-running reporting and BI queries from operational traffic. They carry a replica-set tag (`nodeType: ANALYTICS`) so a client can target them explicitly with a tagged read preference. Remember the replica-set ceiling: a set may have at most seven voting members, so "add a vote everywhere" is not an available strategy. Extra copies beyond that must be non-voting. ## What availability actually costs Spreading votes across regions buys survival and charges latency. Once a majority cannot be formed inside one region, a majority-acknowledged write must travel to another region and back before it is confirmed. Your write latency floor becomes the round trip to the nearest member that completes the majority. Choosing regions is therefore partly a geography exercise: three regions on one continent behave very differently from three spread across continents, even though both survive one region loss. This produces a real design tension that a principal-level candidate should name explicitly: - **Tight grouping** (three nearby regions) keeps writes fast and survives a regional failure, but not a wide correlated failure. - **Wide grouping** survives more, but every durable write pays a long round trip, and after a failover the primary may land far from your application tier — which changes application latency, not just database latency. - **Asymmetric designs** — two electable regions plus a small third region holding one electable member purely as a tiebreaker — are a common compromise: the tiebreaker region carries little traffic but restores the ability to form a majority. ## Multi-cloud Atlas can place members of a single replica set across AWS, Google Cloud and Azure. The reasoning is identical — it is still about where the votes are — but the failure domain being defended is a **provider**, not a region. Realistically it also serves procurement and data-residency goals, and it lets you put a read-only node close to workloads that live on another cloud. The costs are the ones you would expect: cross-provider network paths are typically slower and less predictable than intra-provider ones, egress is billed, and your operational tooling now spans three consoles. ## Where the primary should live Priority ordering is the lever that decides where writes are served from in steady state, and it deserves an explicit decision rather than a default. Put the highest priority in the region hosting your write-heavy application tier. If a failover moves the primary elsewhere, be clear about whether you want it to stay there (fewer disruptions) or migrate back (predictable latency) — the priorities encode that choice, and every failback is another election. ## Reads are a separate decision Multi-region node placement gives you the *option* of local reads, not local reads by default. Directing traffic to nearby secondaries or read-only nodes is a client-side read routing decision with its own staleness consequences, and it should be made per workload: a dashboard can tolerate lag that a checkout flow cannot. ## The judgment an interviewer wants Start from the failure you are actually defending against, not from a diagram. "Which single failure must we survive without losing writes, and what write latency will we accept to get it?" answers the topology question almost by itself: it fixes the number of vote-bearing regions, their geographic spread, and where the priority sits. Everything else — read-only nodes, analytics nodes, multi-cloud — is optimisation layered on top.

  • Why does a 2+1 two-region Atlas cluster fail to elect a primary when the larger region is lost?
    Because a replica set needs a majority of voting members to elect. With three voting members split two and one, losing the two-node region leaves a single vote out of three. The survivor has all your data and can serve reads, but it cannot become primary, so writes stop until the failed region returns.
  • When would you add read-only nodes instead of more electable nodes?
    When you want a local copy for reads in a region that should not influence elections or write latency. Read-only nodes never vote and never become primary, so they add read capacity and geographic proximity without enlarging the quorum or pulling a distant region into the majority write path.
  • What is the point of Atlas analytics nodes?
    Isolation. They are non-electable replicas dedicated to long-running reporting and BI queries, tagged so clients can target them deliberately with a tagged read preference. Heavy analytical scans then compete for cache and I/O on their own hardware instead of degrading the operational workload on the primary and its secondaries.
  • How do you decide which region holds the primary?
    Set the highest region priority where your write-heavy application tier runs, so the steady-state write path is local. Then decide explicitly whether the primary should fail back after an outage: failing back restores predictable latency but costs another election, while staying put avoids the disruption at the price of a longer write path.

saying these in an interview costs you the question

  • Deploying across two regions and calling it region-fault-tolerant
  • Counting data copies instead of voting members
  • Ignoring that cross-region majority writes add latency
  • Assuming multi-region automatically gives local reads
  • Adding votes everywhere despite the seven-voting-member ceiling

context