skip to content

Replica Sets

Replica sets: how copies stay in sync through the oplog, how a new primary is elected, and how read preference and write concern trade consistency against latency. Interviewers ask because availability and durability guarantees are configured here, not assumed.

part ofMongoDBoverview, primer and where to startread it →
on this pageshow

questions

18

What roles do the primary, secondary, and arbiter play in a MongoDB replica set?

level: juniorimportance: must knowfreq 72%

answer

  1. One member takes writes, the rest follow
  2. The follower list is not the whole membership
  3. One member type stores nothing at all
  4. Two data copies is not three

basics

~20 s

The primary is the only member that accepts writes and records them in its oplog. Secondaries hold full copies of the data, replay that oplog, and can be elected primary. An arbiter stores no data and only votes.

solid answer

~40 s

A replica set is a group of `mongod` processes holding the same data. Exactly one member is the **primary**: it accepts every write and records it in the oplog, a capped collection called `oplog.rs` in the `local` database. **Secondaries** are full data-bearing copies that continuously tail their sync source's oplog and apply the entries, so each carries a complete copy of the data set and is eligible to become primary. By default clients read from the primary; a secondary only serves reads if the read preference asks for it. An **arbiter** (`arbiterOnly: true`) carries no data at all — it exists purely to cast a vote, as a cheap way to make the number of voters odd. MongoDB advises against arbiters, because a primary–secondary–arbiter set has only two copies of the data.

go deeper

for a junior

Be ready to name the three member types and state plainly that only the primary takes writes. Knowing that secondaries hold full copies and arbiters hold none is enough at this level.

for a middle

Explain the mechanics: writes land on the primary, get written to the oplog, and secondaries tail and apply that oplog. Be able to say how a driver discovers which member is primary.

for a senior

Show judgment on topology: argue why a third data-bearing secondary beats an arbiter, know the redundancy consequences of a P-S-A set, and explain when chained replication helps or hurts.

for a principal

Own the composition decision across regions and workloads — how many voting members, where the copies live, which members are hidden or non-voting, and what each choice costs in hardware, latency, and failure tolerance.

## What a replica set is A MongoDB replica set is a group of `mongod` processes that maintain the same data set. Membership is defined by a configuration document you can inspect with `rs.conf()` and change with `rs.reconfig()`. Drivers connect to the set as a whole: you give them a seed list, they run the `hello` command against members, discover the full topology, and keep track of which member is currently primary. That discovery is why an application does not have to be reconfigured when the primary changes. ## The primary At any moment exactly one member is the primary, and it is the only member that accepts client writes. Every write it applies is recorded as one or more entries in the **oplog** — a capped collection named `oplog.rs`, stored in the `local` database. The `local` database is itself never replicated; each member owns its own. The oplog is what makes the set a set: secondaries have no other channel through which changes reach them. The primary also serves reads by default, because the default read preference is `primary`. ## Secondaries A secondary is a full, data-bearing copy of the data set. It replicates by continuously tailing the oplog of a **sync source** and applying the entries it finds there in order. The sync source is usually the primary, but it does not have to be: a secondary may sync from another secondary — chained replication, governed by the `chainingAllowed` setting in the replica set config, which is enabled by default. Chaining reduces load on the primary but can add a hop of latency. Secondaries are hot standbys. They serve reads only when a client explicitly asks for a non-primary read preference, and any electable secondary can be elected primary if the current primary becomes unreachable. ## Arbiters An arbiter is a member configured with `arbiterOnly: true`. It runs a `mongod` process, participates in the set, and casts a vote — but it stores none of the data and applies no oplog entries. Because it stores nothing, it can run on a tiny host, which is exactly why people reach for one: it is a cheap way to give a two-data-node deployment an odd number of voters. Modern MongoDB documentation discourages arbiters, and the reason is redundancy arithmetic. A primary–secondary–arbiter (P-S-A) set contains only **two** copies of the data. Lose one data-bearing member and you are running on a single copy with no redundancy at all until it is replaced, even though the set still has a primary and looks healthy. Writes that must be acknowledged by a majority of data-bearing members also cannot be satisfied while one of the two data nodes is down. The recommended fix is almost always to add a third data-bearing secondary rather than an arbiter. ## Data-bearing members with special jobs Several member options change how a secondary participates without changing the fact that it holds all the data: - `priority: 0` — the member replicates normally but is never eligible to become primary. - `hidden: true` — the member is not advertised to client applications, so no read preference will route queries to it; it implies `priority: 0`. Typical use: a dedicated analytics or backup node. - `secondaryDelaySecs` — the member deliberately lags the primary by a fixed interval, providing a rolling window in which to recover from a destructive mistake. - `buildIndexes: false` — a rarely used option for a dedicated backup member that does not need query indexes. ## Limits A replica set can have up to 50 members, of which at most 7 may be voting members. The remainder are non-voting data-bearing members, useful for spreading read load or placing a copy in an extra region. ## Role versus configuration A common confusion is treating "primary" as something you configure. It is not: `arbiterOnly`, `hidden`, `priority` and `secondaryDelaySecs` are configuration, while primary and secondary are **states** — the current condition of a member, reported by `rs.status()`. Any electable data-bearing member can be primary at different times in the set's life, and a well-designed application never assumes which host that is.

  • Why does MongoDB recommend adding a third data-bearing secondary instead of an arbiter?
    An arbiter adds a vote but no copy of the data. A primary-secondary-arbiter set therefore keeps only two copies: if either data-bearing member is lost you are running with no redundancy, and writes requiring acknowledgement from a majority of data-bearing members cannot complete. A third secondary costs more hardware but gives real fault tolerance and lets the set survive a member loss while still majority-committing writes.
  • Can a secondary replicate from another secondary rather than from the primary?
    Yes. This is chained replication, controlled by the `chainingAllowed` setting in the replica set configuration and enabled by default. A member picks a sync source that is ahead of it and reachable, which may be another secondary. Chaining offloads oplog reads from the primary, at the cost of adding a replication hop, so latency-sensitive deployments sometimes disable it.
  • How does an application find out which member is currently primary?
    The driver maintains its own view of the topology. It runs the `hello` command against known members, learns the full member list and each member's state, and monitors them continuously. When the primary changes, the driver observes the new state and routes subsequent writes to the new primary, retrying eligible operations. Applications should connect with the full seed list rather than a single hostname.

saying these in an interview costs you the question

  • Says secondaries also accept writes directly from clients
  • Believes an arbiter stores a copy of the data
  • Thinks reads always go to secondaries by default
  • Counts a P-S-A set as three copies of the data
  • Treats primary as a fixed host you configure

context

open as a page

What does a MongoDB write concern of w: "majority" actually guarantee about a write?

level: middleimportance: must knowfreq 76%

basics

~20 s

A w: "majority" acknowledgment means the write reached a majority of the replica set's voting, data-bearing members, so it survives an election and cannot be rolled back. It does not mean every member has it.

open as a page

Why does a MongoDB replica set with four voting members tolerate no more failures than one with three?

level: middleimportance: must knowfreq 68%

basics

~20 s

Electing a primary needs a strict majority of the configured voting members. Three members need two votes, four need three, so both survive exactly one loss. The fourth vote adds cost and tie risk without adding fault tolerance.

open as a page

Why does MongoDB rewrite operations into idempotent form before writing them to the oplog?

level: middleimportance: must knowfreq 62%

basics

~20 s

So an entry can be replayed any number of times with the same result. Secondaries and recovering members may re-apply the tail of the oplog after a restart or a sync-source switch, and a relative operation replayed twice would corrupt the data.

open as a page

Which MongoDB settings must you combine so a user always reads their own write from a secondary?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Run both operations inside one causally consistent client session, write with w: "majority" and read with readConcern: "majority". The session carries a cluster timestamp so the secondary waits until it has applied that write before answering.

open as a page

After a MongoDB replica set failover, what is a rollback and which writes does it discard?

level: seniorimportance: must knowfreq 58%

basics

~20 s

A rollback happens when a former primary rejoins and holds writes the new primary never received. Those un-replicated writes are undone and saved to BSON files under the data directory. Only writes acknowledged with w majority are safe from it.

open as a page

How do MongoDB's five readPreference modes differ in which members they route reads to?

level: middleimportance: should knowfreq 66%

basics

~10 s

MongoDB offers primary (default, primary only), primaryPreferred (primary, else secondaries), secondary (secondaries only), secondaryPreferred (secondaries, else primary), and nearest (whichever eligible member has lowest latency, primary or secondary).

open as a page

If wtimeout expires on a MongoDB write with w: "majority", has the write been rolled back?

level: middleimportance: should knowfreq 52%

basics

~10 s

No. wtimeout only bounds how long the primary waits for replication acknowledgment. The write is already applied on the primary and usually replicates anyway; the error means the outcome is unknown, not undone.

open as a page

What does a replica set's electionTimeoutMillis control, and what breaks if you set it too low?

level: middleimportance: should knowfreq 52%

basics

~10 s

electionTimeoutMillis is how long an eligible secondary waits without reaching the primary before calling an election. It defaults to 10000 ms. Lowering it shortens failover detection but makes ordinary latency spikes trigger needless elections.

open as a page

How does members[n].priority affect which MongoDB replica set member becomes primary?

level: middleimportance: should knowfreq 40%

basics

~20 s

Priority expresses a preference, not a guarantee. A member with priority 0 can never be elected and cannot call an election. A member with higher priority than the current primary calls a priority takeover once its oplog is nearly caught up.

open as a page

How does initial sync differ from steady-state oplog tailing in a MongoDB replica set?

level: middleimportance: should knowfreq 52%

basics

~20 s

Initial sync copies an entire data set from a sync source to a brand-new or wiped member, building indexes and applying oplog entries produced during the copy. Steady-state tailing then applies each new oplog entry as it appears, incrementally and forever.

open as a page

What is the difference between MongoDB's readConcern levels local, majority and linearizable?

level: seniorimportance: should knowfreq 48%

basics

~20 s

local returns whatever the queried member has, including writes that may later be rolled back. majority returns only majority-committed data, which is durable but can be slightly behind. linearizable is primary-only and confirms in real time that the returned data is majority-committed.

open as a page

In a MongoDB primary-secondary-arbiter replica set, what goes wrong when the secondary goes down?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The set keeps a primary, since the primary and arbiter are two of three votes, but the arbiter holds no data. So majority write concern can no longer be satisfied, the majority commit point stops advancing, cache pressure builds, and only one copy of the data remains.

open as a page

What are hidden, delayed, and priority-0 members in a MongoDB replica set used for?

level: seniorimportance: should knowfreq 42%

basics

~20 s

They are full data-bearing secondaries carrying special restrictions. A priority-0 member replicates but can never be primary; a hidden member is additionally invisible to client read preference; a delayed member deliberately lags by a fixed interval as a rolling undo window.

open as a page

A MongoDB secondary was offline overnight and now reports RECOVERING — what happened and how do you fix it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

It fell outside the oplog window: the newest entry it applied is older than the oldest entry still retained on every potential sync source, so there is no continuous history to replay. It cannot resume tailing and must be rebuilt by initial sync or seeded from a recent copy.

open as a page

How do you decide how large a MongoDB replica set's oplog should be?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Size it by the time window it buys, not by bytes: measure oplog growth at peak write rate and allocate enough that the window comfortably exceeds your longest expected member outage, maintenance window, or initial-sync duration.

open as a page

How would you choose cluster-wide default read and write concerns for a platform with mixed workloads?

level: principalimportance: should knowfreq 30%

basics

~20 s

Set a safe cluster-wide default with setDefaultRWConcern, keeping majority writes as the floor, and let individual low-value paths opt down explicitly. Tune topology before weakening durability, and watch for arbiters silently lowering the implicit default.

open as a page

How would you configure a MongoDB replica set to meet a strict limit on write unavailability during failover?

level: principalimportance: should knowfreq 30%

basics

~20 s

Break the outage into detection, voting, catch-up and client rediscovery, then attack each. Size the voters odd and keep them on a low-latency network, avoid arbiters, use retryable writes so a failover is latency rather than an error, and rehearse with rs.stepDown() to measure the real number.

open as a page