What roles do the primary, secondary, and arbiter play in a MongoDB replica set?
answer
- One member takes writes, the rest follow
- The follower list is not the whole membership
- One member type stores nothing at all
- Two data copies is not three
basics
~20 sThe primary is the only member that accepts writes and records them in its oplog. Secondaries hold full copies of the data, replay that oplog, and can be elected primary. An arbiter stores no data and only votes.
solid answer
~40 sA replica set is a group of `mongod` processes holding the same data. Exactly one member is the **primary**: it accepts every write and records it in the oplog, a capped collection called `oplog.rs` in the `local` database. **Secondaries** are full data-bearing copies that continuously tail their sync source's oplog and apply the entries, so each carries a complete copy of the data set and is eligible to become primary. By default clients read from the primary; a secondary only serves reads if the read preference asks for it. An **arbiter** (`arbiterOnly: true`) carries no data at all — it exists purely to cast a vote, as a cheap way to make the number of voters odd. MongoDB advises against arbiters, because a primary–secondary–arbiter set has only two copies of the data.
go deeper
Be ready to name the three member types and state plainly that only the primary takes writes. Knowing that secondaries hold full copies and arbiters hold none is enough at this level.
Explain the mechanics: writes land on the primary, get written to the oplog, and secondaries tail and apply that oplog. Be able to say how a driver discovers which member is primary.
Show judgment on topology: argue why a third data-bearing secondary beats an arbiter, know the redundancy consequences of a P-S-A set, and explain when chained replication helps or hurts.
Own the composition decision across regions and workloads — how many voting members, where the copies live, which members are hidden or non-voting, and what each choice costs in hardware, latency, and failure tolerance.
## What a replica set is A MongoDB replica set is a group of `mongod` processes that maintain the same data set. Membership is defined by a configuration document you can inspect with `rs.conf()` and change with `rs.reconfig()`. Drivers connect to the set as a whole: you give them a seed list, they run the `hello` command against members, discover the full topology, and keep track of which member is currently primary. That discovery is why an application does not have to be reconfigured when the primary changes. ## The primary At any moment exactly one member is the primary, and it is the only member that accepts client writes. Every write it applies is recorded as one or more entries in the **oplog** — a capped collection named `oplog.rs`, stored in the `local` database. The `local` database is itself never replicated; each member owns its own. The oplog is what makes the set a set: secondaries have no other channel through which changes reach them. The primary also serves reads by default, because the default read preference is `primary`. ## Secondaries A secondary is a full, data-bearing copy of the data set. It replicates by continuously tailing the oplog of a **sync source** and applying the entries it finds there in order. The sync source is usually the primary, but it does not have to be: a secondary may sync from another secondary — chained replication, governed by the `chainingAllowed` setting in the replica set config, which is enabled by default. Chaining reduces load on the primary but can add a hop of latency. Secondaries are hot standbys. They serve reads only when a client explicitly asks for a non-primary read preference, and any electable secondary can be elected primary if the current primary becomes unreachable. ## Arbiters An arbiter is a member configured with `arbiterOnly: true`. It runs a `mongod` process, participates in the set, and casts a vote — but it stores none of the data and applies no oplog entries. Because it stores nothing, it can run on a tiny host, which is exactly why people reach for one: it is a cheap way to give a two-data-node deployment an odd number of voters. Modern MongoDB documentation discourages arbiters, and the reason is redundancy arithmetic. A primary–secondary–arbiter (P-S-A) set contains only **two** copies of the data. Lose one data-bearing member and you are running on a single copy with no redundancy at all until it is replaced, even though the set still has a primary and looks healthy. Writes that must be acknowledged by a majority of data-bearing members also cannot be satisfied while one of the two data nodes is down. The recommended fix is almost always to add a third data-bearing secondary rather than an arbiter. ## Data-bearing members with special jobs Several member options change how a secondary participates without changing the fact that it holds all the data: - `priority: 0` — the member replicates normally but is never eligible to become primary. - `hidden: true` — the member is not advertised to client applications, so no read preference will route queries to it; it implies `priority: 0`. Typical use: a dedicated analytics or backup node. - `secondaryDelaySecs` — the member deliberately lags the primary by a fixed interval, providing a rolling window in which to recover from a destructive mistake. - `buildIndexes: false` — a rarely used option for a dedicated backup member that does not need query indexes. ## Limits A replica set can have up to 50 members, of which at most 7 may be voting members. The remainder are non-voting data-bearing members, useful for spreading read load or placing a copy in an extra region. ## Role versus configuration A common confusion is treating "primary" as something you configure. It is not: `arbiterOnly`, `hidden`, `priority` and `secondaryDelaySecs` are configuration, while primary and secondary are **states** — the current condition of a member, reported by `rs.status()`. Any electable data-bearing member can be primary at different times in the set's life, and a well-designed application never assumes which host that is.
- Why does MongoDB recommend adding a third data-bearing secondary instead of an arbiter?An arbiter adds a vote but no copy of the data. A primary-secondary-arbiter set therefore keeps only two copies: if either data-bearing member is lost you are running with no redundancy, and writes requiring acknowledgement from a majority of data-bearing members cannot complete. A third secondary costs more hardware but gives real fault tolerance and lets the set survive a member loss while still majority-committing writes.
- Can a secondary replicate from another secondary rather than from the primary?Yes. This is chained replication, controlled by the `chainingAllowed` setting in the replica set configuration and enabled by default. A member picks a sync source that is ahead of it and reachable, which may be another secondary. Chaining offloads oplog reads from the primary, at the cost of adding a replication hop, so latency-sensitive deployments sometimes disable it.
- How does an application find out which member is currently primary?The driver maintains its own view of the topology. It runs the `hello` command against known members, learns the full member list and each member's state, and monitors them continuously. When the primary changes, the driver observes the new state and routes subsequent writes to the new primary, retrying eligible operations. Applications should connect with the full seed list rather than a single hostname.
saying these in an interview costs you the question
- Says secondaries also accept writes directly from clients
- Believes an arbiter stores a copy of the data
- Thinks reads always go to secondaries by default
- Counts a P-S-A set as three copies of the data
- Treats primary as a fixed host you configure