After a MongoDB replica set failover, what is a rollback and which writes does it discard?
answer
- Two oplogs that diverge after a shared point
- The elected primary's history wins
- Depends entirely on the write concern used
- Undone documents land in files on disk
- Journalling on one node is not enough
basics
~20 sA rollback happens when a former primary rejoins and holds writes the new primary never received. Those un-replicated writes are undone and saved to BSON files under the data directory. Only writes acknowledged with w majority are safe from it.
solid answer
~50 sWhen a primary accepts writes and then loses contact before those writes reach the rest of the set, a new primary is elected without them. When the old primary comes back it discovers its oplog diverges from the new primary's history after a common point. It cannot force its version on the set, so it undoes every operation after that divergence point — that is a rollback — and rejoins as a secondary. The undone documents are written to BSON files in a `rollback` directory under `dbPath` so an operator can inspect or reapply them; nothing is replayed automatically, and clients that already received an acknowledgement are never told. The writes at risk are exactly those acknowledged before reaching a majority: `w: 1` writes, and `j: true` writes that were only durable on the one node. A write acknowledged with `w: "majority"` is majority-committed and can never be rolled back.
code
javascript · 5 lines// Exposed to rollback: acknowledged by the primary alone
db.orders.insertOne({ _id: 1 }, { writeConcern: { w: 1 } })
// Safe from rollback: majority-committed before acknowledgement
db.orders.insertOne({ _id: 2 }, { writeConcern: { w: "majority" } })go deeper
Recall that a MongoDB failover can discard writes the old primary accepted but never replicated, and that the write concern used decides whether a write is exposed.
Explain the mechanism: diverging oplogs, the majority commit point as the exact boundary, the elected primary's history winning, and the rejoining node reverting to the common point.
Show that you have handled one — find the rollback files under dbPath, reconcile them with the application owner, and explain why journalling on a single node is not cluster durability.
Own the policy: which write paths get majority acknowledgement and which knowingly accept loss, what the latency cost is, and how you detect and account for lost acknowledged writes after an incident.
## The situation that produces a rollback A replica set has one primary. It applies a write, records it in the oplog, and — depending on the write concern — may acknowledge the client immediately, before any secondary has copied that oplog entry. Now the primary is partitioned away or crashes. The surviving members hold an election, and one of them becomes primary. That new primary's oplog does not contain the entries that never made it across. The set now has two histories that agree up to some point and diverge after it. MongoDB resolves this asymmetrically and deliberately: the history held by the elected primary wins. When the old primary returns, it compares oplogs, finds the last entry the two share, and reverts everything it applied after that. It transitions through a `ROLLBACK` state, then `RECOVERING`, then rejoins as a `SECONDARY` following the new primary. This is not a bug or a corruption event. It is how MongoDB converges to a single history without ever allowing two divergent primaries to persist. ## Which writes are actually at risk The precise boundary is the **majority commit point**: the position in the oplog up to which a majority of data-bearing voting members are known to have replicated. Anything at or before that point is majority-committed and durable across a failover. Anything after it, on a node that then loses the election, is a rollback candidate. That maps onto write concerns directly: - `w: 1` — acknowledged by the primary alone. **Rollback candidate.** - `j: true` with `w: 1` — flushed to that one node's journal. Durable against that node crashing and restarting, but *not* against that node losing the election. **Still a rollback candidate.** This is the most common misconception in the room: journalling is per-node durability, not cluster durability. - `w: "majority"` — the primary waits until a majority of data-bearing voting members have applied the write before acknowledging. **Never rolled back**, because any node that later wins an election must have that write, since two majorities overlap. Since MongoDB 5.0 the implicit default write concern for a replica set is `w: "majority"`, so a modern deployment is protected by default unless the application or a set with arbiters overrides it. Applications that deliberately downgrade to `w: 1` for latency are accepting the possibility of losing acknowledged writes. ## What actually happens to the data The reverted documents are not simply thrown away. MongoDB writes rollback files — BSON files, one set per affected collection — into a `rollback` directory under the node's `dbPath`. They contain the documents as they existed on the rolled-back node. Since MongoDB 4.0 these files are always created, so a rollback always leaves an audit trail on disk. Recovering that data is a manual operation. You read the BSON files, decide which of them represent business-meaningful writes that were genuinely lost, and reapply them — carefully, because the rest of the world has moved on and blindly reinserting can violate invariants that later writes established. There is no automatic replay, and there should not be: only the application knows whether reapplying a stale order or a stale balance adjustment is correct. Modern MongoDB (4.0 and later) implements rollback by recovering the storage engine to a stable timestamp rather than by walking the oplog backwards operation by operation, which makes it much faster and removes the old fixed data-size cap on how much could be rolled back. A configurable time limit bounds how far back a node is willing to roll; beyond it the node requires a full resync instead. ## The client's perspective — why this stings A client that issued a `w: 1` write received a successful acknowledgement. Nothing ever tells it the write was later reverted. The write simply is not there any more. If that acknowledgement was used to update a UI, send a confirmation email, or trigger a downstream call, the system is now inconsistent with its own database, and no amount of retry logic will notice. This is the practical argument for majority write concern on anything that matters. The latency cost is one round trip to the second-fastest replica; the benefit is that an acknowledgement means what the caller thinks it means. ## Reducing exposure - Use `w: "majority"` for writes whose loss would be visible or damaging, and treat `w: 1` as a deliberate, documented choice for low-value high-volume data. - Keep the voting members close enough that the majority commit point advances quickly; a lagging secondary widens the window of un-committed writes on the primary. - Avoid configurations that make majority acknowledgement fragile — most notably a primary-secondary-arbiter set, where the loss of the single secondary makes `w: "majority"` unsatisfiable and leaves every write exposed. - Prefer a planned `rs.stepDown()` over killing a primary during maintenance; a step-down gives secondaries a catch-up period first, which shrinks the divergence. - Monitor for rollback files appearing on any node. Their presence means acknowledged writes were lost and somebody should look. ## Diagnosing after the fact After an unplanned failover, check each rejoined member's log for the rollback transition and look for a `rollback` directory under its `dbPath`. The files name the affected collections and contain the affected documents. That is your complete inventory of what was lost, and it is the only one you will get.
- A write used j: true so it was flushed to the primary's journal. Can it still be rolled back?Yes. The `j` option is about durability on the node that acknowledged the write — it survives that process crashing and restarting. It says nothing about replication. If that node loses the election while holding the write and the new primary never received it, the entry is after the majority commit point and gets rolled back. Only `w: "majority"` makes a write survive a failover.
- How do you recover data from rollback files?Locate the `rollback` directory under the node's `dbPath` and read the BSON files it contains, one set per affected collection. Then decide per document whether reapplying it is still correct — the rest of the system has continued running, and blindly reinserting can violate invariants established by later writes. There is no automatic replay; treat it as a manual reconciliation with the application owner.
- How would you shrink the window of writes exposed to rollback without moving to majority write concern everywhere?Keep the majority commit point advancing quickly: place voters on a low-latency network, keep secondaries from lagging, and avoid a set where a single secondary loss makes majority acknowledgement impossible. Use planned `rs.stepDown()` for maintenance so the primary gives secondaries a catch-up period first. Then apply majority write concern selectively to the operations whose loss would actually be visible.
- Why does MongoDB let the newly elected primary's history win rather than preferring the node with more data?Because the elected primary is the only node guaranteed to hold every majority-committed write — a candidate must be at least as up to date as a majority to win the vote. Preferring the node with more entries would resurrect writes that a majority never saw and could reintroduce data the cluster had already agreed to discard, breaking convergence.
saying these in an interview costs you the question
- Thinks j true protects a write across failover
- Believes rolled-back writes replay automatically when the node rejoins
- Says clients are notified when their write is reverted
- Confuses rollback with a restore from backup
- Assumes any write on the primary is safe once it returns OK