skip to content

A MongoDB secondary was offline overnight and now reports RECOVERING — what happened and how do you fix it?

level: seniorimportance: should knowfreq 55%

answer

  1. The log it needs to replay is finite
  2. Time versus bytes
  3. The stale member's own oplog is not the problem
  4. Rebuilding one member can endanger another

basics

~20 s

It fell outside the oplog window: the newest entry it applied is older than the oldest entry still retained on every potential sync source, so there is no continuous history to replay. It cannot resume tailing and must be rebuilt by initial sync or seeded from a recent copy.

solid answer

~50 s

Replication resumes by replaying oplog entries from where the member left off. The oplog is a **capped** collection, so it only holds a recent window of history. If a member is down long enough — or lags long enough — that its last applied entry has already been overwritten on every member it could sync from, the chain is broken: there is no way to get from its stale state to current, because the intervening changes are gone. MongoDB puts it in `RECOVERING` and it stays there. The recovery options are: wipe its `dbPath` and let it run a full **initial sync**, or seed its `dbPath` from a recent snapshot or backup of a healthy member whose position is still inside the window, then let it tail from there. The second is far cheaper on a large data set. The real fix is preventative: size the oplog so its window comfortably exceeds your longest realistic outage or maintenance window.

code

javascript · 5 lines
javascript
// current oplog window on a healthy member
rs.printReplicationInfo()

// per-secondary lag against the primary
rs.printSecondaryReplicationInfo()

go deeper

for a junior

Know that the oplog is finite and that a member gone too long cannot simply resume — it has to be rebuilt from a full copy.

for a middle

Explain the mechanism: last applied position older than the oldest retained entry means no replayable history, hence RECOVERING. Name both recovery routes and the role of the capped oplog.

for a senior

This is your tier. Diagnose it from rs.status() and the oplog window, choose between snapshot-seeding and full initial sync on cost grounds, and articulate how a rebuild's load can cascade into a second stale member.

for a principal

Own the prevention policy: oplog sizing and minimum retention as a fleet standard, snapshot cadence tied to rebuild time, and alerting expressed as headroom against the window rather than raw lag.

## Why this happens at all A secondary applies changes by replaying its sync source's oplog from the position it last reached. The oplog is a fixed-size capped collection: when it fills, the oldest entries are overwritten. The span of time between the oldest and newest entry still present is the **oplog window** — the maximum amount of history the set can replay to a member that has fallen behind. If a member's last applied entry is older than the oldest entry any sync source still holds, the history it would need no longer exists. There is no partial fix: the member cannot skip ahead, because skipping would leave it permanently divergent from the rest of the set. MongoDB refuses, marks the member `RECOVERING`, and logs that it could not find a member to sync from. The window is measured in **time**, but it is bought in **bytes**. A 50 GB oplog gives you a 24-hour window at one write rate and a 40-minute window at another. This is why the failure so often arrives during an incident: write volume spikes, the window silently collapses, and members that were comfortably inside it are suddenly outside it. ## The two ways in **A member was down.** Hardware failure, an OS patch that took longer than planned, a node cordoned during a rolling upgrade. The clock runs while it is gone. **A member was up but too slow.** Undersized disk, a heavy analytics query pinning resources, an index build, or simply a write rate the secondary's storage cannot sustain. Lag grows monotonically until it exceeds the window. This variant is more dangerous because dashboards show the member as `SECONDARY` and healthy right up until it is not. ## Getting it back **Option 1 — full initial sync.** Stop the `mongod`, delete the contents of its `dbPath`, restart it. As a member of the set with no data, it performs initial sync automatically: clone every database, build every index, apply the oplog produced during the clone. Correct and hands-off, but on a multi-terabyte data set it is hours to days, and it puts sustained read load on whichever member it syncs from. **Option 2 — seed from a recent copy.** Take a file-system snapshot or a consistent backup of a healthy member, place it in the stale member's `dbPath`, and start it. If the snapshot's oplog position is still inside the source's window, the member simply tails from there and is current in minutes. This is the standard technique at scale, and it is why teams that operate large sets keep recent snapshots specifically for member rebuilds, not just for disaster recovery. Neither option is "increase the oplog on the stale member". The window that matters belongs to the **sync source**, not to the member that fell behind. Enlarging the broken member's own oplog changes nothing about the history available to it. ## The resync storm The reason interviewers like this scenario is the second-order effect. Rebuilding one member drives a full data-set read against a healthy member and a full data-set write to the rebuilding one, for hours. That extra load slows the source. If the set was already close to the edge — say the write rate that shrank the window in the first place has not gone away — the source's own lag grows, another member drops outside the window, and you now have two rebuilds competing for the same disks. Sets have been taken from degraded to unavailable this way. Deliberately choosing the sync source, throttling the rebuild, or doing it during a genuine traffic trough are all mitigations; so is not starting two rebuilds at once. ## Preventing it - Size the oplog so the window exceeds your longest realistic outage, and re-derive that at peak write rate rather than average. `rs.printReplicationInfo()` reports the current window. - Set a floor on retention time so the window cannot silently collapse under a write burst. `storage.oplogMinRetentionHours` keeps entries for at least the configured number of hours regardless of the size limit. - Alert on the **window itself and on lag as a fraction of it**, not just on raw lag seconds. A 20-minute lag is fine with a 24-hour window and an emergency with a 30-minute one. - Keep recent snapshots so a rebuild never has to be a full clone.

  • Would enlarging the stale member's own oplog let it catch up?
    No. Catching up means replaying entries that must still exist on the member it syncs from. The stale member's own oplog only records changes it has itself applied. The window that matters belongs to the sync source, so raising the oplog size on healthy members helps future incidents, and raising it on the broken one helps nothing.
  • How would you rebuild a stale member on a multi-terabyte data set without hurting the primary?
    Seed its dbPath from a recent snapshot of a healthy secondary so it only has to tail, not clone. If a full initial sync is unavoidable, run it against a secondary rather than the primary, schedule it in a traffic trough, and rebuild one member at a time so you never have two full-data-set transfers competing for the same storage.
  • What alert would have caught this before the member went stale?
    Alert on the oplog window itself and on each member's lag expressed as a fraction of that window, not on lag in raw seconds. Lag of ten minutes is harmless with a day-long window and near-fatal with a fifteen-minute one. Pair it with an alert on the window shrinking, which catches write-rate spikes before any member is actually at risk.

saying these in an interview costs you the question

  • Suggests increasing the oplog size on the stale member
  • Assumes MongoDB will skip ahead and self-heal
  • Thinks restarting mongod repeatedly will fix it
  • Runs a full initial sync straight off the primary during an incident
  • Alerts only on lag seconds, never on the oplog window

context