skip to content

How does initial sync differ from steady-state oplog tailing in a MongoDB replica set?

level: middleimportance: should knowfreq 52%

answer

  1. A new member has no baseline to apply changes to
  2. One is a one-off, one runs forever
  3. Writes keep arriving during the copy
  4. rs.status() shows a distinct state for it

basics

~20 s

Initial sync copies an entire data set from a sync source to a brand-new or wiped member, building indexes and applying oplog entries produced during the copy. Steady-state tailing then applies each new oplog entry as it appears, incrementally and forever.

solid answer

~40 s

A member with no data — freshly added, or restarted with an emptied `dbPath` — cannot tail an oplog, because it has no baseline to apply changes to. It performs **initial sync**: it picks a sync source, clones every database except `local`, builds indexes, and applies the oplog entries the source produced while the clone was running, so that it finishes at a consistent point in time. Its state shows as `STARTUP2` throughout. Once it catches up it transitions to `SECONDARY` and switches to **steady-state replication**: an open tailing cursor on the source's `oplog.rs`, pulling batches of new entries and applying them with multiple writer threads. Initial sync is a one-off, hours-long, bandwidth-heavy operation that reads the whole data set; tailing is continuous and proportional only to the write rate.

go deeper

for a junior

Recall that a brand-new member must first copy the whole data set before it can follow changes, and that this takes far longer than normal replication.

for a middle

Explain the phases: pick a source, clone, build indexes, apply the oplog produced during the clone, then transition to tailing. Know which state rs.status() shows at each point.

for a senior

Speak to cost and blast radius: how long a sync takes at your data size, why you steer it away from the primary, and how seeding from a snapshot avoids a full clone during an incident.

for a principal

Own the capacity view — how data growth and index count set your recovery time for losing a member, and what oplog sizing, backup cadence, and member placement you fund so a rebuild is never on the critical path.

## The bootstrapping problem Replication in MongoDB works by replaying an oplog — a log of changes — onto an existing copy of the data. That works only if the member already has a copy that matches some point in the log. A brand-new member has nothing. Initial sync is how it acquires that baseline. Initial sync happens automatically when a `mongod` joins a replica set with an empty `dbPath`, and it is also the standard remedy when an existing member has fallen so far behind that the entries it needs are no longer in any sync source's oplog. ## What logical initial sync does The default method is *logical* initial sync — the syncing member reads documents through the normal query interface and writes them locally. Broadly: 1. **Choose a sync source.** The member picks a reachable, sufficiently up-to-date member. It need not be the primary; a secondary is often chosen, which is desirable because cloning is heavy. 2. **Record a start point.** It notes where the source's oplog stands before cloning begins. 3. **Clone the data.** It copies every database except `local`, collection by collection, and builds the indexes for each collection as it goes. 4. **Apply the oplog written during the clone.** Cloning a large data set takes hours, and the source keeps accepting writes throughout. Those writes are pulled from the source's oplog and applied, so the member converges on a consistent snapshot rather than a smeared mixture of old and new documents. 5. **Transition to SECONDARY.** Once it has caught up, the member starts ordinary tailing. The replica set state during all of this is `STARTUP2`, which is how you tell from `rs.status()` that a member is syncing rather than merely lagging. ## Failure and resumability Historically an interrupted initial sync had to start over from zero — brutal on a multi-terabyte data set. Recent versions can resume after transient network failures within a bounded retry period, and will restart the attempt a limited number of times before giving up. MongoDB Enterprise also offers a **file-copy-based** initial sync (selected with the `initialSyncMethod` parameter) that copies the storage-engine files directly instead of reading documents; it is typically faster on large data sets but has its own constraints. ## Steady-state tailing Once it is a `SECONDARY`, the member's behaviour is completely different in shape. It holds a tailing cursor open against the sync source's `oplog.rs` and receives batches of new entries as they are written. It applies each batch using several writer threads in parallel, preserving order per document, then records its progress so the primary and the rest of the set can see how current it is. The sync source is not fixed. If chained replication is enabled — it is by default, via the `chainingAllowed` setting — a secondary may sync from another secondary that is ahead of it, and it will change source if its current one becomes unsuitable. ## Comparing the two | | Initial sync | Steady-state tailing | |---|---|---| | Trigger | Empty `dbPath`, or a member too stale to catch up | Normal operation of a `SECONDARY` | | Data moved | The entire data set | Only new oplog entries | | Duration | Hours to days at scale | Continuous | | Cost driver | Total data size and index build time | Write rate on the primary | | Member state | `STARTUP2` | `SECONDARY` | ## Operational consequences The practical point is that initial sync is *expensive*, and every design decision that makes it more likely — undersized oplogs, long maintenance windows, aggressive restarts — buys you a multi-hour whole-data-set transfer at the least convenient moment. Choosing a secondary as the sync source, or seeding a new member's `dbPath` from a recent snapshot of another member so it only needs to tail rather than clone, are the usual ways to keep the primary out of harm's way. Initial sync also builds every index from scratch, which on an index-heavy collection can dominate the elapsed time — so index count, not just data size, is part of the cost.

  • What replica set state does rs.status() report while a member is performing initial sync?
    `STARTUP2`. That distinguishes a member building its first copy of the data from a `SECONDARY` that merely has replication lag, and from `RECOVERING`, which indicates a member that cannot currently serve reads and may be too stale to resume tailing. Watching `STARTUP2` linger for hours on a large data set is normal, not a fault.
  • Why does initial sync apply oplog entries taken during the clone phase?
    Because the source keeps accepting writes for the whole hours-long clone. Without replaying those changes the new member would hold a mixture of documents copied at different moments — internally inconsistent, and inconsistent with the source. Applying the oplog written during the clone converges the member on a single consistent point in time before it becomes a secondary.
  • How can you add a new member to a large replica set without a full initial sync?
    Seed its `dbPath` from a recent file-system snapshot or backup of an existing member, then start it in the set. If the snapshot's position is still inside the sync source's oplog window, the member just tails from there and catches up in minutes instead of cloning terabytes. If the snapshot is older than the window, it falls back to a full initial sync.

saying these in an interview costs you the question

  • Thinks a new member starts by tailing the oplog immediately
  • Believes initial sync must read from the primary
  • Says writes are blocked on the source during cloning
  • Assumes an interrupted initial sync always restarts from zero
  • Confuses STARTUP2 with ordinary replication lag

context