How do you decide how large a MongoDB replica set's oplog should be?
answer
- You configure bytes but you consume time
- Peak write rate, not the daily average
- The default is a percentage of free disk, capped
- There is a command that resizes it live
- A retention floor guards against bursts
basics
~20 sSize it by the time window it buys, not by bytes: measure oplog growth at peak write rate and allocate enough that the window comfortably exceeds your longest expected member outage, maintenance window, or initial-sync duration.
solid answer
~50 sThe default on WiredTiger is 5% of free disk space, bounded to at least 990 MB and at most 50 GB — a starting guess, not a sizing decision. The number that matters is the **oplog window**: the time span between the oldest and newest retained entry, which `rs.printReplicationInfo()` reports. Measure it at your *peak* write rate, not the daily average, because a burst of bulk writes can collapse a day-long window to minutes. Then size so the window exceeds the longest interval a member might be behind: an OS patch cycle, a node replacement, a backup, or the duration of an initial sync elsewhere in the set. You can change it on a running WiredTiger member with the `replSetResizeOplog` admin command, no restart required, and pin a floor with `storage.oplogMinRetentionHours` so entries are kept for a minimum time regardless of the size limit.
code
javascript · 6 lines// resize to 16 GB on this member, keep at least 24h of history
db.adminCommand({
replSetResizeOplog: 1,
size: 16384,
minRetentionHours: 24
})go deeper
Know the oplog is a fixed-size capped collection and that its size determines how much recent history the set can replay to a member that fell behind.
Explain the bytes-to-time conversion and why write rate drives it. Be able to name the command that reports the current window and the one that resizes the oplog live.
Derive a target window from real intervals — patch cycles, node replacement, backup windows, initial-sync duration elsewhere in the set — and justify the headroom. Know the minimum-retention floor and its disk-space consequence.
Set the fleet standard: what window every tier must maintain, how it is monitored as headroom rather than raw lag, and the disk budget that buys it — weighed against the cost of an unplanned full resync during peak.
## What you are actually buying The oplog is a capped collection: a fixed byte allocation that overwrites its oldest entries when full. You configure **bytes**; what you consume is **time**. The oplog window — the span between the oldest and newest entry still present — is the real product, because it defines how long a member can be behind and still catch up by replaying rather than by rebuilding from scratch. That conversion is entirely determined by write rate. The same 50 GB allocation might hold three days of history for a read-heavy application and forty minutes for one that ingests bulk loads. Sizing from the average write rate is the classic mistake: the window is smallest exactly when you need it most, during the ETL job or the migration or the traffic spike that also happens to be when a member falls over. ## Defaults On WiredTiger, a new deployment allocates 5% of free disk space, with a lower bound of 990 MB and an upper bound of 50 GB. Two things follow. First, the number depends on how empty the disk was at deployment time, which is not a property that has anything to do with your workload. Second, on a large host it silently caps at 50 GB, so "we left the defaults" does not mean "it scaled with our data". Note also that MongoDB may let the oplog grow past its configured size when truncating entries would discard history the set still needs, so the configured size is a target rather than a rigid ceiling. ## Measuring what you have `rs.printReplicationInfo()` reports the configured size, the timestamps of the first and last entries, and the resulting log length in hours. Run it on the primary and on secondaries — the windows differ, because members may have different oplog sizes and different histories. `rs.printSecondaryReplicationInfo()` shows how far each secondary is behind. The pair of numbers is what matters: lag is only meaningful as a fraction of the window. Thirty minutes of lag with a 36-hour window is a curiosity; the same lag with a 45-minute window is an incident. ## Choosing a target Work from the longest interval a member could plausibly be behind and still be expected to rejoin without a rebuild: - **Planned maintenance.** How long does an OS patch, a storage migration, or a rolling version upgrade take per node, including the queue in front of it? - **Unplanned outage you would wait out.** A host that reboots and comes back in two hours should not need a full resync. - **Elsewhere-in-the-set load.** While one member is doing an hours-long initial sync, the others are working harder and may lag more. - **Backup windows.** A member taken out of rotation for a snapshot is a member accumulating lag. Then add real headroom — a window that only just covers your worst case has no margin for the write burst that arrives during it. ## Changing it On a running WiredTiger member you can resize the oplog without a restart: `db.adminCommand({ replSetResizeOplog: 1, size: 16384 })` The size is in megabytes and applies to that member only, so you roll the change across the set member by member. The same command accepts `minRetentionHours`, and the equivalent startup setting is `storage.oplogMinRetentionHours`: entries younger than that are retained even if the collection has reached its size limit. That floor is the single most useful guard against a write burst quietly collapsing the window, at the cost of the oplog consuming more disk than configured during such a burst — so you must have the disk headroom to honour it. ## Costs of over-sizing The oplog is not free: it occupies disk that could hold data, and a very large oplog on a member with tight storage is a genuine risk. But relative to the cost of an unnecessary multi-hour initial sync during an incident, oplogs are cheap. In practice teams under-size far more often than they over-size, and the fix is to decide the window you want, measure the bytes-per-hour your workload produces at peak, multiply, and buy the disk.
- Which command reports the current oplog window, and what does it show?`rs.printReplicationInfo()`, run against the member you care about. It reports the configured oplog size, the timestamps of the oldest and newest retained entries, and the resulting log length in hours. Pair it with `rs.printSecondaryReplicationInfo()`, which shows how far each secondary is behind, since lag is only interpretable as a fraction of the available window.
- How do you enlarge the oplog on a running production replica set?Run `db.adminCommand({ replSetResizeOplog: 1, size: <MB> })` against each WiredTiger member; no restart is needed. It applies per member, so roll it across the set. The same command takes `minRetentionHours` to set a time floor, and `storage.oplogMinRetentionHours` sets the equivalent at startup.
- What does storage.oplogMinRetentionHours change about capped-collection behaviour?It stops the oplog from truncating entries younger than the configured number of hours, even when the collection has reached its size limit. The oplog is then allowed to exceed its configured size to honour the floor. That protects the window during a write burst, but only if the disk has room — otherwise you have converted a resync risk into a disk-full risk.
saying these in an interview costs you the question
- Sizes the oplog from average rather than peak write rate
- Thinks the default 5% scales with the data set
- Believes resizing the oplog requires a restart
- Treats lag in seconds as meaningful without knowing the window
- Assumes the oplog can never exceed its configured size