skip to content

Why does MongoDB rewrite operations into idempotent form before writing them to the oplog?

level: middleimportance: must knowfreq 62%

answer

  1. Replay after a crash is normal, not exceptional
  2. Think about what happens if it runs twice
  3. $inc is the classic example
  4. The log records effects, not statements
  5. updateMany does not produce one entry

basics

~20 s

So an entry can be replayed any number of times with the same result. Secondaries and recovering members may re-apply the tail of the oplog after a restart or a sync-source switch, and a relative operation replayed twice would corrupt the data.

solid answer

~50 s

The oplog is a capped collection, `oplog.rs` in the `local` database, holding one entry per logical change. Secondaries replicate by tailing it and applying entries, and there is no exactly-once delivery: a member that restarts, switches sync source, or resumes after a network blip will re-apply entries it may already have applied. MongoDB therefore records each operation in a form whose result does not depend on how many times it runs. A relative update such as `{ $inc: { qty: 1 } }` is not stored as `$inc`; it is stored as the **resulting absolute value**, so replaying it is a no-op. A multi-document `updateMany` is expanded into one entry per affected document, each keyed by `_id`, so that partially-applied batches converge correctly. Idempotency is what lets replication recovery be "just replay from a known safe point".

code

javascript · 2 lines
javascript
// application write against a document where qty is 11
db.items.updateOne({ _id: 7 }, { $inc: { qty: 1 } })

go deeper

for a junior

Know that the oplog is the log of changes secondaries replay, and that it lives in a fixed-size capped collection. The word idempotent — same result no matter how many times it runs — is worth having ready.

for a middle

This is your tier. Explain the $inc rewrite concretely, say why re-application happens during normal operation, and describe how a multi-document update fans out into per-document entries.

for a senior

Connect idempotency to operations: crash recovery, sync-source switches, and change-stream consumers all rely on it. Be able to explain what would break if the log recorded statements instead of effects.

for a principal

Frame the design tradeoff: a physical, effect-based log is trivially replayable but larger and version-coupled, whereas a logical statement log is compact but demands determinism everywhere. Know which properties your downstream consumers depend on.

## What the oplog is The oplog ("operations log") is a special **capped collection** named `oplog.rs`, stored in the `local` database of every data-bearing replica set member. Capped means it is a fixed-size ring: once full, the oldest entries are overwritten to make room for new ones. `local` is never itself replicated, so each member owns its own oplog. The primary appends an entry for every change it applies. Each entry carries, among other fields, a timestamp `ts` that uniquely orders it, an operation type `op` (`i` insert, `u` update, `d` delete, `c` command, `n` no-op), the namespace `ns`, and the operation payload `o` (plus `o2` carrying the `_id` for updates). ## Why replay is not exactly-once Secondaries replicate by continuously tailing their sync source's oplog and applying the entries. Nothing in that flow gives exactly-once semantics: - A secondary restarts. It resumes applying from the last point it knows is durable on disk, which may be slightly behind what it actually applied, so the tail gets replayed. - A secondary switches sync source after a failover or a network problem and re-reads overlapping entries. - The applier crashes mid-batch and re-runs the batch from its start. Because replay is normal rather than exceptional, replication correctness cannot depend on each entry being applied exactly once. It depends on each entry being **idempotent**: applying it once and applying it five times must leave the document in the same state. ## How MongoDB achieves it The key move is that the oplog does not record the *statement the client sent*; it records the *effect on each document*, in absolute terms. **Relative operators become absolute values.** An application that runs `{ $inc: { qty: 1 } }` against a document where `qty` was 11 does not produce an oplog entry containing `$inc`. It produces an entry that sets `qty` to 12. Replay that entry a hundred times and `qty` is still 12. Had the literal `$inc` been logged, each replay would have incremented again and the secondary would silently diverge from the primary. **Multi-document writes are fanned out.** A single `updateMany` or `deleteMany` that touches 500 documents does not become one oplog entry. It becomes 500 entries, each targeting one document by `_id` in the `o2` field. This matters for two reasons: a re-run of a partially applied batch converges, and an entry never depends on re-evaluating a query predicate whose matching set might have changed since. **Non-deterministic values are resolved at logging time.** Anything the primary computed — a generated `_id`, a server-side date, the result of an upsert deciding to insert rather than update — is baked into the entry as a concrete value, so the secondary does not recompute and get a different answer. **Recent versions log a compact delta.** Newer MongoDB releases record updates as a versioned delta (a `$v: 2` entry) rather than a full `$set` document, to save space on wide documents. The encoding changed; the property did not — the delta still expresses the resulting field values, not a relative adjustment. ## What idempotency does not give you Idempotency is per-entry. It does not mean order is irrelevant: entries affecting the same document must be applied in oplog order, and MongoDB's applier — which uses multiple writer threads for throughput — guarantees exactly that per document while parallelising across documents. Nor does idempotency make the oplog a general-purpose audit log: it is a fixed-size ring that overwrites its oldest entries, so it is not a durable history of everything that ever happened. ## Why interviewers ask this It is the cleanest single question that proves a candidate understands *how* replication works rather than that it exists. The tell of a strong answer is the `$inc` example: recognising that logging the user's operator verbatim would be a correctness bug, and that logging the outcome instead is what makes crash recovery a simple replay. It also sets up everything downstream — change streams, initial sync, and rollback all rest on the same log.

  • How is a single updateMany that modifies 500 documents represented in the oplog?
    As 500 separate entries, one per modified document, each identifying its target by `_id` in the `o2` field. The query predicate is never replayed. This keeps each entry idempotent and independent, so a partially applied batch converges on re-run and a secondary never has to re-evaluate a filter whose matching set may have changed.
  • Does idempotency mean a secondary can apply oplog entries in any order?
    No. Idempotency covers repeated application, not reordering. Entries touching the same document must be applied in oplog order or the final state would be wrong. MongoDB's applier parallelises a batch across multiple writer threads for throughput while preserving per-document ordering, which is why it can go faster than a strictly serial replay.
  • Can you use the oplog as a permanent audit trail of every change?
    No. The oplog is a capped, fixed-size collection: once it is full, the oldest entries are overwritten. It only ever holds a recent window of history, sized by disk allocation and write rate. For durable change history you consume it continuously — change streams are the supported interface — and persist the events elsewhere.

saying these in an interview costs you the question

  • Says the oplog stores the client's original update statement
  • Thinks each oplog entry is applied exactly once
  • Claims replaying $inc twice is harmless
  • Describes the oplog as a permanent change history
  • Says one updateMany produces one oplog entry

context