skip to content

If wtimeout expires on a MongoDB write with w: "majority", has the write been rolled back?

level: middleimportance: should knowfreq 52%

answer

  1. the clock bounds the wait, not the change
  2. the primary applied it before waiting started
  3. an error here is not an undo
  4. think 'unknown outcome', then idempotent retry

basics

~10 s

No. wtimeout only bounds how long the primary waits for replication acknowledgment. The write is already applied on the primary and usually replicates anyway; the error means the outcome is unknown, not undone.

solid answer

~50 s

`wtimeout` caps the wait, not the write. By the time the primary is waiting for other members it has already applied the change locally, so an expired `wtimeout` produces a **write concern error** while the write itself remains in place and normally reaches a majority a moment later. There is no rollback triggered by the timeout. The only way that write disappears is the ordinary rollback path: the primary steps down before the write ever becomes majority-committed, and the ex-primary discards it on rejoin. So treat the error as "outcome unknown": the operation may already have succeeded, so a blind retry must be idempotent or use retryable writes and a unique key. Two related details matter — `wtimeout` applies only when `w` is greater than one, and without it a majority write on a degraded set waits indefinitely.

code

javascript · 10 lines
javascript
try {
  db.orders.insertOne(
    { _id: "ord-7", total: 199 },
    { writeConcern: { w: "majority", wtimeout: 3000 } }
  )
} catch (e) {
  // e.writeConcernError set, e.writeErrors empty:
  // the document is on the primary; only confirmation timed out
  print(e.code)
}

go deeper

for a junior

Remember the headline: a timed-out write concern does not undo anything. The document is already on the primary, and the error is about confirmation.

for a middle

Explain the sequence — apply locally, then wait for other members — and why that ordering makes rollback-on-timeout impossible, plus why the error arrives as a write concern error rather than a write error.

for a senior

Demonstrate the operational response: treat the outcome as unknown, make retries idempotent, alert on the error rate as a replication-health signal, and know why omitting the timeout turns a degraded set into an application-wide hang.

for a principal

Set the house rule: every majority write carries a bounded timeout, write concern errors are a monitored signal rather than an application exception, and no service is allowed to run compensation logic that assumes a timed-out write failed.

## What `wtimeout` actually times The order of events for a write with `{ w: "majority", wtimeout: 500 }` is: 1. The primary validates and applies the change and appends an entry to its oplog. 2. The primary starts waiting for enough other members to report that they have applied that oplog entry. 3. Secondaries pull and apply the entry and report their positions. 4. Once a majority is reached, the primary answers the client. `wtimeout` only bounds **step 2**. Step 1 has already happened. That is why an expired `wtimeout` cannot mean "the write did not happen" — it means "the write happened here, but I could not confirm within your deadline that enough other members also have it". ## What the client sees The server does not report this as an ordinary command failure. The command result carries a **write concern error** alongside an otherwise successful write result, and drivers surface it as a distinct error type (a write concern error rather than a write error). That distinction is the whole point: the data operation succeeded, the durability promise was not confirmed. Nothing in this path undoes anything. In the overwhelmingly common case — a secondary that was briefly slow, a burst of replication traffic, a member restarting — replication catches up milliseconds later and the write becomes majority-committed after the client already saw an error. ## When the write really can vanish The write is at risk only through the normal rollback mechanism, which has nothing to do with your timeout value: if the primary steps down while that write is still not majority-committed, and a member without the write is elected, the ex-primary discards it when it rejoins (writing it to a rollback file rather than replaying it). An expired `wtimeout` is a *signal* that you are in the window where that risk exists, not the cause of it. ## How to handle the error correctly Treat it as **indeterminate**, exactly as you would treat a network timeout on any request: - Do not tell the user the operation failed; it very likely succeeded. - Do not blindly retry a non-idempotent write. Retry only writes that are naturally idempotent, that carry a unique key (so the retry fails cleanly with a duplicate-key error), or that go through MongoDB's retryable-writes mechanism. - If you must know, re-read the document — with a read concern strong enough to be meaningful — rather than guessing. - Alert on the rate. A steady trickle of write concern errors is a health signal about replication lag or a degraded member, not an application bug. ## Choosing a value Two rules keep people out of trouble. First, `wtimeout` has no effect when `w` is `1` or `0` — there is no waiting-for-others phase to bound. It only means something for `w` greater than one, including `w: "majority"`. Second, **omitting it is a real risk**. Without `wtimeout`, a majority write on a set that has lost a data-bearing member blocks indefinitely, which turns a replication problem into an application-wide thread or connection exhaustion problem. A bounded `wtimeout` converts an unbounded hang into an error you can observe, shed load on, and page about. Pick a value comfortably above your normal replication acknowledgment time — typically a few seconds — so that ordinary jitter never trips it. ## A related trap Because the write survives the timeout, code that reacts to a write concern error by "cleaning up" — deleting the document, reversing the balance, marking the order failed — can actively destroy a write that was about to be perfectly durable. Compensation logic must first establish what actually happened, not assume failure. ## Summary `wtimeout` is a deadline on confirmation, not a transaction abort. The write stays. What you have lost is the guarantee, so the correct response is to treat the outcome as unknown and to make your retry path safe under duplication.

  • Given that the write is not rolled back, what is the point of setting wtimeout at all?
    It bounds the failure. Without it, a majority write on a set that has lost a data-bearing member waits indefinitely, tying up application threads and connections until the whole service stalls. With it, you get a prompt, observable error you can log, rate-limit or shed load on, while the write itself continues toward majority commit in the background.
  • How should application code retry after a write concern error?
    Only in a way that is safe if the first write already landed: use a client-supplied unique `_id` or unique index so a duplicate retry fails cleanly, rely on MongoDB's retryable writes for supported single-document operations, or re-read the document to establish the real state. Never issue a blind retry of a non-idempotent update such as an unconditional increment.
  • Does wtimeout do anything when the write concern is w: 1?
    No. `wtimeout` bounds only the phase where the primary waits for other members to acknowledge, so it is meaningful only when `w` is greater than one, including `w: "majority"`. With `w: 1` the primary answers as soon as it has applied the write locally, and with `w: 0` it does not answer at all.

saying these in an interview costs you the question

  • Says the write is rolled back when wtimeout expires
  • Retries a non-idempotent write blindly after the error
  • Runs compensation logic that deletes the document on a write concern error
  • Omits wtimeout and lets majority writes hang forever
  • Confuses a write concern error with a duplicate-key or validation write error

context