Which MongoDB write operations does retryWrites=true actually retry after a failover?
answer
- Think about what the server has to remember
- The eligible set is narrower than people assume
- One document per write is the dividing line
- updateMany sits on the wrong side of it
- Session id plus transaction number identifies the statement
basics
~10 sOnly single-document writes: insertOne, insertMany, updateOne, replaceOne, deleteOne, the findOneAndX family, and bulkWrite made only of those. Multi-document commands such as updateMany and deleteMany are never retried, and neither are unacknowledged writes.
solid answer
~40 sWith `retryWrites=true` — the default in current drivers — the driver attaches the session id and a transaction number to each eligible write, and the primary records that statement's outcome. If the driver hits a retryable error (a network failure, a stepdown, an election), it retries once; the new primary recognises the same session and transaction number and returns the recorded result rather than applying the write a second time. Eligibility is the point interviewers probe: **single-document** writes qualify — `insertOne`, `insertMany`, `updateOne`, `replaceOne`, `deleteOne`, `findOneAndUpdate`/`Replace`/`Delete`, and a `bulkWrite` containing only such operations. `updateMany` and `deleteMany` do not, because they are not safely re-appliable. Unacknowledged writes (`w: 0`) are excluded, and the deployment must be a replica set or sharded cluster. Application errors like a duplicate key are never retried.
code
javascript · 8 lines// retryWrites is on by default; shown explicitly
// mongodb://host1,host2,host3/?replicaSet=rs0&retryWrites=true
await orders.updateOne( // retryable
{ _id: id }, { $inc: { attempts: 1 } });
await orders.updateMany( // NOT retryable
{ status: "pending" }, { $inc: { attempts: 1 } });go deeper
Recall that modern drivers retry writes automatically after a network blip or failover, and that this covers single-document writes only — not updateMany or deleteMany.
Explain the mechanism: a session id plus a transaction number lets the server record the statement's outcome, so a retry returns the recorded result rather than applying the write twice.
Show where the guarantee ends in production: driver-level retry covers one hop, so client resubmissions still need a real idempotency key, and bulk changes need idempotent updates or resumable batches.
Set the platform policy — retryWrites on everywhere, an agreed idempotency-key convention at API boundaries, and a documented pattern for bulk mutations so teams do not assume the driver has already solved exactly-once.
## The problem retryable writes solve A driver sends `insertOne` to the primary. The connection drops before a response arrives. Did the insert apply? The client cannot tell. Retrying blindly may duplicate the document; not retrying surfaces an error for a write that may well have succeeded. Elections make this routine: every primary stepdown kills in-flight operations on a healthy cluster doing a rolling restart. ## The mechanism When `retryWrites=true` (the default in modern drivers, and settable in the connection string), the driver: 1. Associates the write with a **logical session id** (`lsid`) — implicit if you did not create one. 2. Attaches a monotonically increasing **transaction number** (`txnNumber`) to the write command. The primary, before acknowledging, records the outcome of that `(lsid, txnNumber)` statement in an internal collection in the `config` database. If the driver retries after a retryable error, the write arrives at the (possibly new) primary carrying the *same* session id and transaction number. If the server already has a record for it — the record replicates with the rest of the data — it returns the stored result instead of applying the write again. That is what makes the retry exactly-once rather than at-least-once. Drivers retry **once**. If the retry also fails retryably, the error surfaces to the application. Session records are not kept forever: logical sessions time out after a period of inactivity (30 minutes by default), which is far longer than any retry window. ## Which writes are eligible Eligible (single-document semantics): - `insertOne`, `insertMany` - `updateOne`, `replaceOne`, `deleteOne` - `findOneAndUpdate`, `findOneAndReplace`, `findOneAndDelete` - `bulkWrite` whose operations are all of the above Not eligible: - `updateMany`, `deleteMany`, and any `bulkWrite` containing them - writes with write concern `w: 0`, because there is no acknowledgement to lose The reason multi-document commands are excluded is that the server cannot record a single result for a command that may have been half-applied across many documents; replaying it could apply non-idempotent updates (`$inc`, `$push`) twice to documents that already got them. The deployment matters too: retryable writes require a replica set or a sharded cluster. A standalone `mongod` has no mechanism for them. ## Which errors are retried Network errors, and server errors carrying the `RetryableWriteError` label — not primary, node shutting down, election in progress, and similar transient conditions. Explicitly **not** retried: duplicate key errors, document validation failures, authorization errors, and anything else that reflects the state of your data rather than the state of the cluster. Retrying those would just fail again. ## Relationship to transactions Retryable writes and multi-document transactions share plumbing — both ride on a session and a transaction number — but they are different tools. A retryable write covers one write command through one failover. A transaction covers many operations and gives them all-or-nothing semantics and a read snapshot. Inside a transaction the retry unit changes: individual operations are not independently retried; instead the driver retries the *whole* transaction on a transient error, or re-sends the *commit* when the commit outcome is unknown. Turning on `retryWrites` is cheap and near-universal; opening a transaction is a deliberate design decision with real cost. ## What retryable writes do not give you - **They are not application-level idempotency.** They cover the driver's own one-hop retry. If your service times out and a *user* or an upstream retries the request, the second attempt is a new operation with a new transaction number and will apply again. For that you need a natural idempotency key — a deterministic `_id`, an upsert on a business key, or a stored request id. - **They do not make `updateMany` safe.** If you need failover-safe semantics for a bulk change, either make the update naturally idempotent (`$set` to a fixed value rather than `$inc`), chunk it into single-document writes, or drive it from a resumable job that can re-scan for unfinished work. - **They do not change durability.** Whether the write survives a subsequent failure is still write concern's job. ## What an interviewer wants The give-away weak answer is "the driver just retries writes when the network fails." The strong answer names the session id plus transaction number, explains the server-side record that makes the retry non-duplicating, draws the single-document eligibility line with `updateMany` on the wrong side of it, and separates driver-level retry from application-level idempotency.
- Why can updateMany not be made retryable the way updateOne is?The server records one result per statement, keyed by session id and transaction number, so a retry can return the recorded outcome. A multi-document command may have been applied to an arbitrary subset before the failure, and there is no single recorded outcome that describes it. Replaying it would re-apply non-idempotent operators such as `$inc` or `$push` to documents that already received them, so the drivers exclude it outright.
- Your API endpoint times out and the client resubmits. Do retryable writes protect you from a duplicate order?No. Retryable writes only cover the driver's own retry of one command inside a single request. A fresh request from the client carries a new session and transaction number, so the write applies again. You need application-level idempotency: derive `_id` deterministically from a client-supplied idempotency key, or upsert on a unique business key so the second attempt is rejected or absorbed.
- Which errors will the driver refuse to retry even with retryWrites enabled?Anything that reflects the data rather than the cluster: duplicate key violations, document validation failures, authorization errors, and malformed commands. Retrying those would deterministically fail again. Retries are reserved for network errors and server errors carrying the retryable-write label — stepdowns, elections, and nodes shutting down.
It is like a parcel with a tracking number: if you resend it, the depot sees the same number, recognises it has already been delivered, and hands back the delivery receipt instead of delivering twice.
saying these in an interview costs you the question
- Believes every write is retried, including updateMany
- Says retries can duplicate the document
- Treats retryable writes as full application-level idempotency
- Thinks retryable writes work against a standalone mongod
- Expects duplicate key errors to be retried