A bulkWrite of 10,000 operations fails partway with a duplicate-key error — what has been applied, and how does ordered versus unordered change it?
answer
- Not a transaction
- One mode stops, the other keeps going
- What is left behind after the failure
- The error reports indexes for a reason
- Design so replay is harmless
basics
~20 sEverything already applied stays applied — a bulkWrite is not atomic across its operations. Ordered (the default) executes in array order and stops at the first error, leaving the tail untouched. Unordered continues past failures and reports all errors together.
solid answer
~50 s`bulkWrite` sends a batch of operations, but it is **not** a transaction: writes that succeeded before the failure remain in the collection and are not rolled back. With `ordered: true` — the default — the server executes the operations in array order and **stops at the first write error**, so you know a prefix applied and the remainder did not; the error tells you the index of the failing operation. With `{ ordered: false }` the server may execute them in any order and **keeps going**, so all valid operations apply and the error object carries a `writeErrors` array listing every failure by index. Recovery is therefore: read the error's indexes, decide which operations really failed, and re-run only those. The practical rule is to make each operation idempotent — upserts keyed on identity, `$set` rather than `$inc`, `$addToSet` rather than `$push` — so simply replaying the whole batch is safe.
code
javascript · 7 linesdb.events.bulkWrite([
{ insertOne: { document: { _id: 1 } } },
{ insertOne: { document: { _id: 1 } } }, // duplicate key
{ insertOne: { document: { _id: 2 } } }
], { ordered: false })
// _id 1 and _id 2 exist; the error reports index 1 in writeErrors
// with the default ordered: true, _id 2 would never be attemptedgo deeper
Know that bulkWrite sends many write operations in one call, that ordered is the default, and that it is not all-or-nothing.
Explain that ordered stops at the first write error while unordered continues and reports every failure, and describe what the writeErrors indexes let you reconstruct.
Diagnose a partial failure in production: classify the error codes, map indexes back to source records, retry only what failed, and design operations to be idempotent so replay is safe.
Set the ingestion contract — batch sizing, idempotency requirements on every write path, which failures are expected and swallowed, and how partial application is reconciled and observed.
## What bulkWrite is `db.collection.bulkWrite(operations, options)` takes an array of write operations — `insertOne`, `updateOne`, `updateMany`, `replaceOne`, `deleteOne`, `deleteMany` — and sends them to the server together instead of one round trip each. Its purpose is throughput: fewer network round trips and fewer command dispatches for the same amount of work. What it is **not** is an all-or-nothing unit. Each operation is applied independently. If the fiftieth fails, the forty-nine before it are already durable and stay that way. Candidates who describe `bulkWrite` as "a transaction" have the most important property backwards. ## Ordered: the default With `ordered: true`, which is the default when you pass no option, the server executes the operations serially in the order of the array and halts at the first **write error** — a duplicate key, a document-validation failure, an immutable-field violation. Operations after the failing index are not attempted at all. That gives a clean mental model for recovery: index `i` failed, indexes `0..i-1` applied, indexes `i..n-1` did not. Resuming means re-running from `i`. Ordered is what you want when later operations depend on earlier ones — insert the parent, then update the child that references it — because the ordering guarantee is what makes that dependency meaningful. ## Unordered With `{ ordered: false }` the server is free to execute the operations in any order, potentially grouping them, and it **continues past failures**. Every failure is collected, and the driver raises a bulk write error whose `writeErrors` array names each failed operation by its index in the array you submitted, along with the error code and message. All the operations that did not fail have applied. Unordered is normally faster, and the gap widens on a sharded cluster, where operations destined for different shards can be dispatched without waiting for each other — an ordered batch cannot do that, because preserving order across shards means serializing. It is the right default for independent work: ingesting a page of events, applying a set of unrelated per-document updates. ## What has actually been applied after a failure This is the heart of the question. In both modes, successful operations are permanent. The difference is only *which* ones were attempted: - Ordered: a prefix applied; everything from the failure onward is untouched. - Unordered: everything except the explicitly reported failures applied — and "except" is exactly what `writeErrors` enumerates. So the diagnosis after an incident is not "did the batch go through" but "which indexes are in `writeErrors`, and what does that imply about the source data". A duplicate-key error usually means the input contained a record that already exists, which for an ingest pipeline is often benign and for a ledger is not. ## Recovering Three steps in practice: 1. **Inspect the error.** Read `writeErrors` (or, in ordered mode, the single failure and its index) and classify the codes. Duplicate key is a different situation from a validation failure or an unauthorized write. 2. **Retry the failures only.** Because the error identifies operations by index into the array you submitted, you can map them back to the source records precisely. 3. **Prefer idempotent operations so a blind replay is safe.** An `updateOne` with `upsert: true` keyed on the record's natural identity can be replayed any number of times; so can `$set` and `$addToSet`. `insertOne` and `$inc` cannot — replaying them duplicates rows or double-counts. Designing the batch out of idempotent operations turns a partial failure from a reconciliation exercise into "run it again". When a duplicate-key error is *expected* — re-ingesting an overlapping window of events, for instance — the idiomatic handling is to run unordered and treat duplicate-key errors as successes, since the document you wanted is already there. ## Batching and size Drivers split a large operation array into batches according to the server's advertised `maxWriteBatchSize` and the maximum command size, so a 10,000-operation call already becomes several server-side batches. That is invisible to the ordered/unordered semantics but relevant to failure analysis: a batch is not a unit of atomicity either. ## Operational judgment Very large single calls make failure handling coarse and hold resources longer; moderate batch sizes with retry logic around each one are easier to operate. Choose ordered only when order genuinely matters — the guarantee costs throughput, especially when sharded — and make it a conscious decision rather than an accident of leaving the default in place.
- Is bulkWrite atomic across its operations?No. Each operation applies independently, and those that succeeded before a failure are durable and are not rolled back. Only the individual document write is indivisible. Grouping the operations into a multi-document transaction is a separate, explicit choice with its own costs; a plain bulkWrite offers no such guarantee.
- Which is faster, ordered or unordered, and why?Unordered, usually — the server is not required to preserve submission order, so it can group operations and does not have to stop and wait at each one. The difference is largest on a sharded cluster, where operations for different shards can proceed independently instead of being serialized to honour the ordering guarantee.
- How do you make a batch safe to replay after a partial failure?Build it from idempotent operations: `updateOne` with `upsert: true` keyed on the record's natural identity instead of `insertOne`, `$set` instead of `$inc`, `$addToSet` instead of `$push`. Then a blind re-run of the whole array converges to the same state, and you do not have to reconcile which indexes applied.
- When is ordered genuinely the right choice?When operations later in the array depend on earlier ones — a document must exist before another operation updates or references it, or a delete must precede an insert that reuses a unique key. There the ordering guarantee is the point, and stopping at the first error is the desired behaviour because continuing would compound the damage.
saying these in an interview costs you the question
- Calls bulkWrite atomic or a transaction
- Thinks a failure rolls back the operations already applied
- Believes unordered stops at the first error too
- Assumes ordered is faster because it is the default
- Retries the whole batch of non-idempotent inserts after a partial failure