skip to content

How do you add an index to a huge MongoDB collection on a live replica set safely?

level: seniorimportance: should knowfreq 46%

answer

  1. The old background option no longer means anything
  2. A brief exclusive lock at each end, not throughout
  3. All members build; a quorum must agree to commit
  4. Taking a member out means it stops replicating
  5. Watch the oplog window between steps

basics

~20 s

In current versions a normal createIndex already runs without blocking reads and writes for most of the build, taking an exclusive lock only at the start and end. When the resource cost on the primary is still unacceptable, build member by member with the rolling procedure.

solid answer

~50 s

Since MongoDB 4.2 every index build behaves the way the old background builds did: an exclusive collection lock is taken briefly at the beginning and again at the commit, and in between reads and writes proceed normally. The build runs on all data-bearing members simultaneously and the primary only commits once a **commit quorum** — `votingMembers` by default — has finished, so a failure on one secondary aborts it everywhere. What a plain build does still cost is CPU, disk I/O and cache, which on a saturated primary can show up as latency. The heavier alternative is a **rolling build**: take one secondary out of the set at a time, restart it standalone, build the index, put it back and let it catch up from the oplog, repeat, then step down the primary and do the last one. That trades operational complexity and reduced redundancy for keeping the build load off the serving primary.

code

javascript · 6 lines
javascript
// Modern build: no background option, optional commit quorum
db.orders.createIndex(
  { customerId: 1, createdAt: -1 },
  { name: "cust_created" },
  "votingMembers"
)

go deeper

for a junior

Know that createIndex on a live system does not block the collection for the whole build in current versions, and that the old background option is no longer meaningful.

for a middle

Explain the brief exclusive lock at start and end, the simultaneous build across members, the commit quorum, and that resource cost — CPU, I/O, cache — remains even without locking.

for a senior

Walk the rolling procedure step by step and name its risks: oplog window, reduced redundancy, the step-down at the end, and stopping the balancer on a sharded cluster.

for a principal

Own the decision itself: whether the index is worth its permanent write and cache cost, whether an existing index already covers the query, and what maintenance window and rollback the change deserves.

## Start with what a plain build already does The first thing to say is that the old foreground/background distinction is gone. From MongoDB 4.2 onwards, `createIndex` uses a single build process that holds an exclusive collection lock only **at the start and at the end** of the build. In between, it yields, so reads and writes against the collection continue; writes that arrive during the build are captured and applied to the new index before it commits. The `background: true` option that older tutorials tell you to pass is obsolete. The build is also a replica-set-wide operation. It starts on every data-bearing member at roughly the same time, and the primary does not commit the index until a **commit quorum** of members reports that their build finished. The default quorum is `votingMembers`, and you can override it per build: ```javascript db.orders.createIndex({ customerId: 1, createdAt: -1 }, { name: "cust_created" }, "votingMembers") ``` The practical benefit is that a failure — a duplicate key on a unique index, out of disk on one node — aborts the build everywhere rather than leaving members with different index sets. ## What it still costs No lock does not mean no impact. A build over a large collection reads the whole collection, sorts keys with a bounded amount of memory (governed by the `maxIndexBuildMemoryUsageMegabytes` server parameter, spilling to disk beyond it), and writes a new index structure. That is sustained CPU and I/O on every member at once, plus cache pressure that evicts hot documents. On a cluster already near its ceiling, that can be enough to push p99 latency past your SLO even though nothing is blocked. You can watch progress with `db.currentOp()` filtered to index builds, and abort a build with `killOp` if it is hurting. ## The rolling build When you cannot afford that load on the serving members, the rolling procedure keeps it off all but one node at a time. In outline, for each secondary: 1. Stop the member and restart it as a **standalone** — no replica set name, on a different port so application traffic and other members cannot reach it. 2. Build the index locally. Nothing is serving from this node, so the build can go as hard as the hardware allows. 3. Shut it down and restart it as a normal member of the set on its usual port. 4. Wait for it to catch up from the oplog and return to a healthy state before touching the next member. When every secondary carries the index, you `rs.stepDown()` the primary, wait for the new primary to be elected, and repeat the procedure on the demoted node. ## The risks you must name An interviewer is listening for whether you understand what the procedure gives up: - **Oplog window.** While a member is standalone, it is not replicating. If the build takes longer than the primary's oplog retains, the member cannot catch up when it rejoins and needs a full initial resync — which is far more expensive than the build was. Size the oplog against the expected build time before you start, and check the window between steps. - **Reduced redundancy.** With one member out of the set, you are one failure away from losing majority. On a three-member set that is a real exposure for the whole duration; do it during a low-traffic window and never take down two members. - **A step-down is a disruption.** The final step forces an election, so in-flight writes see errors and clients reconnect. Retryable writes absorb most of it, but it is not free. - **Sharded clusters need the balancer stopped.** A chunk migration in the middle of a rolling build moves documents between shards while some members have the index and others do not. Disable the balancer for the duration and re-enable it afterwards. - **Capacity while a member is out.** The remaining members carry the read and oplog load of the missing one. ## The cost that outlives the build The build is a one-time event; the index is forever. Every additional index is another structure written on each insert, and on each update that touches its key fields, plus index pages competing with documents for cache and a longer initial sync for any member you rebuild later. That is why the build conversation belongs with an index-review conversation: before you spend a rolling build, confirm from `explain` that the new index changes the plan you care about, and look for an existing index it makes redundant — the cheapest index to build is the one you do not need because a compound index already covers its prefix.

  • What is the commit quorum for an index build, and why does it exist?
    It is the number of data-bearing voting members that must finish building before the primary commits the index; the default is votingMembers. It exists so that a failure on any one member — a duplicate key on a unique index, a full disk — aborts the build everywhere instead of leaving the set with divergent index sets that would break failover behaviour.
  • Why must you check the oplog window before taking a member standalone for a rolling build?
    A standalone member stops replicating. When it rejoins it catches up by replaying oplog entries from the primary, so if the build outlasts the oplog's retention the entries it needs have aged out and the member requires a full initial resync. That is usually far more expensive and slower than the index build you were trying to make cheap.
  • On a sharded cluster, what must you do before a rolling index build?
    Stop the balancer for the duration. Otherwise a chunk migration can move documents to a shard whose members are in inconsistent index states mid-procedure. Re-enable it once every shard carries the index. You also run the whole procedure shard by shard, since each shard is its own replica set.

saying these in an interview costs you the question

  • Recommends background: true as the safe way to build
  • Says a normal build blocks all reads and writes for its duration
  • Claims a rolling build carries no availability risk
  • Ignores the oplog window while a member is standalone
  • Forgets that each extra index costs on every subsequent write

context