skip to content

In MongoDB, what does a unique index enforce, and what happens when a write violates it?

level: juniorimportance: must knowfreq 70%

answer

  1. It is a constraint, not just an access path
  2. The write fails rather than overwriting
  3. Error code 11000
  4. Compound means the tuple is unique, not each field
  5. Missing field is indexed as null

basics

~20 s

A unique index rejects any write that would create a second document with the same indexed value, returning an E11000 duplicate key error. On a compound unique index the whole key combination must be unique, not each field on its own.

solid answer

~40 s

You create one with `db.users.createIndex({ email: 1 }, { unique: true })`. After that, any insert or update that would produce a second document with the same `email` fails with an **E11000 duplicate key error** and the write is rejected — nothing is overwritten or silently skipped. On a compound unique index such as `{ tenantId: 1, slug: 1 }` the constraint applies to the whole combination, so the same `slug` under two different tenants is allowed. Building the index on a collection that already contains duplicates fails, so the data has to be cleaned first. The `_id` index is unique by definition and cannot be dropped or altered. One classic gotcha: a document missing the indexed field is indexed as `null`, so only one such document can exist under a plain unique index.

code

javascript · 6 lines
javascript
db.users.createIndex({ email: 1 }, { unique: true })

db.users.insertOne({ email: "[email protected]" })
db.users.insertOne({ email: "[email protected]" })
// MongoServerError: E11000 duplicate key error
//   collection: app.users index: email_1 dup key: { email: "[email protected]" }

go deeper

for a junior

Be able to write the createIndex call with unique: true, name the E11000 duplicate key error, and state that the offending write is rejected rather than overwriting anything.

for a middle

Explain compound uniqueness over the tuple, why the build fails on pre-existing duplicates, and why an absent field is indexed as null and therefore allowed only once.

for a senior

Show that the index is the only real concurrency-safe guarantee, that the driver error must be mapped to a domain error, and how you would repair duplicates on a live collection before building it.

for a principal

Own the design consequence: uniqueness constraints interact with the shard key, so a constraint you promise the business may dictate the partitioning scheme or force a separate keyed collection to hold it.

## What a unique index is A MongoDB index is normally just an access path: a sorted structure that lets the planner find documents without scanning the whole collection. Adding `unique: true` turns that access path into a **constraint** as well. The server refuses to store two index entries with the same key, and because the index is maintained inside the same write that modifies the document, the check is atomic with the write — there is no window in which a duplicate briefly exists. ```javascript db.users.createIndex({ email: 1 }, { unique: true }) ``` ## What a violation looks like An `insertOne`, `updateOne`, `replaceOne`, an upsert or a `bulkWrite` operation that would produce a second entry with an existing key fails with a **duplicate key error, code 11000**, usually printed as `E11000 duplicate key error collection: app.users index: email_1 dup key: { email: "[email protected]" }`. The important part is what does *not* happen: the existing document is not overwritten, the duplicate is not silently dropped, and the index is not left in a half-built state. The offending write simply does not apply. In an unordered `bulkWrite` the other operations in the batch still run and the duplicates come back in the write-error list; in an ordered bulk write, processing stops at the first failure. Application code normally catches error code 11000 specifically and turns it into a domain-level "already taken" response rather than a 500. ## Compound unique indexes On `{ tenantId: 1, slug: 1 }` with `unique: true`, uniqueness is over the **tuple**, not over each field. These two documents coexist happily: ```json { "tenantId": "acme", "slug": "welcome" } { "tenantId": "globex", "slug": "welcome" } ``` and only a second `{ "tenantId": "acme", "slug": "welcome" }` is rejected. This is the standard way to express "unique per tenant", "one vote per user per poll", and similar scoped constraints. Candidates who assume a compound unique index makes each column unique individually get this wrong in interviews constantly. ## Building on existing data `createIndex(..., { unique: true })` scans the collection first. If it finds two documents that already share a key, the build **fails** and no index is created; you have to find and repair the duplicates before retrying. Very old MongoDB versions had a `dropDups` option that deleted the offending documents automatically — it was removed long ago, and proposing it is a dated answer. ## Missing fields and null A field that is absent from a document is still indexed, as the value `null`. Under a plain unique index that means at most **one** document in the collection may omit the field, which is almost never what you want for an optional attribute such as an optional `taxId`. The modern fix is a partial index whose `partialFilterExpression` restricts the index to documents that actually carry the field, so uniqueness applies only inside that subset. ## Arrays and multikey uniqueness If the indexed field holds an array, the index is multikey: one entry per array element. A unique multikey index therefore forbids the *same element value* from appearing in two different documents, and also forbids a single document from containing the same value twice in that array. That is occasionally exactly what you want (globally unique tags) and more often a surprise. ## Sharded collections On a sharded collection, a unique index can only be enforced if its key is prefixed by the shard key. The reason is mechanical: each shard maintains its own indexes, so nothing on shard A can see the keys stored on shard B. Only when the shard key is a prefix does the routing guarantee that all documents with a given key land on the same shard, which makes a local index check sufficient. Wanting `email` globally unique on a collection sharded by `userId` is one of the most common design walls teams hit; the usual answers are to choose the shard key so it prefixes the constraint, or to keep a small separate collection whose `_id` is the value that must be unique and write to it first. ## Practical notes The `_id` index exists on every collection, is unique, and cannot be dropped, hidden, or made partial. Application-level "check then insert" is not a substitute for a unique index: two concurrent requests can both pass the check. Treat the index as the source of truth and the pre-check as a UX nicety.

  • Your service checks whether an email exists before inserting. Why is a unique index still needed?
    Because the check and the insert are two separate operations. Two concurrent requests can both find no existing row and both insert, producing a duplicate. Only the index makes the constraint atomic with the write. The pre-check is useful for a friendly error message, but the index is what actually guarantees the invariant, and the code should still catch error 11000.
  • Why can a unique index on email not be enforced on a collection sharded by userId?
    Each shard maintains only its own indexes and cannot see keys held by other shards. Uniqueness can therefore only be checked locally, which is sound only when all documents sharing the constrained value are guaranteed to be on one shard — that is, when the shard key is a prefix of the unique index key. Otherwise MongoDB refuses to create the unique index.
  • What happens if you create a unique index on a field where many documents simply do not have that field?
    The absent field is indexed as null, so the first such document succeeds and every subsequent one fails with a duplicate key error. The fix is a partial index with a partialFilterExpression such as { taxId: { $exists: true } }, so only documents that actually carry the field are indexed and constrained.

saying these in an interview costs you the question

  • Says the duplicate is silently ignored or the old document overwritten
  • Thinks a compound unique index makes each field unique separately
  • Believes an application-level existence check replaces the index
  • Assumes documents missing the field are skipped by a plain unique index
  • Suggests dropDups to clean duplicates during the build

context