skip to content

What rules does MongoDB enforce on the _id field of every document?

level: juniorimportance: must knowfreq 78%

answer

  1. one field is mandatory in every document
  2. the database builds one index for you unasked
  3. you cannot drop that index
  4. try to $set it and the write is rejected
  5. one BSON type is forbidden there

basics

~20 s

Every document must have an _id. MongoDB creates a unique index on it that cannot be dropped, the value is immutable after insert, and it may hold any BSON type except an array. If you omit it, an ObjectId is generated.

solid answer

~50 s

`_id` is the primary key of a MongoDB document and it is mandatory: if you insert without one, the driver (or failing that the server) generates an `ObjectId`. Every collection gets a unique index named `_id_` automatically, and that index cannot be dropped — a duplicate insert fails with a duplicate-key error (E11000). The value is **immutable**: no update operator can change it, so "changing an id" means deleting the document and inserting it again under the new value. The type is up to you — a string, a number, a `Date`, a binary UUID, even an embedded document acting as a compound key — with the notable exception that `_id` may not be an array. Using a meaningful natural key as `_id` saves you a second unique index; using an ObjectId keeps ids opaque and generatable client-side.

code

javascript · 6 lines
javascript
db.settings.insertOne({ _id: "rate-limit", perMinute: 120 })
db.settings.insertOne({ _id: "rate-limit", perMinute: 60 })
// E11000 duplicate key error on index _id_

db.settings.updateOne({ _id: "rate-limit" }, { $set: { _id: "rl" } })
// rejected: the _id field is immutable

go deeper

for a junior

Know that _id is required, unique, auto-generated as an ObjectId when omitted, and cannot be changed later. Being able to state those four facts cleanly is what this question checks.

for a middle

Explain the mechanics: the automatic id unique index that cannot be dropped, the duplicate-key error it raises, and why immutability forces a delete-and-reinsert to change an identifier.

for a senior

Demonstrate judgment about identifier choice — natural key in _id as a database-enforced idempotency guard versus an opaque ObjectId — and call out the anti-pattern of carrying both an ObjectId and a separately indexed natural key.

for a principal

Own identity across services: which system mints ids, whether they must be meaningful or opaque, the migration cost of an id that turns out to be mutable, and how identifier shape affects storage and index footprint at scale.

## The one required field Every MongoDB document has an `_id` field. It is the document's primary key, and it is not optional: if your insert does not contain one, the driver adds an `ObjectId` client-side before the write leaves the process, and if the field is still missing when the server receives it, the server supplies one. There is no way to store a document without `_id`. ## The automatic unique index When a collection is created, MongoDB builds a unique index on `_id`, conventionally named `_id_`. You cannot drop it and you cannot make it non-unique. Two consequences follow: - **Duplicate inserts fail.** Inserting a second document with an `_id` that already exists raises a duplicate-key error (error code 11000). This is frequently used deliberately as an idempotency guard: derive `_id` from a natural business key, and a retried write cannot create a second copy. - **Lookups by `_id` are always indexed.** `findOne({ _id: x })` never degrades into a collection scan, regardless of what other indexes exist. ## Immutability The `_id` value cannot be modified once the document exists. `$set` on `_id` is rejected, and so is any update that would change it. To "rename" a document's identity you read it, insert a copy under the new `_id`, and delete the original — and you own the consistency of that two-step yourself, or wrap it in a transaction. This immutability is what lets the rest of the system treat `_id` as a stable handle: references stored in other documents, cached keys, and change-stream consumers all rely on it never moving. ## Which types are allowed `_id` may hold values of essentially any BSON type — the important exception is that it **may not be an array**, since an array `_id` would make the primary key ambiguous. Common choices: - **`ObjectId`** — the default. 12 bytes, generated without coordination, carries a creation timestamp. - **A natural key** such as a SKU, an email address, or a slug. Saves a separate unique index and makes lookups by that key trivially fast, at the price of an identifier that can change meaning over time (and it cannot be updated in place). - **A UUID**, which belongs in a BSON binary value of the UUID subtype rather than a 36-character string: 16 bytes instead of 36, and correctly typed comparisons. - **An embedded document** used as a compound key, e.g. `{ _id: { tenant: "acme", sku: "A-1" } }`. Be aware that BSON document comparison is **field-order sensitive**, so `{ tenant, sku }` and `{ sku, tenant }` are different `_id` values — always build the sub-document in a fixed field order. - **A number or a date**, which is fine, though monotonically increasing numeric ids give up the coordination-free generation an ObjectId provides. ## Choosing between an ObjectId and a natural key Ask what the identifier is for. If other documents and other systems reference this record, an opaque, immutable, coordination-free `ObjectId` is the safe default — nothing about the business can invalidate it. If the record *is* the natural key (a configuration entry keyed by name, a per-day rollup keyed by date, an idempotency record keyed by a request id), putting that key in `_id` removes an index, removes a lookup, and turns duplicate suppression into a database-enforced guarantee. A mistake worth avoiding is carrying **both**: an ObjectId `_id` plus a separate unique-indexed natural key that everything actually queries by. That is two indexes and two identities for one document; if nothing references the ObjectId, promote the natural key into `_id`. ## Sorting by _id Because the `_id` index is sorted and an ObjectId begins with a timestamp, `sort({ _id: 1 })` approximates insertion order and `_id` range predicates approximate time ranges. That is a property of ObjectId values, not of `_id` itself — with string or natural-key ids, sorting by `_id` sorts by that key and tells you nothing about time.

  • How would you store a UUID as _id without wasting space?
    Put it in a BSON binary value of the UUID subtype rather than as a hex string. That is 16 bytes instead of a 36-character string, keeps comparisons type-correct, and lets drivers hand it back as a native UUID object. Storing the string form roughly doubles both document and index size for the same identifier.
  • Can _id be an embedded document, and what should you watch out for?
    Yes — an embedded document is a legitimate compound key, for example { tenant, period }. The trap is that BSON document comparison is field-order sensitive, so the same fields written in a different order are a different _id and will not match. Always construct the sub-document in one fixed field order.
  • What happens if you insert two documents with the same _id?
    The second insert fails with a duplicate-key error (code 11000), because the automatic _id index is unique and cannot be relaxed. Teams use this deliberately: derive _id from a business key so that a retried or duplicated request cannot create a second document.

saying these in an interview costs you the question

  • Thinks _id is optional and can be left out entirely
  • Says you can update _id with $set like any other field
  • Believes you must create the unique index on _id yourself
  • Claims _id must always be an ObjectId
  • Stores a UUID or ObjectId as a string without noticing the cost

context