skip to content

Document Model & BSON

How documents are typed and shaped, and the embed-versus-reference decision that drives everything downstream. Interviewers ask because document modelling is the skill that does not transfer directly from relational design.

part ofMongoDBoverview, primer and where to startread it →
on this pageshow

questions

17

What rules does MongoDB enforce on the _id field of every document?

level: juniorimportance: must knowfreq 78%

answer

  1. one field is mandatory in every document
  2. the database builds one index for you unasked
  3. you cannot drop that index
  4. try to $set it and the write is rejected
  5. one BSON type is forbidden there

basics

~20 s

Every document must have an _id. MongoDB creates a unique index on it that cannot be dropped, the value is immutable after insert, and it may hold any BSON type except an array. If you omit it, an ObjectId is generated.

solid answer

~50 s

`_id` is the primary key of a MongoDB document and it is mandatory: if you insert without one, the driver (or failing that the server) generates an `ObjectId`. Every collection gets a unique index named `_id_` automatically, and that index cannot be dropped — a duplicate insert fails with a duplicate-key error (E11000). The value is **immutable**: no update operator can change it, so "changing an id" means deleting the document and inserting it again under the new value. The type is up to you — a string, a number, a `Date`, a binary UUID, even an embedded document acting as a compound key — with the notable exception that `_id` may not be an array. Using a meaningful natural key as `_id` saves you a second unique index; using an ObjectId keeps ids opaque and generatable client-side.

code

javascript · 6 lines
javascript
db.settings.insertOne({ _id: "rate-limit", perMinute: 120 })
db.settings.insertOne({ _id: "rate-limit", perMinute: 60 })
// E11000 duplicate key error on index _id_

db.settings.updateOne({ _id: "rate-limit" }, { $set: { _id: "rl" } })
// rejected: the _id field is immutable

go deeper

for a junior

Know that _id is required, unique, auto-generated as an ObjectId when omitted, and cannot be changed later. Being able to state those four facts cleanly is what this question checks.

for a middle

Explain the mechanics: the automatic id unique index that cannot be dropped, the duplicate-key error it raises, and why immutability forces a delete-and-reinsert to change an identifier.

for a senior

Demonstrate judgment about identifier choice — natural key in _id as a database-enforced idempotency guard versus an opaque ObjectId — and call out the anti-pattern of carrying both an ObjectId and a separately indexed natural key.

for a principal

Own identity across services: which system mints ids, whether they must be meaningful or opaque, the migration cost of an id that turns out to be mutable, and how identifier shape affects storage and index footprint at scale.

## The one required field Every MongoDB document has an `_id` field. It is the document's primary key, and it is not optional: if your insert does not contain one, the driver adds an `ObjectId` client-side before the write leaves the process, and if the field is still missing when the server receives it, the server supplies one. There is no way to store a document without `_id`. ## The automatic unique index When a collection is created, MongoDB builds a unique index on `_id`, conventionally named `_id_`. You cannot drop it and you cannot make it non-unique. Two consequences follow: - **Duplicate inserts fail.** Inserting a second document with an `_id` that already exists raises a duplicate-key error (error code 11000). This is frequently used deliberately as an idempotency guard: derive `_id` from a natural business key, and a retried write cannot create a second copy. - **Lookups by `_id` are always indexed.** `findOne({ _id: x })` never degrades into a collection scan, regardless of what other indexes exist. ## Immutability The `_id` value cannot be modified once the document exists. `$set` on `_id` is rejected, and so is any update that would change it. To "rename" a document's identity you read it, insert a copy under the new `_id`, and delete the original — and you own the consistency of that two-step yourself, or wrap it in a transaction. This immutability is what lets the rest of the system treat `_id` as a stable handle: references stored in other documents, cached keys, and change-stream consumers all rely on it never moving. ## Which types are allowed `_id` may hold values of essentially any BSON type — the important exception is that it **may not be an array**, since an array `_id` would make the primary key ambiguous. Common choices: - **`ObjectId`** — the default. 12 bytes, generated without coordination, carries a creation timestamp. - **A natural key** such as a SKU, an email address, or a slug. Saves a separate unique index and makes lookups by that key trivially fast, at the price of an identifier that can change meaning over time (and it cannot be updated in place). - **A UUID**, which belongs in a BSON binary value of the UUID subtype rather than a 36-character string: 16 bytes instead of 36, and correctly typed comparisons. - **An embedded document** used as a compound key, e.g. `{ _id: { tenant: "acme", sku: "A-1" } }`. Be aware that BSON document comparison is **field-order sensitive**, so `{ tenant, sku }` and `{ sku, tenant }` are different `_id` values — always build the sub-document in a fixed field order. - **A number or a date**, which is fine, though monotonically increasing numeric ids give up the coordination-free generation an ObjectId provides. ## Choosing between an ObjectId and a natural key Ask what the identifier is for. If other documents and other systems reference this record, an opaque, immutable, coordination-free `ObjectId` is the safe default — nothing about the business can invalidate it. If the record *is* the natural key (a configuration entry keyed by name, a per-day rollup keyed by date, an idempotency record keyed by a request id), putting that key in `_id` removes an index, removes a lookup, and turns duplicate suppression into a database-enforced guarantee. A mistake worth avoiding is carrying **both**: an ObjectId `_id` plus a separate unique-indexed natural key that everything actually queries by. That is two indexes and two identities for one document; if nothing references the ObjectId, promote the natural key into `_id`. ## Sorting by _id Because the `_id` index is sorted and an ObjectId begins with a timestamp, `sort({ _id: 1 })` approximates insertion order and `_id` range predicates approximate time ranges. That is a property of ObjectId values, not of `_id` itself — with string or natural-key ids, sorting by `_id` sorts by that key and tells you nothing about time.

  • How would you store a UUID as _id without wasting space?
    Put it in a BSON binary value of the UUID subtype rather than as a hex string. That is 16 bytes instead of a 36-character string, keeps comparisons type-correct, and lets drivers hand it back as a native UUID object. Storing the string form roughly doubles both document and index size for the same identifier.
  • Can _id be an embedded document, and what should you watch out for?
    Yes — an embedded document is a legitimate compound key, for example { tenant, period }. The trap is that BSON document comparison is field-order sensitive, so the same fields written in a different order are a different _id and will not match. Always construct the sub-document in one fixed field order.
  • What happens if you insert two documents with the same _id?
    The second insert fails with a duplicate-key error (code 11000), because the automatic _id index is unique and cannot be relaxed. Teams use this deliberately: derive _id from a business key so that a retried or duplicated request cannot create a second document.

saying these in an interview costs you the question

  • Thinks _id is optional and can be left out entirely
  • Says you can update _id with $set like any other field
  • Believes you must create the unique index on _id yourself
  • Claims _id must always be an ObjectId
  • Stores a UUID or ObjectId as a string without noticing the cost

context

open as a page

How do you make a MongoDB collection reject documents that lack required fields?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Attach a validator to the collection — normally a $jsonSchema object listing required fields and their bsonType — using createCollection or collMod. MongoDB then checks every insert and update against it and rejects writes that break the rules.

open as a page

What are the 12 bytes of a MongoDB ObjectId, and what can you infer from one?

level: middleimportance: must knowfreq 75%

basics

~20 s

An ObjectId is 12 bytes: a 4-byte Unix timestamp in seconds, a 5-byte value random per process, and a 3-byte incrementing counter. From one you can read its creation time to the second and infer rough insertion order.

open as a page

What is MongoDB's bucket pattern, and when does bucketing sensor readings beat one document per reading?

level: middleimportance: must knowfreq 62%

basics

~20 s

The bucket pattern stores many small, time-ordered measurements in one document — an hour of readings for one sensor, say — instead of one document per reading. It cuts document count, index entries and per-document overhead for time-series workloads.

open as a page

What is MongoDB's computed pattern, and when does precomputing a total beat aggregating at read time?

level: middleimportance: must knowfreq 52%

basics

~20 s

The computed pattern stores the result of a calculation — a count, sum, average or rollup — in the document and maintains it on write, so reads fetch a value instead of recomputing it. It pays off when reads far outnumber writes.

open as a page

What is the difference between validationLevel strict and moderate in MongoDB?

level: middleimportance: must knowfreq 52%

basics

~20 s

validationLevel decides which writes are checked. strict, the default, validates every insert and update. moderate validates inserts and updates to documents that already satisfy the validator, but skips updates to documents that currently fail it. off disables validation.

open as a page

In MongoDB's subset pattern, what stays embedded in the main document and how do you cap it?

level: juniorimportance: should knowfreq 45%

basics

~20 s

The subset pattern embeds only the small, hot slice of a large related set — the ten newest reviews, say — and keeps the complete set in its own collection. A $push with the $slice modifier keeps the embedded slice capped.

open as a page

How does BSON store a Date, and how does it differ from the BSON Timestamp type?

level: middleimportance: should knowfreq 42%

basics

~20 s

A BSON Date is a signed 64-bit count of milliseconds since the Unix epoch, in UTC, with no timezone stored. BSON Timestamp is a separate internal type used for replication — application data should always use Date.

open as a page

Which numeric types does BSON provide, and when is double the wrong choice?

level: middleimportance: should knowfreq 58%

basics

~20 s

BSON has int (32-bit), long (64-bit), double (64-bit binary float) and decimal (Decimal128). Double is wrong for money and for integers beyond 2^53, because binary floating point cannot represent decimal fractions or large integers exactly.

open as a page

How does MongoDB compare values of different BSON types when matching and sorting?

level: middleimportance: should knowfreq 48%

basics

~20 s

MongoDB never coerces types: a number never equals a string. Across types it uses a fixed BSON order — MinKey, null, numbers, string, object, array, binData, ObjectId, boolean, date, timestamp, regex, MaxKey — with all numeric types compared by value.

open as a page

What does MongoDB's attribute pattern do to a document with dozens of optional fields, and why?

level: middleimportance: should knowfreq 40%

basics

~20 s

The attribute pattern turns many rarely-queried, per-type fields into an array of key/value subdocuments such as specs: [{k, v}]. One compound multikey index on specs.k and specs.v then serves searches on any attribute instead of one index per field.

open as a page

What does validationAction warn do when a MongoDB write violates the collection validator?

level: middleimportance: should knowfreq 36%

basics

~10 s

With validationAction set to warn, MongoDB lets the write succeed and records the violation in the mongod log instead of failing it. The default, error, rejects the write and returns a document-failed-validation write error.

open as a page

What breaks as a MongoDB document approaches the 16 MB BSON limit, and how do you avoid it?

level: seniorimportance: should knowfreq 70%

basics

~20 s

MongoDB caps a BSON document at 16 MB. A write that would exceed it fails outright, and performance suffers long before that, because whole documents are read and written. Split unbounded arrays into their own documents and use GridFS for large binaries.

open as a page

For a MongoDB category tree, how do array-of-ancestors and materialized-path documents differ when querying a subtree?

level: seniorimportance: should knowfreq 36%

basics

~20 s

An array of ancestors makes a subtree an equality match on an indexed array — find({ ancestors: "books" }). A materialized path stores the chain as one string and needs an anchored prefix regex. Both rewrite every descendant when a node moves.

open as a page

What kinds of rules can a MongoDB $jsonSchema validator not enforce?

level: seniorimportance: should knowfreq 33%

basics

~20 s

A validator sees only the single document being written, so it cannot enforce uniqueness, cross-collection references, or anything about other documents. It also never supplies defaults or coerces types, and privileged writes can bypass it entirely.

open as a page

How would you add a $jsonSchema validator to a collection that already holds non-conforming documents?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Roll it out in stages: attach the schema with collMod at validationLevel moderate and validationAction warn, measure the offenders by querying with $nor plus $jsonSchema, backfill them in batches while fixing the writers, then flip to strict and error.

open as a page

What is MongoDB's outlier pattern, and how does it stop a few huge documents from shaping the whole schema?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The outlier pattern keeps the common document shape optimal and handles the rare extreme case separately: a flag field marks documents whose data overflows into extra documents, and only flagged documents pay for the second lookup.

open as a page