skip to content

MongoDB

9 roadmaps112 questionsupdated

The default document database: BSON documents in collections, a query and aggregation API, indexes, replica sets and sharding. Interviewers ask because it is where most candidates' NoSQL experience actually lives, and its trade-offs are well defined enough to test.

on this pageshow

guide

overview

~1 min

MongoDB is a document database: it stores BSON documents in collections, queries them through a filter-and-operator API and an aggregation pipeline, keeps copies in replica sets and spreads data across shards. Interviewers use it because it is where most candidates have actually met NoSQL, and because each of its trade-offs is concrete enough to test: what a document boundary buys, what an index can and cannot serve, what an acknowledged write survives. A strong MongoDB answer names the access pattern first and the feature second. The hub follows the life of the data. [Document Model & BSON](/topics/db-mongodb-data-model) is where design starts: types, `_id`, validation and the embed-or-reference choice. [CRUD & Query Operators](/topics/db-mongodb-crud-query) is the everyday API, including sessions and multi-document transactions. [Aggregation Pipeline](/topics/db-mongodb-aggregation) covers everything past a simple lookup, from grouping to `$lookup` joins. [Indexing & Performance](/topics/db-mongodb-indexing) is where reads become fast or slow, and where `explain()` settles the argument. [Replica Sets](/topics/db-mongodb-replication) decide availability and what "the write succeeded" means, and [Sharding & Scalability](/topics/db-mongodb-sharding) decides how the system grows past one machine. Junior rounds stay close to the API: operators, projections, what a query matches and whether it used an index. Senior and principal rounds turn into design and incident conversations: a schema that grows without bound, a compound index in the wrong order, a shard key that funnels every insert to one shard, a failover that discarded writes. Start with the document model and the query API, then indexing. Replication and sharding make sense once you can say which queries matter and how they are served.

primer

### The document is the unit of design A document holds a whole entity, nested objects and arrays included, and it is what MongoDB reads and writes as one atomic piece. Modelling is therefore driven by the questions the application asks, not by normal forms: data read together tends to live together, and data that grows without limit or is shared by many owners tends to live apart. Deciding when to embed and when to reference has no direct relational equivalent, and interviewers use it to separate people who have modelled in MongoDB from people who have only queried it. ### Flexible schema is a decision, not an absence Nothing requires two documents in a collection to share a shape, so the schema lives in application code, in validators, or nowhere. Strong answers say where it lives and what enforces it. Types matter more than they look: BSON distinguishes integer, double and decimal, and a number stored as a string never matches a numeric filter, so type drift silently shrinks query results. ### Arrays are first-class Queries, updates and indexes all treat arrays specially. A condition on an array field is tested against its elements, update operators can target one element or many, and an index on an array holds an entry per element. Most "why did this match" and "why is this index huge" questions come back to that. ### Indexes decide performance A query without a usable index reads the whole collection. A compound index serves only queries that use its leading fields, and the order of its fields decides whether it can also filter and sort in one pass. `explain()` is how you prove any of it; the ratio of work done to documents returned is the number interviewers want you to read. ### One document, one atomic write A write to one document is all-or-nothing however much of it changes. Multi-document transactions exist, but they cost latency and have limits, so the usual first answer is a model where the invariant fits inside one document. ### Durability and freshness are chosen per operation Write concern says how many members must hold a write before it is acknowledged; read concern and read preference say which data a read may see and which member serves it. The defaults are reasonable, but an interviewer will ask what each weaker or stronger setting risks during a failover. ### The shard key is the most permanent choice In a sharded cluster, the shard key decides where each document lives, which shards a query must visit, and whether writes spread or pile up. Changing it later is possible only as a heavyweight operation, so it is defended at design time from the query patterns, not the data size.

Document
A BSON record stored in a collection, holding fields, nested documents and arrays. The unit MongoDB reads, writes and replicates atomically.
BSON
The binary, typed serialization MongoDB stores and sends. It extends JSON with types such as ObjectId, date, 64-bit integer and Decimal128.
ObjectId
The default 12-byte _id value generated when a document is inserted without one; it embeds a creation timestamp.
Embedding and referencing
The two ways to relate data: nest the related data inside the parent document, or store it separately and link it by an identifier.
Schema validator
A rule attached to a collection, usually written in $jsonSchema, that inserts and updates are checked against before they are accepted.
Aggregation pipeline
An ordered list of stages, such as $match, $group and $lookup, that transforms a stream of documents on the server.
Compound index
An index over several fields in a declared order. Queries must use its leading fields, and field order decides filtering and sorting power.
Multikey index
An index on a field that holds arrays, storing one index entry per array element rather than one per document.
Covered query
A query answered entirely from index keys, without loading any document, because the filter and returned fields all live in one index.
Replica set
A group of mongod processes holding the same data: one primary accepts writes, secondaries copy them and can take over after an election.
Oplog
The capped collection where the primary records every change in replayable form; secondaries tail it to stay in sync.
Write concern
The acknowledgement level a write requests, from the primary alone up to a majority of members, which decides whether it can be rolled back.
Read preference
The rule that decides which replica set members may serve a read, trading freshness against latency and load on the primary.
Shard key
The indexed field or fields that decide which shard owns each document in a sharded collection and which shards a query must contact.
Chunk
A contiguous range of shard-key values owned by one shard; the balancer migrates these ranges to even out data between shards.
mongos
The query router of a sharded cluster. Applications connect to it, and it forwards each operation to the shards that can hold matches.

Follow a write and then a read through a sharded deployment. The application connects to a `mongos`, which reads the shard key from the document, consults routing metadata cached from the config servers, and forwards the insert to the shard owning that key range. Each shard is itself a replica set: its primary checks the document against the collection's validator, if one is set, applies the write, records it in the oplog, and secondaries copy it. The write concern decides when the client hears back. A read takes the reverse path. If its filter carries the shard key, `mongos` contacts one shard, or a few; otherwise it asks all of them and merges the results. On each shard the query planner picks an index, and read preference decides whether a primary or a secondary answers. An aggregation runs the same way with more steps: its leading stages are pushed to the shards and use their indexes, and later stages may run on a merging node. That is why the sections depend on each other. The data model fixes which queries exist; the queries fix which indexes are worth their write cost; the hottest query shape should also carry the shard key. One design decision often serves all three: ```javascript // ESR order: customerId (equality), createdAt (sort), total (range) db.orders.createIndex({ customerId: 1, createdAt: -1, total: 1 }) // the same leading field is the shard key, so this index also supports it sh.shardCollection("shop.orders", { customerId: 1 }) // targeted to one shard, index-backed, sorted without a separate sort step db.orders.find({ customerId: 42, total: { $gte: 100 } }) .sort({ createdAt: -1 }) ``` Replication sits underneath all of it. Durability depends on write concern, a failover can roll back writes that never reached a majority, and reading from secondaries trades freshness for spare capacity. Transactions ride on the same machinery: they need a replica set, and across shards they add a coordinator and extra round trips.

  1. Document Model & BSON →

    Documents, BSON types and embed-or-reference: the model every query, index and shard key is designed around.

  2. CRUD & Query Operators →

    The everyday read and write API, its operators, and where single-document atomicity ends and transactions begin.

  3. Indexing & Performance →

    How queries get fast: index types, compound field order, and reading explain() output as evidence.

  4. Aggregation Pipeline →

    Grouping, joining and reshaping on the server, and why stage order decides whether a pipeline uses indexes.

  5. Replica Sets →

    What an acknowledged write survives, how failover works, and which member a read is allowed to hit.

  6. Sharding & Scalability →

    Horizontal scale, which only makes sense once queries, indexes and replication are clear.

  • Treating flexible schema as no schema: without a validator or a single model in code, type drift means a query on a number silently misses documents stored as strings.

  • Embedding an array that grows without bound, such as every comment or event, until documents get slow to rewrite and approach the size limit.

  • Porting a normalized relational schema one table per collection, then rebuilding every read with $lookup joins that the document model was meant to avoid.

  • Paging deep result sets with skip(); the server still walks every skipped entry, so page 10,000 costs far more than page 1.

  • Reaching for multi-document transactions by default, when a model that keeps the invariant inside one document is cheaper and needs no retry logic.

  • Assuming an acknowledged write is safe after failover when it was only acknowledged by the primary — see rollback after failover.

  • Reading from secondaries to scale reads without accepting that they lag, so a user may not see the write they just made.

  • Choosing a monotonically increasing field such as a timestamp or ObjectId as a ranged shard key, which sends every insert to the same shard.

  • Sharding before the query shapes are known; a filter without the shard key is broadcast to every shard, and adding shards makes that worse, not better.

This guide assumes a current MongoDB release, 7.0 or later, running the WiredTiger storage engine. Interviewers still ask about the releases where the rules changed, because older clusters and old blog posts describe different behaviour: - **4.0** added multi-document transactions on replica sets, and **4.2** extended them to sharded clusters. - **4.4** made it possible to refine a shard key by adding suffix fields. - **5.0** introduced time series collections and live resharding, so a bad shard key is no longer fixable only by dumping and reloading. It also made `w: "majority"` the default write concern for most deployments. - **6.x** reworked balancing around data size per shard rather than chunk counts, which is why newer documentation talks about ranges more than chunks. When an answer depends on these — whether transactions work across shards, whether a shard key can change, what the default write concern is — say which version you mean. Answers written against pre-4.0 MongoDB, such as "MongoDB has no transactions", are the most common outdated claims interviewers hear.

Interviewers expect you to place MongoDB, not just use it. Its closest competitors are other document and key-value stores: Couchbase, Amazon DynamoDB and Azure Cosmos DB, the last of which also offers a MongoDB-compatible API. The comparison that comes up most is against a relational database with JSON columns, typically PostgreSQL with `jsonb`: MongoDB favours data that is read and written as whole aggregates, with scaling built into the product, while a relational store favours many-to-many relationships, joins and constraints across tables. Wide-column stores such as Apache Cassandra fit write-heavy, query-known workloads at larger scale with fewer query features. Around it sit the layers you will be asked about in application interviews. Object-document mappers such as Mongoose for Node.js and Spring Data MongoDB for Java add schemas and repositories on the client side. Change streams feed events to messaging systems such as Kafka. MongoDB Atlas, the managed service from MongoDB, Inc., removes server operations but not modelling, indexing or shard key design; those questions follow you to any hosting choice.

explore

report an issue with this guide →

questions

112 · 6 sections

What rules does MongoDB enforce on the _id field of every document?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Every document must have an _id. MongoDB creates a unique index on it that cannot be dropped, the value is immutable after insert, and it may hold any BSON type except an array. If you omit it, an ObjectId is generated.

open as a page

How do you make a MongoDB collection reject documents that lack required fields?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Attach a validator to the collection — normally a $jsonSchema object listing required fields and their bsonType — using createCollection or collMod. MongoDB then checks every insert and update against it and rejects writes that break the rules.

open as a page

What are the 12 bytes of a MongoDB ObjectId, and what can you infer from one?

level: middleimportance: must knowfreq 75%
basics
~20 s

An ObjectId is 12 bytes: a 4-byte Unix timestamp in seconds, a 5-byte value random per process, and a 3-byte incrementing counter. From one you can read its creation time to the second and infer rough insertion order.

open as a page

What is MongoDB's bucket pattern, and when does bucketing sensor readings beat one document per reading?

level: middleimportance: must knowfreq 62%
basics
~20 s

The bucket pattern stores many small, time-ordered measurements in one document — an hour of readings for one sensor, say — instead of one document per reading. It cuts document count, index entries and per-document overhead for time-series workloads.

open as a page

What is MongoDB's computed pattern, and when does precomputing a total beat aggregating at read time?

level: middleimportance: must knowfreq 52%
basics
~20 s

The computed pattern stores the result of a calculation — a count, sum, average or rollup — in the document and maintains it on write, so reads fetch a value instead of recomputing it. It pays off when reads far outnumber writes.

open as a page

In a MongoDB find() projection, can you mix included and excluded fields?

level: juniorimportance: must knowfreq 72%
basics
~10 s

No. A find() projection is either an inclusion list or an exclusion list, and mixing the two raises an error. The single exception is _id, which may be excluded inside an otherwise inclusive projection.

open as a page

In MongoDB, what does the find filter { tags: "red" } match when tags holds an array?

level: juniorimportance: must knowfreq 78%
basics
~20 s

It matches documents where tags is exactly the string "red" and documents where tags is an array containing "red" as an element. MongoDB applies the equality predicate to the field's value and to each array element.

open as a page

In MongoDB, what is the difference between updateOne with $set and replaceOne?

level: juniorimportance: must knowfreq 78%
basics
~20 s

updateOne with $set changes only the named fields and leaves every other field intact. replaceOne swaps the whole document for the one you pass, so any field the replacement omits disappears. _id cannot be changed either way.

open as a page

Why can $elemMatch return different results from listing the same two conditions on an array field?

level: middleimportance: must knowfreq 72%
basics
~10 s

Two conditions written side by side may be satisfied by different array elements; $elemMatch requires one single element to satisfy all of them at once. On a non-array field the two forms are equivalent.

open as a page

Which MongoDB write operations are atomic without opening an explicit transaction?

level: middleimportance: must knowfreq 78%
basics
~20 s

Every write that touches a single document is atomic in MongoDB, no matter how many fields, embedded documents or array elements it changes. Atomicity stops at the document boundary: a write spanning several documents needs an explicit transaction to be all-or-nothing.

open as a page

What does _id define in a $group stage, and how do you total a field per group?

level: juniorimportance: must knowfreq 85%
basics
~10 s

In $group, _id is the grouping-key expression: one output document per distinct _id value. Setting it to null puts everything in one group. Totals come from accumulators, for example revenue: { $sum: "$amount" }.

open as a page

What does MongoDB's $lookup equality form return for each input document?

level: juniorimportance: must knowfreq 80%
basics
~20 s

$lookup is a left outer join: for each input document it adds an array field, named by as, holding every document from the from collection whose foreignField equals the input's localField. No match yields an empty array, never a dropped document.

open as a page

What does $unwind do to a document whose array field is empty or missing?

level: middleimportance: must knowfreq 72%
basics
~20 s

$unwind emits one document per array element, copying the other fields. A document whose array is empty, missing or null produces no output at all — it is dropped — unless you use the object form with preserveNullAndEmptyArrays: true, which passes it through once.

open as a page

How do you filter or reshape the joined side of a $lookup using let and a sub-pipeline?

level: middleimportance: must knowfreq 62%
basics
~20 s

Use the pipeline form: let binds fields of the input document to variables, and the sub-pipeline runs against the foreign collection with those variables available as $$name. Compare them inside a $match with $expr, then $project or $limit to trim what comes back.

open as a page

Why should $match come early in an aggregation pipeline, and when does MongoDB reorder it for you?

level: middleimportance: must knowfreq 76%
basics
~20 s

Only a leading $match becomes an indexed query on the collection, so filtering first shrinks the stream before expensive stages run. MongoDB's optimizer moves $match ahead of $project, $addFields or $sort only when the filter uses no field those stages create.

open as a page

In MongoDB explain() output, what is the difference between a COLLSCAN and an IXSCAN stage?

level: juniorimportance: must knowfreq 78%
basics
~20 s

COLLSCAN means the server read every document in the collection; IXSCAN means it walked index keys and touched only matching entries. A COLLSCAN on a filtered query over a large collection signals a missing index.

open as a page

In MongoDB, what does a unique index enforce, and what happens when a write violates it?

level: juniorimportance: must knowfreq 70%
basics
~20 s

A unique index rejects any write that would create a second document with the same indexed value, returning an E11000 duplicate key error. On a compound unique index the whole key combination must be unique, not each field on its own.

open as a page

What is the ESR rule for ordering the fields of a MongoDB compound index?

level: middleimportance: must knowfreq 68%
basics
~20 s

ESR orders compound-index keys as Equality first, then Sort fields, then Range fields. Equality pins the scan to a narrow contiguous key range, the sort fields then arrive already ordered, and range predicates go last because they destroy ordering for the keys after them.

open as a page

What do totalKeysExamined, totalDocsExamined and nReturned reveal about a MongoDB query?

level: middleimportance: must knowfreq 72%
basics
~20 s

They are the work-versus-result ratio of a query. nReturned is what the client got; totalKeysExamined is index entries read; totalDocsExamined is documents loaded. Ratios near 1:1:1 mean a well-matched index; large gaps mean wasted work.

open as a page

How does a MongoDB TTL index with expireAfterSeconds actually delete expired documents?

level: middleimportance: must knowfreq 66%
basics
~20 s

A background thread wakes roughly every 60 seconds, finds documents whose indexed date is older than expireAfterSeconds, and deletes them with ordinary delete operations. Expiry is therefore approximate, not instant, and only the primary performs the deletes.

open as a page

What roles do the primary, secondary, and arbiter play in a MongoDB replica set?

level: juniorimportance: must knowfreq 72%
basics
~20 s

The primary is the only member that accepts writes and records them in its oplog. Secondaries hold full copies of the data, replay that oplog, and can be elected primary. An arbiter stores no data and only votes.

open as a page

What does a MongoDB write concern of w: "majority" actually guarantee about a write?

level: middleimportance: must knowfreq 76%
basics
~20 s

A w: "majority" acknowledgment means the write reached a majority of the replica set's voting, data-bearing members, so it survives an election and cannot be rolled back. It does not mean every member has it.

open as a page

Why does a MongoDB replica set with four voting members tolerate no more failures than one with three?

level: middleimportance: must knowfreq 68%
basics
~20 s

Electing a primary needs a strict majority of the configured voting members. Three members need two votes, four need three, so both survive exactly one loss. The fourth vote adds cost and tie risk without adding fault tolerance.

open as a page

Why does MongoDB rewrite operations into idempotent form before writing them to the oplog?

level: middleimportance: must knowfreq 62%
basics
~20 s

So an entry can be replayed any number of times with the same result. Secondaries and recovering members may re-apply the tail of the oplog after a restart or a sync-source switch, and a relative operation replayed twice would corrupt the data.

open as a page

Which MongoDB settings must you combine so a user always reads their own write from a secondary?

level: seniorimportance: must knowfreq 55%
basics
~20 s

Run both operations inside one causally consistent client session, write with w: "majority" and read with readConcern: "majority". The session carries a cluster timestamp so the secondary waits until it has applied that write before answering.

open as a page

What is a chunk in a sharded MongoDB collection, and what does the balancer do with chunks?

level: middleimportance: must knowfreq 70%
basics
~20 s

A chunk (range) is a contiguous span of shard-key values owned by exactly one shard and recorded in config-server metadata. The balancer migrates ranges from the shard holding the most data for a collection to the shard holding the least.

open as a page

Which query filters let mongos target one shard, and which force a broadcast to every shard?

level: middleimportance: must knowfreq 72%
basics
~20 s

A filter containing the shard key — or a leading prefix of a compound shard key — lets mongos map values to chunks and contact only the owning shards. Any other filter, including one on _id when _id is not the shard key, is broadcast to all shards.

open as a page

What makes a good MongoDB shard key, and how do cardinality, frequency and monotonicity affect it?

level: middleimportance: must knowfreq 78%
basics
~10 s

A good MongoDB shard key has many distinct values, no single value that dominates the data, values that do not grow monotonically, and it appears in common query filters so requests reach one shard.

open as a page

During a MongoDB chunk migration, what happens on the donor and recipient shards, and how does live traffic feel it?

level: seniorimportance: must knowfreq 55%
basics
~20 s

The recipient clones the range's documents and then catches up on changes; the donor enters a short critical section that blocks writes to that range while the config metadata is committed; afterwards the donor deletes the moved documents in the background as orphans.

open as a page

After sharding, p99 read latency got worse and every shard is busy — how do you confirm broadcast queries are the cause?

level: seniorimportance: must knowfreq 58%
basics
~20 s

Run explain() through mongos on the hot query shapes: a SHARD_MERGE stage with every shard named in the winning plan means scatter-gather. Compare per-shard keysExamined against nReturned, then fix the filter to carry the shard key rather than adding shards.

open as a page