MongoDB
The default document database: BSON documents in collections, a query and aggregation API, indexes, replica sets and sharding. Interviewers ask because it is where most candidates' NoSQL experience actually lives, and its trade-offs are well defined enough to test.
on this pageshowhide
guide
overview
~1 minMongoDB is a document database: it stores BSON documents in collections, queries them through a filter-and-operator API and an aggregation pipeline, keeps copies in replica sets and spreads data across shards. Interviewers use it because it is where most candidates have actually met NoSQL, and because each of its trade-offs is concrete enough to test: what a document boundary buys, what an index can and cannot serve, what an acknowledged write survives. A strong MongoDB answer names the access pattern first and the feature second. The hub follows the life of the data. [Document Model & BSON](/topics/db-mongodb-data-model) is where design starts: types, `_id`, validation and the embed-or-reference choice. [CRUD & Query Operators](/topics/db-mongodb-crud-query) is the everyday API, including sessions and multi-document transactions. [Aggregation Pipeline](/topics/db-mongodb-aggregation) covers everything past a simple lookup, from grouping to `$lookup` joins. [Indexing & Performance](/topics/db-mongodb-indexing) is where reads become fast or slow, and where `explain()` settles the argument. [Replica Sets](/topics/db-mongodb-replication) decide availability and what "the write succeeded" means, and [Sharding & Scalability](/topics/db-mongodb-sharding) decides how the system grows past one machine. Junior rounds stay close to the API: operators, projections, what a query matches and whether it used an index. Senior and principal rounds turn into design and incident conversations: a schema that grows without bound, a compound index in the wrong order, a shard key that funnels every insert to one shard, a failover that discarded writes. Start with the document model and the query API, then indexing. Replication and sharding make sense once you can say which queries matter and how they are served.
primer
### The document is the unit of design A document holds a whole entity, nested objects and arrays included, and it is what MongoDB reads and writes as one atomic piece. Modelling is therefore driven by the questions the application asks, not by normal forms: data read together tends to live together, and data that grows without limit or is shared by many owners tends to live apart. Deciding when to embed and when to reference has no direct relational equivalent, and interviewers use it to separate people who have modelled in MongoDB from people who have only queried it. ### Flexible schema is a decision, not an absence Nothing requires two documents in a collection to share a shape, so the schema lives in application code, in validators, or nowhere. Strong answers say where it lives and what enforces it. Types matter more than they look: BSON distinguishes integer, double and decimal, and a number stored as a string never matches a numeric filter, so type drift silently shrinks query results. ### Arrays are first-class Queries, updates and indexes all treat arrays specially. A condition on an array field is tested against its elements, update operators can target one element or many, and an index on an array holds an entry per element. Most "why did this match" and "why is this index huge" questions come back to that. ### Indexes decide performance A query without a usable index reads the whole collection. A compound index serves only queries that use its leading fields, and the order of its fields decides whether it can also filter and sort in one pass. `explain()` is how you prove any of it; the ratio of work done to documents returned is the number interviewers want you to read. ### One document, one atomic write A write to one document is all-or-nothing however much of it changes. Multi-document transactions exist, but they cost latency and have limits, so the usual first answer is a model where the invariant fits inside one document. ### Durability and freshness are chosen per operation Write concern says how many members must hold a write before it is acknowledged; read concern and read preference say which data a read may see and which member serves it. The defaults are reasonable, but an interviewer will ask what each weaker or stronger setting risks during a failover. ### The shard key is the most permanent choice In a sharded cluster, the shard key decides where each document lives, which shards a query must visit, and whether writes spread or pile up. Changing it later is possible only as a heavyweight operation, so it is defended at design time from the query patterns, not the data size.
- Document
- A BSON record stored in a collection, holding fields, nested documents and arrays. The unit MongoDB reads, writes and replicates atomically.
- BSON
- The binary, typed serialization MongoDB stores and sends. It extends JSON with types such as ObjectId, date, 64-bit integer and Decimal128.
- ObjectId
- The default 12-byte _id value generated when a document is inserted without one; it embeds a creation timestamp.
- Embedding and referencing
- The two ways to relate data: nest the related data inside the parent document, or store it separately and link it by an identifier.
- Schema validator
- A rule attached to a collection, usually written in $jsonSchema, that inserts and updates are checked against before they are accepted.
- Aggregation pipeline
- An ordered list of stages, such as $match, $group and $lookup, that transforms a stream of documents on the server.
- Compound index
- An index over several fields in a declared order. Queries must use its leading fields, and field order decides filtering and sorting power.
- Multikey index
- An index on a field that holds arrays, storing one index entry per array element rather than one per document.
- Covered query
- A query answered entirely from index keys, without loading any document, because the filter and returned fields all live in one index.
- Replica set
- A group of mongod processes holding the same data: one primary accepts writes, secondaries copy them and can take over after an election.
- Oplog
- The capped collection where the primary records every change in replayable form; secondaries tail it to stay in sync.
- Write concern
- The acknowledgement level a write requests, from the primary alone up to a majority of members, which decides whether it can be rolled back.
- Read preference
- The rule that decides which replica set members may serve a read, trading freshness against latency and load on the primary.
- Shard key
- The indexed field or fields that decide which shard owns each document in a sharded collection and which shards a query must contact.
- Chunk
- A contiguous range of shard-key values owned by one shard; the balancer migrates these ranges to even out data between shards.
- mongos
- The query router of a sharded cluster. Applications connect to it, and it forwards each operation to the shards that can hold matches.
Follow a write and then a read through a sharded deployment. The application connects to a `mongos`, which reads the shard key from the document, consults routing metadata cached from the config servers, and forwards the insert to the shard owning that key range. Each shard is itself a replica set: its primary checks the document against the collection's validator, if one is set, applies the write, records it in the oplog, and secondaries copy it. The write concern decides when the client hears back. A read takes the reverse path. If its filter carries the shard key, `mongos` contacts one shard, or a few; otherwise it asks all of them and merges the results. On each shard the query planner picks an index, and read preference decides whether a primary or a secondary answers. An aggregation runs the same way with more steps: its leading stages are pushed to the shards and use their indexes, and later stages may run on a merging node. That is why the sections depend on each other. The data model fixes which queries exist; the queries fix which indexes are worth their write cost; the hottest query shape should also carry the shard key. One design decision often serves all three: ```javascript // ESR order: customerId (equality), createdAt (sort), total (range) db.orders.createIndex({ customerId: 1, createdAt: -1, total: 1 }) // the same leading field is the shard key, so this index also supports it sh.shardCollection("shop.orders", { customerId: 1 }) // targeted to one shard, index-backed, sorted without a separate sort step db.orders.find({ customerId: 42, total: { $gte: 100 } }) .sort({ createdAt: -1 }) ``` Replication sits underneath all of it. Durability depends on write concern, a failover can roll back writes that never reached a majority, and reading from secondaries trades freshness for spare capacity. Transactions ride on the same machinery: they need a replica set, and across shards they add a coordinator and extra round trips.
- Document Model & BSON →
Documents, BSON types and embed-or-reference: the model every query, index and shard key is designed around.
- CRUD & Query Operators →
The everyday read and write API, its operators, and where single-document atomicity ends and transactions begin.
- Indexing & Performance →
How queries get fast: index types, compound field order, and reading explain() output as evidence.
- Aggregation Pipeline →
Grouping, joining and reshaping on the server, and why stage order decides whether a pipeline uses indexes.
- Replica Sets →
What an acknowledged write survives, how failover works, and which member a read is allowed to hit.
- Sharding & Scalability →
Horizontal scale, which only makes sense once queries, indexes and replication are clear.
Treating flexible schema as no schema: without a validator or a single model in code, type drift means a query on a number silently misses documents stored as strings.
Embedding an array that grows without bound, such as every comment or event, until documents get slow to rewrite and approach the size limit.
Porting a normalized relational schema one table per collection, then rebuilding every read with
$lookupjoins that the document model was meant to avoid.Paging deep result sets with
skip(); the server still walks every skipped entry, so page 10,000 costs far more than page 1.Reaching for multi-document transactions by default, when a model that keeps the invariant inside one document is cheaper and needs no retry logic.
Assuming an acknowledged write is safe after failover when it was only acknowledged by the primary — see rollback after failover.
Reading from secondaries to scale reads without accepting that they lag, so a user may not see the write they just made.
Choosing a monotonically increasing field such as a timestamp or ObjectId as a ranged shard key, which sends every insert to the same shard.
Sharding before the query shapes are known; a filter without the shard key is broadcast to every shard, and adding shards makes that worse, not better.
This guide assumes a current MongoDB release, 7.0 or later, running the WiredTiger storage engine. Interviewers still ask about the releases where the rules changed, because older clusters and old blog posts describe different behaviour: - **4.0** added multi-document transactions on replica sets, and **4.2** extended them to sharded clusters. - **4.4** made it possible to refine a shard key by adding suffix fields. - **5.0** introduced time series collections and live resharding, so a bad shard key is no longer fixable only by dumping and reloading. It also made `w: "majority"` the default write concern for most deployments. - **6.x** reworked balancing around data size per shard rather than chunk counts, which is why newer documentation talks about ranges more than chunks. When an answer depends on these — whether transactions work across shards, whether a shard key can change, what the default write concern is — say which version you mean. Answers written against pre-4.0 MongoDB, such as "MongoDB has no transactions", are the most common outdated claims interviewers hear.
Interviewers expect you to place MongoDB, not just use it. Its closest competitors are other document and key-value stores: Couchbase, Amazon DynamoDB and Azure Cosmos DB, the last of which also offers a MongoDB-compatible API. The comparison that comes up most is against a relational database with JSON columns, typically PostgreSQL with `jsonb`: MongoDB favours data that is read and written as whole aggregates, with scaling built into the product, while a relational store favours many-to-many relationships, joins and constraints across tables. Wide-column stores such as Apache Cassandra fit write-heavy, query-known workloads at larger scale with fewer query features. Around it sit the layers you will be asked about in application interviews. Object-document mappers such as Mongoose for Node.js and Spring Data MongoDB for Java add schemas and repositories on the client side. Change streams feed events to messaging systems such as Kafka. MongoDB Atlas, the managed service from MongoDB, Inc., removes server operations but not modelling, indexing or shard key design; those questions follow you to any hosting choice.
explore
- Document Model & BSON17 questions
- BSON Types & _id6 questions
- Schema Validation5 questions
- MongoDB Schema Patterns6 questions
- CRUD & Query Operators23 questions
- Query Operators & Matching6 questions
- Writes & Update Operators6 questions
- Projections, Cursors & Pagination5 questions
- Sessions & Multi-Document Transactions6 questions
- Aggregation Pipeline18 questions
- Core Stages & Expressions6 questions
- $lookup, Graph & Set Stages6 questions
- Pipeline Optimization & Output6 questions
- Indexing & Performance18 questions
- Index Types6 questions
- Index Options & Lifecycle6 questions
- explain(), Selectivity & ESR6 questions
- Replica Sets18 questions
- Replica Set Topology & Oplog6 questions
- Elections & Failover6 questions
- Write/Read Concerns & Read Preference6 questions
- Sharding & Scalability18 questions
- Shard Key Selection6 questions
- Chunks, Balancing & Zones6 questions
- Query Routing: Targeted vs Scatter-Gather6 questions
questions
112 · 6 sectionsWhat rules does MongoDB enforce on the _id field of every document?
basics
~20 sEvery document must have an _id. MongoDB creates a unique index on it that cannot be dropped, the value is immutable after insert, and it may hold any BSON type except an array. If you omit it, an ObjectId is generated.
How do you make a MongoDB collection reject documents that lack required fields?
basics
~20 sAttach a validator to the collection — normally a $jsonSchema object listing required fields and their bsonType — using createCollection or collMod. MongoDB then checks every insert and update against it and rejects writes that break the rules.
What are the 12 bytes of a MongoDB ObjectId, and what can you infer from one?
basics
~20 sAn ObjectId is 12 bytes: a 4-byte Unix timestamp in seconds, a 5-byte value random per process, and a 3-byte incrementing counter. From one you can read its creation time to the second and infer rough insertion order.
What is MongoDB's bucket pattern, and when does bucketing sensor readings beat one document per reading?
basics
~20 sThe bucket pattern stores many small, time-ordered measurements in one document — an hour of readings for one sensor, say — instead of one document per reading. It cuts document count, index entries and per-document overhead for time-series workloads.
What is MongoDB's computed pattern, and when does precomputing a total beat aggregating at read time?
basics
~20 sThe computed pattern stores the result of a calculation — a count, sum, average or rollup — in the document and maintains it on write, so reads fetch a value instead of recomputing it. It pays off when reads far outnumber writes.
In a MongoDB find() projection, can you mix included and excluded fields?
basics
~10 sNo. A find() projection is either an inclusion list or an exclusion list, and mixing the two raises an error. The single exception is _id, which may be excluded inside an otherwise inclusive projection.
In MongoDB, what does the find filter { tags: "red" } match when tags holds an array?
basics
~20 sIt matches documents where tags is exactly the string "red" and documents where tags is an array containing "red" as an element. MongoDB applies the equality predicate to the field's value and to each array element.
In MongoDB, what is the difference between updateOne with $set and replaceOne?
basics
~20 supdateOne with $set changes only the named fields and leaves every other field intact. replaceOne swaps the whole document for the one you pass, so any field the replacement omits disappears. _id cannot be changed either way.
Why can $elemMatch return different results from listing the same two conditions on an array field?
basics
~10 sTwo conditions written side by side may be satisfied by different array elements; $elemMatch requires one single element to satisfy all of them at once. On a non-array field the two forms are equivalent.
Which MongoDB write operations are atomic without opening an explicit transaction?
basics
~20 sEvery write that touches a single document is atomic in MongoDB, no matter how many fields, embedded documents or array elements it changes. Atomicity stops at the document boundary: a write spanning several documents needs an explicit transaction to be all-or-nothing.
What does _id define in a $group stage, and how do you total a field per group?
basics
~10 sIn $group, _id is the grouping-key expression: one output document per distinct _id value. Setting it to null puts everything in one group. Totals come from accumulators, for example revenue: { $sum: "$amount" }.
What does MongoDB's $lookup equality form return for each input document?
basics
~20 s$lookup is a left outer join: for each input document it adds an array field, named by as, holding every document from the from collection whose foreignField equals the input's localField. No match yields an empty array, never a dropped document.
What does $unwind do to a document whose array field is empty or missing?
basics
~20 s$unwind emits one document per array element, copying the other fields. A document whose array is empty, missing or null produces no output at all — it is dropped — unless you use the object form with preserveNullAndEmptyArrays: true, which passes it through once.
How do you filter or reshape the joined side of a $lookup using let and a sub-pipeline?
basics
~20 sUse the pipeline form: let binds fields of the input document to variables, and the sub-pipeline runs against the foreign collection with those variables available as $$name. Compare them inside a $match with $expr, then $project or $limit to trim what comes back.
Why should $match come early in an aggregation pipeline, and when does MongoDB reorder it for you?
basics
~20 sOnly a leading $match becomes an indexed query on the collection, so filtering first shrinks the stream before expensive stages run. MongoDB's optimizer moves $match ahead of $project, $addFields or $sort only when the filter uses no field those stages create.
In MongoDB explain() output, what is the difference between a COLLSCAN and an IXSCAN stage?
basics
~20 sCOLLSCAN means the server read every document in the collection; IXSCAN means it walked index keys and touched only matching entries. A COLLSCAN on a filtered query over a large collection signals a missing index.
In MongoDB, what does a unique index enforce, and what happens when a write violates it?
basics
~20 sA unique index rejects any write that would create a second document with the same indexed value, returning an E11000 duplicate key error. On a compound unique index the whole key combination must be unique, not each field on its own.
What is the ESR rule for ordering the fields of a MongoDB compound index?
basics
~20 sESR orders compound-index keys as Equality first, then Sort fields, then Range fields. Equality pins the scan to a narrow contiguous key range, the sort fields then arrive already ordered, and range predicates go last because they destroy ordering for the keys after them.
What do totalKeysExamined, totalDocsExamined and nReturned reveal about a MongoDB query?
basics
~20 sThey are the work-versus-result ratio of a query. nReturned is what the client got; totalKeysExamined is index entries read; totalDocsExamined is documents loaded. Ratios near 1:1:1 mean a well-matched index; large gaps mean wasted work.
How does a MongoDB TTL index with expireAfterSeconds actually delete expired documents?
basics
~20 sA background thread wakes roughly every 60 seconds, finds documents whose indexed date is older than expireAfterSeconds, and deletes them with ordinary delete operations. Expiry is therefore approximate, not instant, and only the primary performs the deletes.
What roles do the primary, secondary, and arbiter play in a MongoDB replica set?
basics
~20 sThe primary is the only member that accepts writes and records them in its oplog. Secondaries hold full copies of the data, replay that oplog, and can be elected primary. An arbiter stores no data and only votes.
What does a MongoDB write concern of w: "majority" actually guarantee about a write?
basics
~20 sA w: "majority" acknowledgment means the write reached a majority of the replica set's voting, data-bearing members, so it survives an election and cannot be rolled back. It does not mean every member has it.
Why does a MongoDB replica set with four voting members tolerate no more failures than one with three?
basics
~20 sElecting a primary needs a strict majority of the configured voting members. Three members need two votes, four need three, so both survive exactly one loss. The fourth vote adds cost and tie risk without adding fault tolerance.
Why does MongoDB rewrite operations into idempotent form before writing them to the oplog?
basics
~20 sSo an entry can be replayed any number of times with the same result. Secondaries and recovering members may re-apply the tail of the oplog after a restart or a sync-source switch, and a relative operation replayed twice would corrupt the data.
Which MongoDB settings must you combine so a user always reads their own write from a secondary?
basics
~20 sRun both operations inside one causally consistent client session, write with w: "majority" and read with readConcern: "majority". The session carries a cluster timestamp so the secondary waits until it has applied that write before answering.
What is a chunk in a sharded MongoDB collection, and what does the balancer do with chunks?
basics
~20 sA chunk (range) is a contiguous span of shard-key values owned by exactly one shard and recorded in config-server metadata. The balancer migrates ranges from the shard holding the most data for a collection to the shard holding the least.
Which query filters let mongos target one shard, and which force a broadcast to every shard?
basics
~20 sA filter containing the shard key — or a leading prefix of a compound shard key — lets mongos map values to chunks and contact only the owning shards. Any other filter, including one on _id when _id is not the shard key, is broadcast to all shards.
What makes a good MongoDB shard key, and how do cardinality, frequency and monotonicity affect it?
basics
~10 sA good MongoDB shard key has many distinct values, no single value that dominates the data, values that do not grow monotonically, and it appears in common query filters so requests reach one shard.
During a MongoDB chunk migration, what happens on the donor and recipient shards, and how does live traffic feel it?
basics
~20 sThe recipient clones the range's documents and then catches up on changes; the donor enters a short critical section that blocks writes to that range while the config metadata is committed; afterwards the donor deletes the moved documents in the background as orphans.
After sharding, p99 read latency got worse and every shard is busy — how do you confirm broadcast queries are the cause?
basics
~20 sRun explain() through mongos on the hot query shapes: a SHARD_MERGE stage with every shard named in the winning plan means scatter-gather. Compare per-shard keysExamined against nReturned, then fix the filter to carry the shard key rather than adding shards.