Document
Databases whose primary unit is a schema-flexible JSON or BSON document rather than a row in a fixed table. Interviewers ask because the modelling instincts differ from SQL: you embed what you read together instead of normalising it away.
on this pageshowhide
explore
- MongoDB (has its own guide)112 questions
- Document Model & BSON17 questions
- CRUD & Query Operators23 questions
- Aggregation Pipeline18 questions
- Indexing & Performance18 questions
- Replica Sets18 questions
- Sharding & Scalability18 questions
- MongoDB Atlas17 questions
- Clusters, Tiers & Serverless5 questions
- Atlas Search & Vector Search6 questions
- Atlas Operations & Security6 questions
- CouchDB12 questions
- Documents, Revisions & Views6 questions
- Replication Protocol & Offline Sync6 questions
- Couchbase12 questions
- Buckets, Collections & N1QL6 questions
- Clustering, Durability & XDCR6 questions
- RethinkDB5 questions
- Document Data Modeling28 questions
- Embedding vs Referencing5 questions
- Aggregate-Oriented Design6 questions
- Denormalization & Duplication Trade-offs6 questions
- Schema Flexibility & Validation6 questions
- Schemaless Migrations & Document Versioning5 questions
→ has its own guide
- Backend Developerroleanchors this topic
- Data Engineerroleanchors this topic
- Full Stack Developerroleanchors this topic
- Java Backend Developerroleanchors this topic
- Kotlin Backend Developerroleanchors this topic
- Software Architectroleanchors this topic
- Forward Deployed Engineerrole
- MongoDBskill
- Server-Side Game Developerrole
questions
186 · 6 sectionsWhat does _id define in a $group stage, and how do you total a field per group?
basics
~10 sIn $group, _id is the grouping-key expression: one output document per distinct _id value. Setting it to null puts everything in one group. Totals come from accumulators, for example revenue: { $sum: "$amount" }.
What does MongoDB's $lookup equality form return for each input document?
basics
~20 s$lookup is a left outer join: for each input document it adds an array field, named by as, holding every document from the from collection whose foreignField equals the input's localField. No match yields an empty array, never a dropped document.
In a MongoDB find() projection, can you mix included and excluded fields?
basics
~10 sNo. A find() projection is either an inclusion list or an exclusion list, and mixing the two raises an error. The single exception is _id, which may be excluded inside an otherwise inclusive projection.
In MongoDB, what does the find filter { tags: "red" } match when tags holds an array?
basics
~20 sIt matches documents where tags is exactly the string "red" and documents where tags is an array containing "red" as an element. MongoDB applies the equality predicate to the field's value and to each array element.
In MongoDB, what is the difference between updateOne with $set and replaceOne?
basics
~20 supdateOne with $set changes only the named fields and leaves every other field intact. replaceOne swaps the whole document for the one you pass, so any field the replacement omits disappears. _id cannot be changed either way.
In MongoDB Atlas, how do database users differ from the users who sign in to the Atlas console?
basics
~20 sAtlas console accounts are organization and project members who manage the deployment through the UI or Admin API. Database users are separate credentials, created per project, that authenticate to the cluster itself and carry MongoDB roles such as readWrite on named databases.
In Atlas Backup, how do scheduled cloud snapshots differ from continuous cloud backup?
basics
~20 sScheduled cloud snapshots are cloud-provider volume snapshots taken on a policy you define, so you can restore only to a snapshot instant. Continuous cloud backup additionally captures oplog entries, letting you restore to any timestamp inside a configured restore window.
Why must $search be the first stage of an Atlas aggregation pipeline?
basics
~20 s$search runs on mongot, a separate Lucene process, and produces the pipeline's result stream from its own index instead of filtering documents handed to it. Being the source of the stream, it cannot follow another stage.
In $vectorSearch, what does numCandidates control and how do you choose it?
basics
~20 snumCandidates is how many near neighbours the approximate search examines before the best limit are returned. Raising it improves recall and costs latency; it must be at least limit, and a common starting point is ten to twenty times limit.
Why must a CouchDB document update send the document's current _rev, and what happens if it is stale?
basics
~20 sCouchDB uses _rev for optimistic concurrency. An update must carry the _rev the client last read; if that is no longer the document's current revision, CouchDB rejects the write with HTTP 409 Conflict instead of overwriting.
When a CouchDB document has conflicting revisions, how does every node deterministically pick the same winner?
basics
~20 sCouchDB stores both branches and picks a winner by a fixed rule: a live leaf beats a deleted one, then the leaf with the highest generation number wins, and equal generations are broken by choosing the higher revision hash. Every node computes the same answer without coordinating.
What does CouchDB's _changes feed return, and how do the normal, longpoll and continuous modes differ?
basics
~20 sThe _changes feed lists documents that changed since a given sequence, one row per document with its current revision. feed=normal returns immediately, feed=longpoll holds the request open until a change arrives, feed=continuous streams rows indefinitely.
Which HTTP calls does a CouchDB replication make to move changes from source to target?
basics
~20 sReplication reads the source's _changes feed, asks the target's _revs_diff which of those revisions it lacks, fetches the missing ones with open_revs, and posts them to the target's _bulk_docs with new_edits set to false. Progress is checkpointed in _local documents.
What does PouchDB's db.sync(remote, {live: true, retry: true}) do in an offline-first app?
basics
~20 sIt starts two continuous replications, local to remote and remote to local, that keep running as changes occur and reconnect automatically after failures. Writes land in the local database first, so the app works offline and catches up when connectivity returns.
In Couchbase 7.x, how do buckets, scopes and collections relate to each other?
basics
~10 sA Couchbase bucket is the top-level container that owns the memory quota, replica count and persistence settings. Inside it, scopes group collections, and collections hold the documents. The full keyspace path is bucket.scope.collection.
Why does a Couchbase SQL++ query fail with "No index available" when the documents exist?
basics
~20 sCouchbase's Query service reaches documents only through an index. With no primary index and no secondary GSI whose leading key matches a predicate, there is no access path, so the statement is rejected rather than scanning the whole collection.
What do Couchbase durability levels majority, majorityAndPersistActive and persistToMajority guarantee?
basics
~20 sIn Couchbase, majority means the mutation is in memory on a majority of the active-plus-replica copies; majorityAndPersistActive adds a disk write on the active; persistToMajority requires it on disk on a majority. Each step up costs latency.
How does Couchbase map a document key to a node using vBuckets?
basics
~20 sCouchbase hashes the document key into one of a bucket's 1024 vBuckets, then looks that vBucket up in the cluster map to find the node holding its active copy. The SDK does this locally, so requests go straight to the owning node.
When should you read a Couchbase document with the KV API instead of a SQL++ query?
basics
~20 sWhenever the document key is known. A KV read hashes the key to the node that owns it and is served from that node's managed cache, with no index lookup, no query planning and no hop through the Query service.
What does a RethinkDB changefeed emit when you call .changes() on a table?
basics
~10 sCalling .changes() returns a cursor that stays open and pushes one document per change, each shaped {old_val, new_val}. An insert has old_val null, a delete has new_val null, an update carries both.
What delivery guarantees does a RethinkDB changefeed give, and what happens if the client disconnects?
basics
~20 sChangefeeds are live-only push, not a durable log. There is no offset to resume from, so anything that happens while a client is disconnected is lost, and consecutive changes to one document may be coalesced into a single before/after pair.
How does RethinkDB's ReQL query language differ from sending SQL query strings?
basics
~20 sReQL is embedded in the host language: you chain driver methods to build a query object, then call run(connection) to execute it. Nothing is sent as a string, and nothing runs until run() is called.
How do you build a live top-10 leaderboard in RethinkDB with a single changefeed?
basics
~10 sAttach a changefeed to an ordered, limited query: orderBy on an index, limit(10), then .changes({includeInitial: true}). The feed sends the current top ten first, then one change each time the window's membership shifts.
What is RethinkDB's maintenance status, and how would you weigh adopting it today?
basics
~20 sRethinkDB the company shut down in 2016; the code was relicensed under Apache 2.0 and moved to the Linux Foundation, maintained since by a small community with infrequent releases. Adopting it now is mainly a sustainability bet, not a technical one.
A product's name is copied into every order document. What breaks when that product is renamed?
basics
~20 sNothing updates the copies automatically: old order documents keep the old name until some process rewrites them. Either the write path owns propagation, or the copy is deliberately frozen as a historical snapshot and named to say so.
When should you embed related data in a document and when should you store a reference?
basics
~10 sEmbed when the related data is bounded in size, read with its parent, and changes with it. Reference when it is unbounded, shared by many parents, edited independently, or queried on its own.
What does calling a document database "schemaless" actually mean for your application?
basics
~20 sIt means the database does not check a document's shape when you write it — not that there is no schema. The schema still exists, implicitly, in the code that writes and reads documents. Enforcement moved; it did not disappear.
Why is a single document the consistency scope in a document database, and how does that shape aggregate boundaries?
basics
~20 sA write to one document lands entirely or not at all, so one document is the widest scope where an invariant can be enforced for free. Any rule that must always hold should sit inside a single aggregate; rules spanning documents need extra machinery or must tolerate lag.
What does query-first, or 'design for the read', modelling mean in a document database?
basics
~20 sQuery-first modelling starts from the list of operations the application must serve and shapes documents so each one is answered by a simple, well-targeted access. The document shape follows the workload, not a diagram of entities.