skip to content

Document

Databases whose primary unit is a schema-flexible JSON or BSON document rather than a row in a fixed table. Interviewers ask because the modelling instincts differ from SQL: you embed what you read together instead of normalising it away.

on this pageshow

explore

→ has its own guide

questions

186 · 6 sections

What does _id define in a $group stage, and how do you total a field per group?

level: juniorimportance: must knowfreq 85%
basics
~10 s

In $group, _id is the grouping-key expression: one output document per distinct _id value. Setting it to null puts everything in one group. Totals come from accumulators, for example revenue: { $sum: "$amount" }.

open as a page

What does MongoDB's $lookup equality form return for each input document?

level: juniorimportance: must knowfreq 80%
basics
~20 s

$lookup is a left outer join: for each input document it adds an array field, named by as, holding every document from the from collection whose foreignField equals the input's localField. No match yields an empty array, never a dropped document.

open as a page

In a MongoDB find() projection, can you mix included and excluded fields?

level: juniorimportance: must knowfreq 72%
basics
~10 s

No. A find() projection is either an inclusion list or an exclusion list, and mixing the two raises an error. The single exception is _id, which may be excluded inside an otherwise inclusive projection.

open as a page

In MongoDB, what does the find filter { tags: "red" } match when tags holds an array?

level: juniorimportance: must knowfreq 78%
basics
~20 s

It matches documents where tags is exactly the string "red" and documents where tags is an array containing "red" as an element. MongoDB applies the equality predicate to the field's value and to each array element.

open as a page

In MongoDB, what is the difference between updateOne with $set and replaceOne?

level: juniorimportance: must knowfreq 78%
basics
~20 s

updateOne with $set changes only the named fields and leaves every other field intact. replaceOne swaps the whole document for the one you pass, so any field the replacement omits disappears. _id cannot be changed either way.

open as a page

What limits do Atlas shared tiers M0, M2 and M5 impose compared with a dedicated cluster?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Atlas shared tiers run on multi-tenant hosts with capped storage (M0 gives 512 MB), throttled throughput and connections, and no dedicated CPU or RAM. Sharding, auto-scaling, multi-region node placement and private networking all require a dedicated M10 or larger cluster.

open as a page

In MongoDB Atlas, how do database users differ from the users who sign in to the Atlas console?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Atlas console accounts are organization and project members who manage the deployment through the UI or Admin API. Database users are separate credentials, created per project, that authenticate to the cluster itself and carry MongoDB roles such as readWrite on named databases.

open as a page

In Atlas Backup, how do scheduled cloud snapshots differ from continuous cloud backup?

level: middleimportance: must knowfreq 58%
basics
~20 s

Scheduled cloud snapshots are cloud-provider volume snapshots taken on a policy you define, so you can restore only to a snapshot instant. Continuous cloud backup additionally captures oplog entries, letting you restore to any timestamp inside a configured restore window.

open as a page

Why must $search be the first stage of an Atlas aggregation pipeline?

level: middleimportance: must knowfreq 70%
basics
~20 s

$search runs on mongot, a separate Lucene process, and produces the pipeline's result stream from its own index instead of filtering documents handed to it. Being the source of the stream, it cannot follow another stage.

open as a page

In $vectorSearch, what does numCandidates control and how do you choose it?

level: middleimportance: must knowfreq 65%
basics
~20 s

numCandidates is how many near neighbours the approximate search examines before the best limit are returned. Raising it improves recall and costs latency; it must be at least limit, and a common starting point is ten to twenty times limit.

open as a page

Why must a CouchDB document update send the document's current _rev, and what happens if it is stale?

level: juniorimportance: must knowfreq 72%
basics
~20 s

CouchDB uses _rev for optimistic concurrency. An update must carry the _rev the client last read; if that is no longer the document's current revision, CouchDB rejects the write with HTTP 409 Conflict instead of overwriting.

open as a page

When a CouchDB document has conflicting revisions, how does every node deterministically pick the same winner?

level: middleimportance: must knowfreq 58%
basics
~20 s

CouchDB stores both branches and picks a winner by a fixed rule: a live leaf beats a deleted one, then the leaf with the highest generation number wins, and equal generations are broken by choosing the higher revision hash. Every node computes the same answer without coordinating.

open as a page

What does CouchDB's _changes feed return, and how do the normal, longpoll and continuous modes differ?

level: middleimportance: must knowfreq 70%
basics
~20 s

The _changes feed lists documents that changed since a given sequence, one row per document with its current revision. feed=normal returns immediately, feed=longpoll holds the request open until a change arrives, feed=continuous streams rows indefinitely.

open as a page

Which HTTP calls does a CouchDB replication make to move changes from source to target?

level: seniorimportance: must knowfreq 55%
basics
~20 s

Replication reads the source's _changes feed, asks the target's _revs_diff which of those revisions it lacks, fetches the missing ones with open_revs, and posts them to the target's _bulk_docs with new_edits set to false. Progress is checkpointed in _local documents.

open as a page

What does PouchDB's db.sync(remote, {live: true, retry: true}) do in an offline-first app?

level: juniorimportance: should knowfreq 40%
basics
~20 s

It starts two continuous replications, local to remote and remote to local, that keep running as changes occur and reconnect automatically after failures. Writes land in the local database first, so the app works offline and catches up when connectivity returns.

open as a page

In Couchbase 7.x, how do buckets, scopes and collections relate to each other?

level: juniorimportance: must knowfreq 70%
basics
~10 s

A Couchbase bucket is the top-level container that owns the memory quota, replica count and persistence settings. Inside it, scopes group collections, and collections hold the documents. The full keyspace path is bucket.scope.collection.

open as a page

Why does a Couchbase SQL++ query fail with "No index available" when the documents exist?

level: middleimportance: must knowfreq 68%
basics
~20 s

Couchbase's Query service reaches documents only through an index. With no primary index and no secondary GSI whose leading key matches a predicate, there is no access path, so the statement is rejected rather than scanning the whole collection.

open as a page

What do Couchbase durability levels majority, majorityAndPersistActive and persistToMajority guarantee?

level: middleimportance: must knowfreq 65%
basics
~20 s

In Couchbase, majority means the mutation is in memory on a majority of the active-plus-replica copies; majorityAndPersistActive adds a disk write on the active; persistToMajority requires it on disk on a majority. Each step up costs latency.

open as a page

How does Couchbase map a document key to a node using vBuckets?

level: middleimportance: must knowfreq 70%
basics
~20 s

Couchbase hashes the document key into one of a bucket's 1024 vBuckets, then looks that vBucket up in the cluster map to find the node holding its active copy. The SDK does this locally, so requests go straight to the owning node.

open as a page

When should you read a Couchbase document with the KV API instead of a SQL++ query?

level: middleimportance: should knowfreq 52%
basics
~20 s

Whenever the document key is known. A KV read hashes the key to the node that owns it and is served from that node's managed cache, with no index lookup, no query planning and no hop through the Query service.

open as a page

What does a RethinkDB changefeed emit when you call .changes() on a table?

level: middleimportance: must knowfreq 70%
basics
~10 s

Calling .changes() returns a cursor that stays open and pushes one document per change, each shaped {old_val, new_val}. An insert has old_val null, a delete has new_val null, an update carries both.

open as a page

What delivery guarantees does a RethinkDB changefeed give, and what happens if the client disconnects?

level: seniorimportance: must knowfreq 55%
basics
~20 s

Changefeeds are live-only push, not a durable log. There is no offset to resume from, so anything that happens while a client is disconnected is lost, and consecutive changes to one document may be coalesced into a single before/after pair.

open as a page

How does RethinkDB's ReQL query language differ from sending SQL query strings?

level: juniorimportance: should knowfreq 50%
basics
~20 s

ReQL is embedded in the host language: you chain driver methods to build a query object, then call run(connection) to execute it. Nothing is sent as a string, and nothing runs until run() is called.

open as a page

How do you build a live top-10 leaderboard in RethinkDB with a single changefeed?

level: middleimportance: should knowfreq 40%
basics
~10 s

Attach a changefeed to an ordered, limited query: orderBy on an index, limit(10), then .changes({includeInitial: true}). The feed sends the current top ten first, then one change each time the window's membership shifts.

open as a page

What is RethinkDB's maintenance status, and how would you weigh adopting it today?

level: principalimportance: nice to knowfreq 30%
basics
~20 s

RethinkDB the company shut down in 2016; the code was relicensed under Apache 2.0 and moved to the Linux Foundation, maintained since by a small community with infrequent releases. Adopting it now is mainly a sustainability bet, not a technical one.

open as a page

A product's name is copied into every order document. What breaks when that product is renamed?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Nothing updates the copies automatically: old order documents keep the old name until some process rewrites them. Either the write path owns propagation, or the copy is deliberately frozen as a historical snapshot and named to say so.

open as a page

When should you embed related data in a document and when should you store a reference?

level: juniorimportance: must knowfreq 85%
basics
~10 s

Embed when the related data is bounded in size, read with its parent, and changes with it. Reference when it is unbounded, shared by many parents, edited independently, or queried on its own.

open as a page

What does calling a document database "schemaless" actually mean for your application?

level: juniorimportance: must knowfreq 78%
basics
~20 s

It means the database does not check a document's shape when you write it — not that there is no schema. The schema still exists, implicitly, in the code that writes and reads documents. Enforcement moved; it did not disappear.

open as a page

Why is a single document the consistency scope in a document database, and how does that shape aggregate boundaries?

level: middleimportance: must knowfreq 62%
basics
~20 s

A write to one document lands entirely or not at all, so one document is the widest scope where an invariant can be enforced for free. Any rule that must always hold should sit inside a single aggregate; rules spanning documents need extra machinery or must tolerate lag.

open as a page

What does query-first, or 'design for the read', modelling mean in a document database?

level: middleimportance: must knowfreq 70%
basics
~20 s

Query-first modelling starts from the list of operations the application must serve and shapes documents so each one is answered by a simple, well-targeted access. The document shape follows the workload, not a diagram of entities.

open as a page