What breaks as a MongoDB document approaches the 16 MB BSON limit, and how do you avoid it?
answer
- one number every MongoDB interview expects
- writes fail, they do not truncate
- aggregation output is bounded by it too
- the cause is almost always one growing array
- binaries have their own supported mechanism
basics
~20 sMongoDB caps a BSON document at 16 MB. A write that would exceed it fails outright, and performance suffers long before that, because whole documents are read and written. Split unbounded arrays into their own documents and use GridFS for large binaries.
solid answer
~50 sThe maximum BSON document size is 16 MB (16,777,216 bytes), and it is a hard limit: an `$push` or `$set` that would take a document past it fails with an error rather than truncating anything. It also bounds documents *returned* — each document a query or aggregation produces must fit, so a `$group` that collects a huge array into one result document fails even though the input was fine. Long before you reach the ceiling, big documents hurt: WiredTiger reads and writes whole documents, so every small update rewrites megabytes, the working set inflates, and a multikey index over a huge array pays index maintenance on every write. The usual cause is an **unbounded array** — events, readings, comments pushed forever into a parent document. The fix is to stop growing one document: give each child its own document, or for large binary payloads use GridFS, which splits a file into chunks (255 kB by default) across two collections.
code
javascript · 6 lines// unbounded: every reading rewrites the whole device document
db.devices.updateOne({ _id: id }, { $push: { readings: r } })
// bounded: one document per reading, parent stays small
db.readings.insertOne({ deviceId: id, ts: new Date(), value: r.value })
db.readings.createIndex({ deviceId: 1, ts: -1 })go deeper
Recall the number: a single BSON document cannot exceed 16 MB, and files larger than that belong in GridFS rather than in a document field.
Explain the failure mode — the write is rejected, not truncated — and that the limit also bounds documents produced by an aggregation, not just documents stored.
Diagnose the real problem: unbounded arrays causing whole-document rewrites, oplog and replication amplification, multikey index churn and working-set pressure long before 16 MB is reached, and describe the split you would perform.
Own the modelling rule that prevents it — bounded arrays embed, unbounded collections get their own documents — and the migration plan for a live collection whose largest documents are already near the ceiling.
## The limit A single BSON document may not exceed **16 MB** — 16,777,216 bytes. The limit applies to documents stored in a collection and equally to documents produced by a query or an aggregation. A related structural cap: BSON documents may not nest more than 100 levels deep. The limit exists because a document is MongoDB's unit of atomicity, transfer and storage. The server must be able to hold one in memory, ship it over the wire, and rewrite it as a unit; an unbounded document size would make memory use and latency unbounded too. ## How it actually fails - **On write.** An update that would push a document past 16 MB is rejected with an error naming the size. Nothing is truncated and nothing is silently dropped — the write simply does not happen. In an append-heavy design this means the feature works perfectly for months and then starts failing for exactly the busiest, most valuable documents, which is the worst possible failure distribution. - **On read.** Every document a cursor returns must fit in 16 MB, which is normally automatic — but an aggregation that builds a document, for instance `$group` with `$push` collecting every member of a large group into one array, can produce an over-limit result document and fail even though every input document was small. - **Not a result-set limit.** A query returning a million small documents is fine; cursors stream them in batches. The cap is per document. ## The degradation before the cliff The cliff at 16 MB is rarely the real problem, because the design is already suffering well below it: - **Whole-document reads and writes.** The storage engine rewrites the document, not the changed field. Appending one 200-byte entry to a 4 MB document costs a 4 MB write, plus the corresponding journal and replication traffic — the oplog entry for the change travels to every secondary. - **Working-set inflation.** Fetching one field of a huge document pulls the whole thing into cache, evicting other data. Projection reduces what crosses the network, not what is read from storage. - **Index maintenance.** An index on an array field is multikey and stores one index entry per array element per document. A 50,000-element array is 50,000 index entries, all revisited as the document changes. - **Contention.** All writers to that one growing document contend on it, where a per-child design would spread writes across many documents. ## Designing so it cannot happen The governing question is: **is this array bounded by something in the domain?** A handful of shipping addresses, a fixed set of translations, the line items of one order — bounded, and embedding is correct. Events, readings, log lines, comments, audit entries, followers — unbounded, and embedding is a bug waiting for a busy customer. For unbounded collections, give each child its own document with a field referencing the parent, and index that field. Reads become a second query, but writes stop rewriting megabytes, deletion becomes a targeted operation, and the parent stops growing. Where you need a bounded slice inside the parent — the last ten comments, say — keep only that slice and hold the full history separately. For large **binary** payloads, do not stuff bytes into a field at all. GridFS is the supported mechanism: it stores file metadata in one collection and fixed-size chunks (255 kB by default) in another, so files of any size can be stored and streamed. If the bytes are really an object-store concern — images, video, backups — storing them outside the database and keeping only a reference is usually better still. ## Diagnosing an existing collection Collection statistics give you average document size, which is a blunt but useful early warning; an aggregation projecting `$bsonSize` on the document (or `$size` on the suspect array) finds the specific offenders and shows the distribution. If the largest documents are orders of magnitude above the average, you have found the unbounded array, and the tail is where the failures will start.
- Does the 16 MB limit apply to the result of a query?It applies per document, not per result set. A cursor can stream millions of documents, but each one it returns must fit in 16 MB. That is why an aggregation can fail at $group even when every input document was small: collecting a large group into one array with $push can build an over-limit output document.
- How do you find which documents in a collection are dangerously large?Run an aggregation projecting $bsonSize of the document, sort descending, and look at the tail. Projecting $size on the suspect array shows which array is driving the growth. Collection statistics give average document size as a coarse early warning, but the distribution is what matters — failures start with the largest documents.
- When is embedding an array still the right call?When the array is bounded by the domain — a person's addresses, an order's line items, a fixed set of translations — and is read with its parent. Bounded plus read-together means embedding saves a query and keeps the update atomic. Unbounded growth driven by time or user activity is the case that must be split out.
saying these in an interview costs you the question
- Thinks an oversized update truncates the document instead of failing
- Believes the 16 MB limit caps the whole result set
- Stores large binary files directly in a document field
- Treats an ever-growing array of events as normal modelling
- Assumes a small update to a huge document is a small write