skip to content

Why do unbounded embedded arrays eventually break a document-database design?

level: middleimportance: must knowfreq 66%

answer

  1. Arrays are stored inside the parent record
  2. Every append rewrites the whole document
  3. There is a hard cap per document
  4. Growth is skewed toward a few parents

basics

~20 s

An array with no ceiling eventually pushes the document past the store's hard size limit, and long before that every read and every update of the parent pays for the whole array. Move those children into their own records.

solid answer

~50 s

An embedded array lives inside the parent document, so its cost is charged to every operation on that parent. Four things degrade as it grows. First, there is a hard maximum document size — MongoDB caps a single document at 16 MB, and other document stores impose their own limit — and once the array reaches it, writes simply fail, in production, on your most popular parent. Second, appending an element usually means rewriting the whole record, so write cost grows with array length. Third, any read of the parent transfers the entire array unless you project it away. Fourth, an index over the array holds one entry per element, so the index inflates too. The distribution makes it worse: growth is skewed, so the design looks fine until one hot parent breaks. Store the children as their own documents carrying the parent's identifier, and keep at most a bounded recent slice inside the parent.

code

json · 8 lines
json
{
  "_id": "post-42",
  "title": "Why documents grow",
  "comments": [
    { "author": "ana", "text": "first" },
    { "author": "bo", "text": "second" }
  ]
}

go deeper

for a junior

Know that a document has a maximum size and that arrays inside it count toward that size, so a list that keeps growing does not belong inside its parent.

for a middle

Explain the mechanics: the document is the unit of write, so appends rewrite it, reads carry it, and an index over the array holds one entry per element. Then name the hard cap.

for a senior

Show that you reason about the distribution, not the average, and describe the migration off an already-fat model: children into their own collection, a bounded slice kept in the parent for the hot read.

for a principal

Own the guardrails — size and array-length monitoring on every collection, a review rule that every embedded array names what bounds it, and a plan for the entities whose growth outruns the original model.

## Where the array actually lives An embedded array is not a separate structure the engine manages on your behalf. It is bytes inside the parent record. That single fact explains every symptom: whatever the array costs, the parent pays, on every operation. ## The hard ceiling Document stores impose a maximum size on a single document. MongoDB's limit is 16 MB; other products set their own. This is not a soft guideline that degrades gracefully — when the document reaches the cap, the write that would exceed it is rejected. The failure has an ugly shape: it appears first on your most successful entity (the viral post, the largest customer, the busiest device), in production, on a write path that has worked for months. There is no incremental warning, and the repair — restructuring the model and migrating the data — is the most expensive kind of change to make under pressure. ## Write amplification Storage engines write documents as units. Appending one small element to a 4 MB array can require rewriting a 4 MB record, and, if the document has grown beyond its allocated space, relocating it. The cost of adding element *n* is proportional to the size of the whole array, so the total cost of building the array is quadratic rather than linear. A workload that appends steadily to one document also concentrates all of that write traffic — and any concurrency control the engine applies per document — on a single record, so appends to different children contend with each other. ## Read amplification Any operation that fetches the parent fetches the array with it unless you explicitly project it away, and every caller has to remember to. A page that just wants a post's title and author transfers the whole comment history across the wire. The same bytes displace other data in cache: one 12 MB document evicts a large number of small, frequently-read documents from the working set, so a single fat entity degrades the hit rate for everything else. ## Index growth An index on a field inside an array stores one index entry per array element per document. A parent with 200,000 elements contributes 200,000 entries. That inflates the index, slows the updates that maintain it, and makes the index itself less likely to stay resident in memory. ## Telling bounded from unbounded The test is not "how big is it today" but "what stops it growing?" A person has a bounded number of mailing addresses because a person has a bounded number of homes. A product has a bounded number of size variants because the catalogue defines them. A post has no bounded number of comments, a device has no bounded number of readings, an account has no bounded number of audit events — nothing in the domain caps them, and the fact that today's median is eleven tells you nothing about the maximum. Growth in these relationships is heavily skewed: the median parent stays small forever while the top of the distribution runs away, which is exactly why the design passes review and fails in production. A second warning sign is time. Anything appended by the passage of time rather than by an act of the domain — events, logs, messages, history — is unbounded by construction. ## What to do instead Give the children their own documents, each carrying the parent's identifier, and index that identifier. Now the children grow without touching the parent, they can be filtered, sorted and paginated on their own, and reads of the parent are cheap again. The cost is a second query on the paths that genuinely need both. Where the hot read needs the parent plus the newest handful of children, keep a bounded slice inside the parent — the last ten comments, the current five alerts — and treat it as a cache of the authoritative collection. Many stores support an update that appends and trims to the last *n* elements in one operation, which keeps the slice bounded without a read-modify-write. Never let that slice be the only copy. A third option, when the children are numerous but individually tiny and are always read as a block, is to group them into batches of fixed size, so one parent maps to many child documents each holding a bounded number of elements. That trades a little query complexity for a bounded document size and far fewer records than one document per element. ## What to say in an interview Name the four costs — the hard cap, write amplification, read amplification, index growth — then say the deciding question is what bounds the array, and then show that you have thought about the skew: the model must survive the largest parent, not the average one.

  • An array is bounded at about 200 elements. Is it safe to embed?
    Usually yes, if the elements are small and read with the parent. The things to check are the byte size of the largest element rather than the count, whether the bound is enforced anywhere or merely observed, and whether every read of the parent really wants all 200. If nothing in the domain or the code prevents the 201st, treat the bound as folklore.
  • You keep the newest ten children inside the parent as a cache. What can go wrong?
    It can drift from the authoritative collection when an append succeeds in one place and fails in the other, and it can go stale when a child is edited or deleted after being copied. Make the slice cheap to rebuild from the source, never treat it as the record of truth, and decide explicitly how much staleness the screen tolerates.
  • How would you detect an array heading for the size limit before it fails?
    Monitor the distribution, not the mean: track the maximum document size and the maximum array length per collection, and alert on a percentile approaching a fraction of the cap. Because growth is skewed, an average-based dashboard will look healthy right up to the failure.

saying these in an interview costs you the question

  • Claims the database splits an oversized document automatically
  • Judges an array by its median length, not its maximum
  • Thinks appending one element costs the same at any length
  • Believes there is no upper limit on document size
  • Adds an index on the array instead of moving the children out

context