skip to content

MongoDB Schema Patterns

The catalogue of named document layouts MongoDB's own guidance teaches — bucketing time-series writes, isolating outliers, precomputing rollups, modelling trees. Naming the right pattern for a scenario is a fast way to show you have modelled documents in production.

part ofMongoDBoverview, primer and where to startread it →
on this pageshow

questions

6

What is MongoDB's bucket pattern, and when does bucketing sensor readings beat one document per reading?

level: middleimportance: must knowfreq 62%

answer

  1. many small records, one document
  2. group by source and time window
  3. fewer _id index entries to cache
  4. $push plus $inc into a rolling document
  5. native time series does it for you

basics

~20 s

The bucket pattern stores many small, time-ordered measurements in one document — an hour of readings for one sensor, say — instead of one document per reading. It cuts document count, index entries and per-document overhead for time-series workloads.

solid answer

~50 s

Bucketing groups events that are always read together — normally one source plus one time window — into a single document holding an array of measurements and a few rollup fields such as `count`, `sum`, `min` and `max`. Instead of a billion tiny documents each carrying an `_id`, an `_id` index entry and a repeated set of field names, you get far fewer, denser documents, so the working set and the index shrink and a range read becomes a handful of document fetches. Writes are an upsert: `$push` the sample, `$inc` the rollups, with a bound like `count: { $lt: 200 }` in the filter so a full bucket stops matching and a fresh one is inserted. The costs are that single-measurement reads and updates now dig into an array, deletion granularity is the bucket, and a bucket must stay comfortably under the 16 MB BSON limit. From MongoDB 5.0 a native time series collection does the same bucketing inside the server.

code

json · 11 lines
json
{
  "sensorId": "s-42",
  "startTs": "2026-08-20T14:00:00Z",
  "count": 60,
  "sumTemp": 1302.5,
  "maxTemp": 23.4,
  "samples": [
    { "ts": "2026-08-20T14:00:03Z", "temp": 21.7 },
    { "ts": "2026-08-20T14:01:03Z", "temp": 21.8 }
  ]
}

go deeper

for a junior

Recall that bucketing packs many small time-ordered records into one document — an hour of readings for one sensor — rather than storing one document per reading.

for a middle

Be ready to say where the saving comes from: fewer documents, fewer _id index entries, fewer repeated field names, fewer fetches per range query. Show the upsert that appends a sample and increments the rollups.

for a senior

Expect to justify a bucket boundary against real write and read rates, cap the array so a bucket stays well under 16 MB, and state honestly what single-measurement reads, updates and retention now cost.

for a principal

Own the choice between hand-rolled buckets and native time series collections — what control you trade for what maintenance — plus retention, deletion granularity and how the decision survives a tenfold growth in ingest.

## The problem bucketing solves Time-series and event workloads emit a torrent of tiny records: one temperature reading, one page view, one price tick. Modelled naively that is one document per event. Each document carries its own `_id`, its own entry in the `_id` index, a full copy of every field name inside the BSON, and per-record overhead in the storage engine. The measurement itself might be twenty bytes; the container around it is often larger. At hundreds of millions of events the fixed costs, not the payload, dominate storage and — more importantly — cache. An index that no longer fits in RAM turns every read into disk work. A second problem is read shape. Nobody asks for one reading. They ask for "the last six hours for sensor s-42", which against one-document-per-reading means thousands of index entries and thousands of document fetches for one chart. ## What the pattern is The bucket pattern collapses many records that are always read together into one document. The grouping key is normally the source plus a time window: ```json { "sensorId": "s-42", "startTs": "2026-08-20T14:00:00Z", "count": 60, "sumTemp": 1302.5, "minTemp": 20.9, "maxTemp": 23.4, "samples": [ { "ts": "2026-08-20T14:00:03Z", "temp": 21.7 } ] } ``` One index entry on `{ sensorId: 1, startTs: 1 }` now locates an hour of data. The rollup fields are the computed pattern applied inside the bucket: the average for the hour is `sumTemp / count` without touching the array at all. ## Writing into buckets The idiomatic write is one upsert per event. The filter identifies the currently open bucket and carries the bound; the update appends and maintains rollups: ```javascript db.readings.updateOne( { sensorId: "s-42", startTs: hourStart, count: { $lt: 200 } }, { $push: { samples: { ts: now, temp: 21.7 } }, $inc: { count: 1, sumTemp: 21.7 }, $max: { maxTemp: 21.7 } }, { upsert: true } ) ``` When the bucket reaches 200 samples nothing matches, so the upsert inserts a new bucket instead of growing the old one past its bound. All of this happens inside one document, so it is atomic without a transaction. ## Choosing the boundary Two bounds matter. The **semantic** bound is the window you query by — bucket by hour if dashboards ask for hours, because a bucket that spans more than the query window forces you to filter the array. The **physical** bound keeps the document small: cap by element count or estimated size so a bucket cannot approach the 16 MB BSON limit, and remember that very large documents make every update more expensive because the storage engine rewrites the document, not just the appended field. ## What you give up Single-event access is no longer a point read: you match the bucket and then filter or `$unwind` the array, and updating one sample needs the positional operators or `arrayFilters`. Deletes and TTL work at bucket granularity, so retention becomes coarser. Uniqueness on an individual measurement is harder to enforce, because a unique index constrains documents, not array elements. And a badly chosen bucket key — bucketing by time alone, mixing every sensor into one document — creates contention on a single hot document. ## Native time series collections From MongoDB 5.0 the server implements this internally. `db.createCollection("readings", { timeseries: { timeField: "ts", metaField: "sensorId", granularity: "hours" } })` gives you a collection you insert one document per measurement into and query as if it were unbucketed, while the storage layer keeps bucketed, column-oriented internals. `expireAfterSeconds` on the collection handles retention. For new time-stamped workloads that is the default answer; hand-rolled bucketing remains the right tool when the grouping is not time-based — batching many small child records under a parent, for example — or when you need the rollups visible as ordinary queryable fields. ## How to answer it Say what the pattern is in one sentence, name where the win comes from (document count, index entries, repeated field names, fewer fetches per range read), show the bounded upsert, then volunteer the costs and the 16 MB ceiling. Finishing with "and on 5.0+ I'd reach for a native time series collection first" is what separates someone who has read the pattern from someone who has run it.

  • What bounds a bucket besides the time window?
    A size bound. Carry a counter in the document and put it in the upsert filter — `count: { $lt: 200 }` — so once the bucket is full no document matches and the upsert inserts a fresh one. Bound by element count or estimated bytes so a bucket never approaches the 16 MB BSON limit, since oversized documents also make every append more expensive to rewrite.
  • How do you read one individual measurement back out of a bucket?
    Match the bucket by its key, then filter the array — `$unwind` followed by `$match` in an aggregation, or `$filter` in a projection. That is strictly more work than a point read, which is why you bucket along the axis you actually query by. If single-event lookup is a primary access pattern, bucketing is the wrong shape.
  • When would you use a native time series collection instead of hand-rolled buckets?
    For anything time-stamped on MongoDB 5.0 or later. `createCollection` with `timeseries: { timeField, metaField, granularity }` makes the server bucket internally while you insert and query single measurements, and `expireAfterSeconds` gives retention. Hand-rolled buckets stay useful when the grouping is not time-based, or when you need the rollup fields as ordinary indexable document fields.

A bucket is a shipping pallet: sending a thousand parcels individually pays the label, the barcode scan and the manifest line a thousand times; palletizing them by destination and hour pays it once.

saying these in an interview costs you the question

  • Says bucketing is only about saving disk space
  • Lets the samples array grow with no count or size bound
  • Claims bucketing makes writes across several sensors atomic
  • Thinks bucket documents are exempt from the 16 MB limit
  • Buckets by time alone, ignoring the metadata key

context

open as a page

What is MongoDB's computed pattern, and when does precomputing a total beat aggregating at read time?

level: middleimportance: must knowfreq 52%

basics

~20 s

The computed pattern stores the result of a calculation — a count, sum, average or rollup — in the document and maintains it on write, so reads fetch a value instead of recomputing it. It pays off when reads far outnumber writes.

open as a page

In MongoDB's subset pattern, what stays embedded in the main document and how do you cap it?

level: juniorimportance: should knowfreq 45%

basics

~20 s

The subset pattern embeds only the small, hot slice of a large related set — the ten newest reviews, say — and keeps the complete set in its own collection. A $push with the $slice modifier keeps the embedded slice capped.

open as a page

What does MongoDB's attribute pattern do to a document with dozens of optional fields, and why?

level: middleimportance: should knowfreq 40%

basics

~20 s

The attribute pattern turns many rarely-queried, per-type fields into an array of key/value subdocuments such as specs: [{k, v}]. One compound multikey index on specs.k and specs.v then serves searches on any attribute instead of one index per field.

open as a page

For a MongoDB category tree, how do array-of-ancestors and materialized-path documents differ when querying a subtree?

level: seniorimportance: should knowfreq 36%

basics

~20 s

An array of ancestors makes a subtree an equality match on an indexed array — find({ ancestors: "books" }). A materialized path stores the chain as one string and needs an anchored prefix regex. Both rewrite every descendant when a node moves.

open as a page

What is MongoDB's outlier pattern, and how does it stop a few huge documents from shaping the whole schema?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

The outlier pattern keeps the common document shape optimal and handles the rare extreme case separately: a flag field marks documents whose data overflows into extra documents, and only flagged documents pay for the second lookup.

open as a page