skip to content

What is MongoDB's bucket pattern, and when does bucketing sensor readings beat one document per reading?

level: middleimportance: must knowfreq 62%

answer

  1. many small records, one document
  2. group by source and time window
  3. fewer _id index entries to cache
  4. $push plus $inc into a rolling document
  5. native time series does it for you

basics

~20 s

The bucket pattern stores many small, time-ordered measurements in one document — an hour of readings for one sensor, say — instead of one document per reading. It cuts document count, index entries and per-document overhead for time-series workloads.

solid answer

~50 s

Bucketing groups events that are always read together — normally one source plus one time window — into a single document holding an array of measurements and a few rollup fields such as `count`, `sum`, `min` and `max`. Instead of a billion tiny documents each carrying an `_id`, an `_id` index entry and a repeated set of field names, you get far fewer, denser documents, so the working set and the index shrink and a range read becomes a handful of document fetches. Writes are an upsert: `$push` the sample, `$inc` the rollups, with a bound like `count: { $lt: 200 }` in the filter so a full bucket stops matching and a fresh one is inserted. The costs are that single-measurement reads and updates now dig into an array, deletion granularity is the bucket, and a bucket must stay comfortably under the 16 MB BSON limit. From MongoDB 5.0 a native time series collection does the same bucketing inside the server.

code

json · 11 lines
json
{
  "sensorId": "s-42",
  "startTs": "2026-08-20T14:00:00Z",
  "count": 60,
  "sumTemp": 1302.5,
  "maxTemp": 23.4,
  "samples": [
    { "ts": "2026-08-20T14:00:03Z", "temp": 21.7 },
    { "ts": "2026-08-20T14:01:03Z", "temp": 21.8 }
  ]
}

go deeper

for a junior

Recall that bucketing packs many small time-ordered records into one document — an hour of readings for one sensor — rather than storing one document per reading.

for a middle

Be ready to say where the saving comes from: fewer documents, fewer _id index entries, fewer repeated field names, fewer fetches per range query. Show the upsert that appends a sample and increments the rollups.

for a senior

Expect to justify a bucket boundary against real write and read rates, cap the array so a bucket stays well under 16 MB, and state honestly what single-measurement reads, updates and retention now cost.

for a principal

Own the choice between hand-rolled buckets and native time series collections — what control you trade for what maintenance — plus retention, deletion granularity and how the decision survives a tenfold growth in ingest.

## The problem bucketing solves Time-series and event workloads emit a torrent of tiny records: one temperature reading, one page view, one price tick. Modelled naively that is one document per event. Each document carries its own `_id`, its own entry in the `_id` index, a full copy of every field name inside the BSON, and per-record overhead in the storage engine. The measurement itself might be twenty bytes; the container around it is often larger. At hundreds of millions of events the fixed costs, not the payload, dominate storage and — more importantly — cache. An index that no longer fits in RAM turns every read into disk work. A second problem is read shape. Nobody asks for one reading. They ask for "the last six hours for sensor s-42", which against one-document-per-reading means thousands of index entries and thousands of document fetches for one chart. ## What the pattern is The bucket pattern collapses many records that are always read together into one document. The grouping key is normally the source plus a time window: ```json { "sensorId": "s-42", "startTs": "2026-08-20T14:00:00Z", "count": 60, "sumTemp": 1302.5, "minTemp": 20.9, "maxTemp": 23.4, "samples": [ { "ts": "2026-08-20T14:00:03Z", "temp": 21.7 } ] } ``` One index entry on `{ sensorId: 1, startTs: 1 }` now locates an hour of data. The rollup fields are the computed pattern applied inside the bucket: the average for the hour is `sumTemp / count` without touching the array at all. ## Writing into buckets The idiomatic write is one upsert per event. The filter identifies the currently open bucket and carries the bound; the update appends and maintains rollups: ```javascript db.readings.updateOne( { sensorId: "s-42", startTs: hourStart, count: { $lt: 200 } }, { $push: { samples: { ts: now, temp: 21.7 } }, $inc: { count: 1, sumTemp: 21.7 }, $max: { maxTemp: 21.7 } }, { upsert: true } ) ``` When the bucket reaches 200 samples nothing matches, so the upsert inserts a new bucket instead of growing the old one past its bound. All of this happens inside one document, so it is atomic without a transaction. ## Choosing the boundary Two bounds matter. The **semantic** bound is the window you query by — bucket by hour if dashboards ask for hours, because a bucket that spans more than the query window forces you to filter the array. The **physical** bound keeps the document small: cap by element count or estimated size so a bucket cannot approach the 16 MB BSON limit, and remember that very large documents make every update more expensive because the storage engine rewrites the document, not just the appended field. ## What you give up Single-event access is no longer a point read: you match the bucket and then filter or `$unwind` the array, and updating one sample needs the positional operators or `arrayFilters`. Deletes and TTL work at bucket granularity, so retention becomes coarser. Uniqueness on an individual measurement is harder to enforce, because a unique index constrains documents, not array elements. And a badly chosen bucket key — bucketing by time alone, mixing every sensor into one document — creates contention on a single hot document. ## Native time series collections From MongoDB 5.0 the server implements this internally. `db.createCollection("readings", { timeseries: { timeField: "ts", metaField: "sensorId", granularity: "hours" } })` gives you a collection you insert one document per measurement into and query as if it were unbucketed, while the storage layer keeps bucketed, column-oriented internals. `expireAfterSeconds` on the collection handles retention. For new time-stamped workloads that is the default answer; hand-rolled bucketing remains the right tool when the grouping is not time-based — batching many small child records under a parent, for example — or when you need the rollups visible as ordinary queryable fields. ## How to answer it Say what the pattern is in one sentence, name where the win comes from (document count, index entries, repeated field names, fewer fetches per range read), show the bounded upsert, then volunteer the costs and the 16 MB ceiling. Finishing with "and on 5.0+ I'd reach for a native time series collection first" is what separates someone who has read the pattern from someone who has run it.

  • What bounds a bucket besides the time window?
    A size bound. Carry a counter in the document and put it in the upsert filter — `count: { $lt: 200 }` — so once the bucket is full no document matches and the upsert inserts a fresh one. Bound by element count or estimated bytes so a bucket never approaches the 16 MB BSON limit, since oversized documents also make every append more expensive to rewrite.
  • How do you read one individual measurement back out of a bucket?
    Match the bucket by its key, then filter the array — `$unwind` followed by `$match` in an aggregation, or `$filter` in a projection. That is strictly more work than a point read, which is why you bucket along the axis you actually query by. If single-event lookup is a primary access pattern, bucketing is the wrong shape.
  • When would you use a native time series collection instead of hand-rolled buckets?
    For anything time-stamped on MongoDB 5.0 or later. `createCollection` with `timeseries: { timeField, metaField, granularity }` makes the server bucket internally while you insert and query single measurements, and `expireAfterSeconds` gives retention. Hand-rolled buckets stay useful when the grouping is not time-based, or when you need the rollup fields as ordinary indexable document fields.

A bucket is a shipping pallet: sending a thousand parcels individually pays the label, the barcode scan and the manifest line a thousand times; palletizing them by destination and hour pays it once.

saying these in an interview costs you the question

  • Says bucketing is only about saving disk space
  • Lets the samples array grow with no count or size bound
  • Claims bucketing makes writes across several sensors atomic
  • Thinks bucket documents are exempt from the 16 MB limit
  • Buckets by time alone, ignoring the metadata key

context