What is MongoDB's computed pattern, and when does precomputing a total beat aggregating at read time?
answer
- do the work once, at write time
- reads vastly outnumber writes
- store the components, not the average
- $inc on one document is atomic
- have a job that recomputes from source
basics
~20 sThe computed pattern stores the result of a calculation — a count, sum, average or rollup — in the document and maintains it on write, so reads fetch a value instead of recomputing it. It pays off when reads far outnumber writes.
solid answer
~50 sRather than aggregating millions of source documents on every page load, you compute once at write time and store the answer where it is read. A product carries `ratingCount` and `ratingSum`, each new review does `$inc: { ratingCount: 1, ratingSum: 4 }`, and the page renders the average by division from one point read. Store the **components** (sum and count) rather than the average itself, because components are incrementally updatable and stay exact. The pattern applies at two scales: in-document counters maintained by `$inc`, which are atomic because they touch a single document; and periodic rollups written into a summary collection with `$merge`, an on-demand materialized view refreshed on a schedule. Its cost is a second copy of the truth, so the counter can drift if a write path forgets to update it or a cross-document update fails halfway. Any serious use of the pattern ships with a recompute job that re-derives the value from the source data.
code
javascript · 11 lines// maintain components on write
db.products.updateOne(
{ _id: "p-991" },
{ $inc: { ratingCount: 1, ratingSum: 4 } }
)
// periodic rebuild from the authoritative source
db.reviews.aggregate([
{ $group: { _id: "$productId", ratingCount: { $sum: 1 }, ratingSum: { $sum: "$rating" } } },
{ $merge: { into: "productStats", on: "_id", whenMatched: "replace", whenNotMatched: "insert" } }
])go deeper
Recall the idea: keep counts, sums and totals in the document and update them when data changes, so the read is a single fetch rather than a calculation.
Explain the read/write ratio argument, why components beat storing a finished average, and that $inc within one document is atomic while an update spanning two documents is not.
Demonstrate the operational half: reconciliation jobs, $merge-based materialized rollups, and the judgment call between an eventually-repaired display value and an invariant that needs a transaction.
Own where derived state should live across the system — in-document counters, summary collections, or a downstream store — and set the policy for staleness, rebuild cost and who owns the reconciliation.
## The trade being made Most application workloads are read-heavy by a wide margin. If a dashboard shows "4,312 reviews, average 4.4" and the underlying reviews collection has millions of documents, computing that on every render means an aggregation over an index scan of thousands of entries — repeated thousands of times per minute for data that changes a few times an hour. The computed pattern moves that work to the write, where it happens once, and stores the answer where the read already is. That is a straight ratio argument: if reads outnumber writes by a hundred to one, doing a little extra work per write to remove work from every read is obviously correct. If the ratio inverts — a counter that is written constantly and read once a day — precomputation is a hotspot, not an optimization, and you should compute on demand. ## In-document counters The simplest form keeps the computed field beside the data it summarizes: ```javascript db.products.updateOne( { _id: "p-991" }, { $inc: { ratingCount: 1, ratingSum: 4 }, $set: { lastReviewAt: now } } ) ``` Two points make or break this. First, **store components, not results**. Keeping `ratingSum` and `ratingCount` lets any new rating be folded in with a single `$inc`; keeping only `avgRating` does not, because you cannot update an average incrementally without knowing the count, and repeatedly recomputing an average from a stored average accumulates error. Second, `$inc` on one document is **atomic** — concurrent increments from many clients cannot lose an update, because MongoDB applies the modification to a single document under its own lock. That is why counters live in a document, not in application memory. The atomicity ends at the document boundary. Inserting the review into `reviews` and incrementing the counter on `products` are two writes to two documents; they are atomic together only inside a transaction. Many systems accept that and reconcile instead. ## Materialized rollups The second form precomputes across many documents into a summary collection: ```javascript db.reviews.aggregate([ { $group: { _id: "$productId", n: { $sum: 1 }, sum: { $sum: "$rating" } } }, { $merge: { into: "productStats", on: "_id", whenMatched: "replace", whenNotMatched: "insert" } } ]) ``` `$merge` writes the pipeline's output into a target collection, updating documents that already exist — MongoDB calls the result an *on-demand materialized view*. Run it on a schedule, or incrementally over the window that changed, and reads hit `productStats` directly. This is the right shape when the computation is too complex for an `$inc`, when it spans documents, or when a small staleness window is acceptable. ## Drift, and how to live with it A computed value is a second copy of a fact, so it can disagree with the source. Causes are mundane: a new write path that forgets the increment, a partial failure between the two writes, a bulk import that bypasses application code, a manual fix in the shell. Treat drift as inevitable and design for it: keep the source data authoritative, keep the derived value cheap to rebuild, and run a periodic reconciliation that recomputes from source and overwrites. If a counter is *load bearing* — inventory that gates a sale, a quota that gates access — either put it in the same document as the thing it constrains so a single atomic update enforces the invariant, or use a transaction. "It's only a display count, we repair nightly" and "it's an invariant, it must be transactional" are both good answers; not distinguishing them is not. ## Where it appears in other patterns The computed pattern rarely arrives alone. Bucket documents carry `count`, `sum`, `min` and `max` for the window so a chart never touches the samples array. Subset documents carry `reviewCount` and `avgRating` beside the embedded slice, so the summary line needs no second query. Recognizing it as an ingredient of the other patterns is a good sign in an interview. ## How to answer it Lead with the ratio ("reads vastly outnumber writes, so move the work to the write"), show the `$inc` of components, note that single-document updates are atomic while the counter-plus-source pair is not, mention `$merge` for cross-document rollups, and finish with the reconciliation job. That is the complete answer in four sentences.
- Why store ratingSum and ratingCount rather than avgRating alone?Because components fold in incrementally: a new rating is one `$inc` on each, and the average is exact division at read time. An average alone cannot be updated without knowing how many values produced it, and repeatedly averaging an average drifts. Storing both also lets you display the count, which the UI usually wants anyway.
- What keeps a precomputed counter from drifting away from the source data?A reconciliation job. Periodically re-derive the value with an aggregation and write it back — `$group` then `$merge` into the summary target — so any missed increment self-heals. Alongside it, funnel every write through one code path, and if the counter enforces a real invariant rather than decorating a page, put it in the same document as the constrained data or use a transaction.
- When is the computed pattern the wrong choice?When writes dominate reads, so you pay per write to serve a value nobody reads; when the precomputed field turns a spread-out write load into contention on one hot document; or when the value must be exactly consistent with data in another document and you are unwilling to pay for a transaction. Compute on demand in those cases.
It is the difference between a shop that counts its till after every customer asks how the day is going, and one that keeps a running total on a slip of paper — provided someone recounts at closing time to catch the slips that went missing.
saying these in an interview costs you the question
- Stores only the average and tries to update it incrementally
- Assumes counter and source update atomically across documents
- Applies it to a write-heavy, rarely-read value
- Has no job to recompute the value from source
- Thinks $inc can lose updates under concurrency