skip to content

How do you keep a precomputed count stored inside a document from drifting away from the rows it summarises?

level: seniorimportance: should knowfreq 52%

answer

  1. two writes where there used to be one
  2. never read-modify-write in application code
  3. retries are the enemy of a blind delta
  4. store what the number covers
  5. some numbers may never be merely approximate

basics

~20 s

Increment it atomically within its own document, make the increment idempotent so retries cannot double-count, store a marker of what has been counted, and run a periodic recompute from the source that detects and repairs drift instead of assuming none.

solid answer

~50 s

An embedded counter drifts because the increment and the event it counts live in different documents, so any crash, retry or partial failure between them changes one and not the other. Three things keep it honest. First, do the increment as an in-place atomic operation on the single document that holds it, never read-modify-write in application code, or concurrent updates will lose each other. Second, make it idempotent: a retried message must not count twice, which usually means recording the id of what was counted or deriving the increment from a state transition that can only happen once. Third, accept that drift still happens and build the reconciliation — a scheduled recompute from the source that compares and repairs, plus a stored `countedThrough` marker so you know what the number covers. If the count is only decorative, saying "approximate, repaired nightly" is a perfectly good answer.

code

json · 7 lines
json
{
  "_id": "post-501",
  "commentCount": 41,
  "countedEventIds": ["evt-9a", "evt-9b", "evt-9c"],
  "countedThrough": "2026-08-19T10:02:00Z",
  "recomputedAt": "2026-08-19T03:00:00Z"
}

go deeper

for a junior

Know that a stored count is a second copy of information that must be updated separately, and that adding one in application code after reading the document races with concurrent updates.

for a middle

Explain in-place atomic increments within a single document, why the separate event write makes retries dangerous, and what idempotency looks like in practice.

for a senior

Demonstrate the full loop: idempotent increments, a stored coverage marker, a scheduled recompute that reports how many documents it repaired, and a refusal to enforce inventory or balances from a rollup.

for a principal

Own the classification across the product — which numbers may be approximate, which must be exact, and what the organisation promises about repair latency for each — and make sure no new write path can bypass the counter unnoticed.

## Why counters drift at all A precomputed rollup — a comment count on a post, a total on a cart, a per-day summary embedded in a user document — exists so a read does not have to scan or aggregate the underlying data. The moment you store it, the number and the thing it counts live in different places, and the system has to make two changes where it used to make one. Anything that can happen between those two changes causes drift: a process crash, a retried message, a failed second write, a delete that skipped the decrement path, a bulk import that bypassed the application entirely. Drift is silent and cumulative. Nothing errors; the number is just a little wrong, then more wrong, and eventually a user notices a post claiming 41 comments that shows 39. ## Make the increment atomic within its document The first requirement is that concurrent increments do not lose each other. Reading the document, adding one in application code and writing the whole document back is a lost-update race: two concurrent handlers both read 41 and both write 42. Document stores give you atomic in-place update operators that apply the delta on the server within the single document, and that is what should be used. Because the counter lives inside one document, that single-document atomicity is enough for the counter itself — no multi-document transaction is needed to protect the number from concurrent updates. ## Make the increment idempotent Single-document atomicity does not solve the harder problem: the increment is a *separate* write from the insert of the comment. At-least-once delivery, client retries after a timeout, and job reruns all mean the same event can be processed twice. An unconditional delta is not safe under retry. The common defences are: - **Record what was counted.** Add the event id to a bounded set inside the document and only increment if it was not already present, so a replay is a no-op. This is exact but the set grows, so it needs a horizon. - **Count a state transition instead of an event.** Increment only as part of the write that flips a field from one value to another; if the transition already happened, no increment occurs. This is the cleanest option when it fits. - **Derive rather than accumulate.** If the underlying items are stored in a bounded array within the same document, the count is a function of that array and can be recomputed in the same atomic update — drift becomes structurally impossible. This only works while the array stays bounded. ## Reconcile on a schedule, and say what the number means Even with the above, imports, manual fixes and bugs will diverge the number eventually. Mature designs treat the counter as a cache of an aggregate and run a periodic recompute: aggregate the source, compare, repair, and — critically — *report* how many documents were wrong. That number is a quality signal; a sudden rise in repairs means a code path started bypassing the counter. Store a marker with the counter saying what it covers: a `countedThrough` timestamp or source version and a `recomputedAt` stamp. Without it, a reader cannot tell a correct 41 from a stale 41, and a recompute cannot safely run concurrently with live increments. ## Decide the accuracy the number actually needs Not every rollup deserves the same rigour. Classify: - **Decorative** (view counts, reaction totals): approximate is fine, repair nightly, and cheap approximate accumulation is acceptable. - **Operational** (unread badges, items in cart): should be right, but a rare off-by-one repaired within minutes is survivable. - **Authoritative** (balances, remaining inventory, quota enforcement): must never be trusted as a cache. Enforce these against the owning record inside a single-document conditional update, or against the source itself — never against a rollup that a reconciliation job repairs later, because by then you have oversold. Saying that last part out loud is what separates a senior answer from a mechanical one: the interesting question is not how to increment, it is which numbers are allowed to be wrong for a while. ## Rollups beyond counters The same reasoning covers embedded sums, min/max, last-N previews and per-period buckets. Bucketed rollups (a document per user per day holding accumulated totals) keep documents bounded, make recomputation cheap because only the affected bucket is rebuilt, and confine contention to the current bucket rather than one perpetually hot document. ## How to answer Walk the three layers — atomic in-place increment, idempotency against retries, scheduled reconciliation with a coverage marker — then close with the classification: decorative numbers may drift and be repaired; authoritative numbers must be enforced at the source, not read from a rollup.

  • Why is a blind increment unsafe even when the update itself is atomic?
    Atomicity protects the number from concurrent updates, not from being applied twice. The comment insert and the increment are separate writes, so a client timeout, a retried message or a rerun job can apply the same increment again. Safety requires idempotency: record the event id in the document and skip if present, or tie the increment to a state transition that can only occur once.
  • Where would you refuse to use a precomputed count at all?
    Anywhere the number authorises an action — remaining inventory, an account balance, a quota check. A rollup repaired by a nightly job is wrong for hours, and by then you have oversold or let a customer overdraw. Enforce those against the owning record with a conditional update that fails when the invariant would break, so the check and the change are one atomic act.
  • How does bucketing a rollup by time period help?
    Storing per-user-per-day totals keeps each document bounded, spreads writes across many documents instead of one perpetually hot one, and makes reconciliation cheap because only the affected bucket is recomputed rather than a lifetime total. Reads sum a small number of buckets, and old buckets become immutable, so they can be trusted without further repair.
  • What should the reconciliation job report besides fixing numbers?
    How many documents it had to repair, and how large the deltas were. That is a code-quality signal: a steady trickle is expected, a sudden jump means some new path is inserting source records without going through the counter update. Silent repair hides the bug that caused the drift, so the metric matters as much as the fix.

saying these in an interview costs you the question

  • Reads the document, adds one, and writes the whole document back
  • Assumes an atomic increment is automatically safe under retries
  • Trusts a precomputed count for inventory or balance checks
  • Has no reconciliation job because 'the increment cannot fail'
  • Stores the count with no marker of what it covers

context