skip to content

Your hot read path and hot write path want different document shapes — how do you draw the aggregate boundary?

level: seniorimportance: should knowfreq 45%

answer

  1. Take the invariants off the table first
  2. Ratios beat opinions
  3. Ask how stale is acceptable
  4. Two honest shapes beat one compromise
  5. Every derived shape needs a rebuild path

basics

~20 s

Measure both paths, then let the dominant one own the primary aggregate and serve the other from a derived shape or extra accesses. Boundaries protect invariants first; beyond that, the heavier and more latency-sensitive path wins.

solid answer

~50 s

First separate what is negotiable. Any invariant that must hold at every instant fixes part of the boundary and is not up for trade. Beyond that, quantify: how many reads per write, what latency budget each path has, and how much staleness the product tolerates. If reads dominate heavily and the write path is tolerant, shape the primary aggregate for the read and let the writer do more work or accept a small lag. If the write path is the constrained one — high-frequency ingestion, a checkout that must not slow down — keep the write aggregate small and focused and serve the read from a second, derived shape maintained afterwards. The failure mode to avoid is a compromise shape that serves neither path well; two clear shapes with an explicit refresh path usually beat one muddled one.

code

json · 6 lines
json
// Write aggregate: one review, small and independently written
{ "_id": "rev-901", "productId": "prd-12", "stars": 4, "body": "Solid." }

// Read aggregate: product carries derived figures, refreshed after a review lands
{ "_id": "prd-12", "name": "Keyboard", "price": 4900,
  "ratingAverage": 4.2, "reviewCount": 318 }

go deeper

for a junior

Be ready to say that document shape is a trade between read convenience and write cost, and that the shape should follow whichever path the workload is actually dominated by.

for a middle

Explain the concrete options — shape for the read, shape for the write, or keep two shapes — and what each costs. Know that unbounded nesting breaks both paths regardless of the trade.

for a senior

Demonstrate the measured approach: invariants first, then read-to-write ratio, latency budgets, contention and staleness tolerance, then a committed choice with its operational cost named, including the rebuild path for any derived shape.

for a principal

Own the standing decision: when the organisation is allowed to introduce a derived shape, who runs and monitors it, and how such shapes are prevented from multiplying into a set of half-maintained copies nobody can rebuild.

## The situation A single entity is pulled in two directions. The read path wants everything in one document so a page loads in one access. The write path wants a small, focused document so writers do not collide and updates stay cheap. Both are legitimate, and no shape satisfies both perfectly. This is one of the recurring senior-level modelling arguments. ## Step one: separate the non-negotiable Before trading, take the invariants off the table. If a rule must be true whenever anyone looks, the data it mentions has to sit in one document, and no throughput argument changes that. Everything that is *not* an invariant — figures that may lag, values that are merely convenient to have nearby — is where the negotiation happens. This step matters because arguments about shape often turn out to be arguments about whether a rule is real. Settle that first and the shape frequently settles itself. ## Step two: quantify both paths Opinions about which path matters are worth little; ratios are worth a lot. Establish: - **Reads per write** for this data. A hundred to one and a one-to-one workload lead to different answers. - **Latency budgets.** A read on the critical rendering path of the main screen is constrained differently from an admin export. A write inside a checkout is constrained differently from a nightly import. - **Concurrency.** How many independent writers touch this document, and do they touch disjoint parts of it? Disjoint writers on one document still serialise. - **Staleness tolerance.** How long may a derived copy lag before someone actually suffers? Answers of "never" should be challenged; answers of "a minute is fine" open up options. - **Growth.** Does either side accumulate without bound? An unbounded accumulation inside a document eventually breaks both paths regardless of the trade. ## Step three: choose a resolution **Read-dominant, tolerant writer.** Shape the aggregate for the read. The writer does the assembly work at write time — more work per write, but writes are rare. This is the common case in content and catalogue systems. **Write-dominant or write-constrained.** Keep the write aggregate small: exactly the fields the write path touches, so writes are cheap and contention is low. Serve the read from a derived shape refreshed after the write, and accept the lag you measured as tolerable. This is common in ingestion, telemetry and high-frequency transactional flows. **Both paths hot and neither tolerant.** Split into two shapes deliberately: a compact write aggregate that owns the truth and a read-oriented shape derived from it. Now you own a refresh path, and it must be explicit — who updates the derived shape, how failures are retried, how it is rebuilt from scratch if it drifts, and how the lag is monitored. This is real operational surface, so take it only when the numbers justify it. **Neither path is actually hot.** Do the simple thing and stop. Plenty of these arguments are about data that sees a handful of operations a minute, where any shape works and the simplest one wins. ## Step four: name the owner Whatever you choose, one shape must be the source of truth. Two shapes that both accept writes and are expected to agree is the hardest thing in this whole area to operate. The derived shape should be reconstructible from the primary one by a job you can actually run, because that job is your recovery path when the refresh mechanism drops an update. ## The compromise trap The tempting middle answer is one document that is a bit denormalised for reading and a bit trimmed for writing. It often ends up worst: the read still needs a second access for the fields that got trimmed, and the write still drags along the fields that got added. When the two paths genuinely pull apart, two honest shapes with a documented refresh beat one shape that is a compromise nobody chose deliberately. ## A concrete case A product page shows a product plus an aggregate rating and a review count. Reviews arrive continuously and are written by many independent users; the product page is read constantly. Nesting reviews inside the product is the naive read-first answer, and it fails: reviews grow without bound, every new review rewrites a document that every page load reads, and writers contend. Keeping reviews entirely separate is the naive write-first answer, and it fails differently: rendering the page needs the rating, which now requires computing over reviews on every read. The resolution is neither extreme. Reviews are their own aggregate — small, independently written, unbounded in number. The product document carries a derived `ratingAverage` and `reviewCount` refreshed after a review lands. The invariant question settles the design: nobody is harmed if the rating is a few seconds behind, so it may be derived, and the boundary holds. ## How to answer this in an interview Do not pick a side immediately. Ask what the invariants are, ask for the read-to-write ratio and the staleness tolerance, then commit to a shape and say what it costs. Mention the recovery path for any derived shape you introduce. The answer interviewers are listening for is not "reads win" — it is a candidate who prices the trade, names a source of truth, and knows they have just taken on an operational job.

  • How do you keep a derived read shape from silently drifting from its source?
    Give it an owner and a rebuild job. Refresh it from an explicit path with retries, monitor the lag as a metric with an alert, and be able to reconstruct it wholesale from the source of truth on demand. Drift is inevitable over a long enough window; the question is whether you detect it and can repair it cheaply.
  • When is a compromise single shape actually the right answer?
    When neither path is hot enough for the difference to matter, or when the extra fields the read wants are small, rarely changing and few. Introducing a derived shape adds a refresh path, a lag metric and a rebuild job; if the measured benefit is marginal, that operational surface is not worth owning.
  • What if the product insists the derived value must never be stale?
    Then it is an invariant, and it must sit in the same document as the writes that change it — which means accepting the contention that brings. Present that cost explicitly. In practice, showing what 'never stale' costs in write throughput often converts the requirement into a tolerable lag measured in seconds.

saying these in an interview costs you the question

  • Picks read or write optimisation without measuring the ratio
  • Nests an unbounded collection to save one read
  • Maintains two shapes with no designated source of truth
  • Adds a derived shape with no rebuild or monitoring path
  • Treats every derived figure as needing instant accuracy

context