skip to content

Why is a single document the consistency scope in a document database, and how does that shape aggregate boundaries?

level: middleimportance: must knowfreq 62%

answer

  1. Think about what a single write covers
  2. All-or-nothing, but only that far
  3. Start from the invariant list
  4. Ask what must never be seen broken
  5. Same boundary decides who contends

basics

~20 s

A write to one document lands entirely or not at all, so one document is the widest scope where an invariant can be enforced for free. Any rule that must always hold should sit inside a single aggregate; rules spanning documents need extra machinery or must tolerate lag.

solid answer

~50 s

Document stores give all-or-nothing behaviour over one document as a basic property of the storage engine — a reader never sees half an update. That makes the document the natural **consistency scope**: any rule that must never be observed broken has to live inside one. So the design move is to ask, for each invariant, *what data must change together in one indivisible step?*, and let that set the boundary. An order total that must always equal the sum of its lines argues for nesting the lines. A customer's lifetime-spend figure that may lag by a minute does not: it can live in the customer document and be refreshed afterwards. Invariants that genuinely span aggregates cost you something extra — a multi-step write with compensation, or an accepted window of divergence — and that cost is exactly what a good boundary minimises.

go deeper

for a junior

Be ready to state that a write to one document is all-or-nothing, and that anything which must stay consistent together should therefore live in the same document.

for a middle

Explain the invariant-driven boundary: list what must never be observed broken, check whether it fits in one document, and describe the three options when it does not.

for a senior

Show judgment about which rules are genuinely invariants and which can lag, and connect the boundary to write contention — a document every request in a hot flow updates is a queue, however correct it is.

for a principal

Own the policy: where the system spends coordination, which invariants justify concentrated contention, and how you stop teams reaching for cross-document coordination as the default fix for a boundary drawn in the wrong place.

## Why the document is the boundary A document database updates a document as a unit. Whatever the engine does internally, the guarantee it exposes is the same: a concurrent reader sees the document either as it was or as it became, never a half-applied mixture. That single property is what makes the document a **consistency scope** — the largest region of data over which you can enforce a rule with no extra machinery. This is not an incidental detail. It is the reason aggregate-oriented modelling exists. If the free unit of atomicity were a whole database, boundaries would be a performance question only. Because the free unit is one document, boundaries decide *what can be guaranteed*. ## Invariants and where they force the line An **invariant** is a statement that must be true whenever anyone looks — not eventually true, but true at every observable moment. "An order's `total` equals the sum of its line amounts." "A booking's `seatsTaken` never exceeds `capacity`." "An account's status is `closed` only if its balance is zero." For each invariant, ask which data it mentions. If all of it fits in one document, the invariant is enforceable: read the document, compute the new state, write it back in one operation, and no observer ever sees a violation. If the invariant mentions data in two documents, no single write covers it, and you are left with three options: 1. **Move the boundary** so both pieces live in one document. This is the cheapest fix when the pieces are small and bounded. 2. **Accept lag.** Reclassify the rule from an invariant to a target: the two documents may disagree briefly and a follow-up write reconciles them. Most reporting figures, counters and cached labels are legitimately in this class. 3. **Add machinery.** Multi-document transactions where the platform offers them, or an application-level sequence with compensation. Both work; both cost latency, complexity and — with transactions — contention, and neither is free enough to be the default. The skill is in deciding which invariants are real. Product people describe most rules as absolute. Many are not: nobody notices if a "posts written" counter is a second behind. A boundary drawn to protect a rule that did not need protecting is a boundary that made every read heavier for nothing. ## Concurrency lives here too The consistency scope is also the contention scope. Two writers touching the same document serialise against each other; two writers touching different documents do not. That cuts both ways. A large aggregate makes more rules enforceable and more writers collide — the classic symptom is a document that every request in a hot flow updates, turning a distributed system into a queue behind one record. A small aggregate spreads writers out but shrinks what you can guarantee. So the boundary decision has a concurrency column as well as a consistency column: *which rules must hold at all times*, and *who writes to this document, how often*. A rule that forces two hot, independent writers onto one document is a rule worth challenging. ## A worked example A seat-booking service must never oversell a screening. `seatsTaken <= capacity` is a genuine invariant — overselling is a refund and an apology. Both values must therefore live in one screening document, and every booking updates that document, conditionally, so a concurrent booking cannot slip past the check. The contention is accepted because the rule is real. The same service shows "tickets sold this week" on an internal dashboard. That number touches every screening. It is *not* an invariant: a dashboard a minute stale is fine. It stays outside the booking aggregate, computed or accumulated separately, and the booking path pays nothing for it. The same reasoning covers the customer's saved payment methods. They are read on the checkout page but no rule ties them to a booking's correctness, so they are a separate aggregate reached by identifier — the checkout flow issues a second read rather than dragging the payment methods into every screening write. ## Common mistakes The first is assuming that because the engine offers cross-document transactions, boundaries stop mattering. They still do: transactions have a latency and contention cost, and reaching for them routinely usually signals that the aggregates were drawn against a diagram rather than against the rules. The second is the reverse — treating every desirable property as an invariant and letting documents grow until they contain half the domain. Ask what actually breaks if the two values disagree for a second. If the answer is "a chart looks slightly off", it is not an invariant. The third is forgetting that a boundary protects reads too. If a rule holds inside a document, any reader of that document sees a coherent state without coordination. That is a large part of why single-document reads are so easy to reason about. ## How to answer this in an interview Lead with the guarantee — one document, all-or-nothing — then say the consequence: the invariant list, not the entity list, draws the boundary. Give one real invariant you would nest for and one pseudo-invariant you would deliberately let lag, and mention that the same boundary is the contention boundary. That combination is what separates a memorised slogan from someone who has drawn these lines on a real system.

  • How do you decide whether a business rule is a real invariant or something that may lag?
    Ask what a user or the business actually experiences if the two values disagree for a few seconds. Overselling a seat means a refund; a dashboard figure being stale means nothing. Real invariants have a concrete failure with a cost attached. Anything whose worst case is 'a number looks slightly off' is a target, not an invariant, and should not distort the boundary.
  • If a rule spans two aggregates, what are the options?
    Move the boundary so both sides sit in one document; accept a divergence window and reconcile with a follow-up write; or coordinate explicitly with a multi-step write that can compensate on failure. Choose by how bad a temporary violation is and how hot the write path is — coordination costs latency and contention on every write, not just on the failing ones.
  • Doesn't a large aggregate that protects many rules hurt throughput?
    Yes — the consistency scope is the contention scope. Every writer of any part of the document serialises with every other. If two unrelated hot flows both update one large document, you have made the boundary a bottleneck. That is a reason to challenge whether all those rules truly need enforcing at every instant.

The document is the sheet of paper you can rewrite in one stroke. Anything written on that sheet can be kept mutually consistent for free; anything on another sheet needs you to put down the pen and pick it up again, and someone may read the other sheet in between.

saying these in an interview costs you the question

  • Assumes multi-document transactions make boundaries irrelevant
  • Treats every business rule as an absolute invariant
  • Ignores that a wide aggregate concentrates write contention
  • Believes a reader can observe a half-applied document update
  • Draws boundaries from entity relationships, never from rules

context