skip to content

You're designing the aggregates for an event-sourced order-management system with CQRS. How do you decide where to draw each aggregate's consistency boundary, and what's the trade-off between making an aggregate bigger, to get more invariants enforced for free inside one transaction, versus splitting it and coordinating across the split via a process manager?

level: principalimportance: should knowfreq 45%

answer

  1. aggregate boundary = one event stream = one optimistic-concurrency unit
  2. true invariant stays in one aggregate, eventually-true rule goes in a separate aggregate plus saga
  3. bigger aggregate = simpler invariants but more write contention, the hot-aggregate problem
  4. smaller aggregate = better concurrency but needs an explicit saga and compensation
  5. Vernon's effective-aggregate-design heuristic: model true invariants

basics

~20 s

An aggregate should be just big enough to enforce the rules that must be true at every single instant. Anything that can tolerate a short delay to become consistent should live in a separate aggregate, coordinated afterward by a saga.

solid answer

~40 s

An aggregate boundary equals a transactional consistency boundary: one aggregate instance maps to one event stream, and only one command can succeed against it at a time via optimistic concurrency, so it inherently serializes writes to that instance. The design heuristic is to draw the boundary around true invariants, rules that must never be violated even momentarily, and to put anything that can become correct slightly after the fact into a separate aggregate coordinated by a saga. Bigger aggregates make invariant enforcement trivial but increase write contention, since every command against that instance competes for the same version; smaller aggregates improve concurrency and throughput but require explicit saga logic, with its own compensation and idempotency cost, for any rule spanning multiple aggregates.

go deeper

for a junior

Understands that an aggregate groups related data that changes together.

for a middle

Can explain that one aggregate equals one transaction and can propose a reasonable boundary for a simple example.

for a senior

Can identify hot-aggregate contention from production symptoms and correctly resize or split a boundary, introducing a saga where needed.

for a principal

Sets the true-invariant-versus-eventually-true-rule heuristic across teams, owns the org-level trade-off between contention and coordination complexity, and reviews aggregate boundary decisions before they calcify into hard-to-migrate event schemas.

## The aggregate is the transactional boundary In domain-driven design terms adopted directly by Event Sourcing, an aggregate is the **transactional consistency boundary**: one aggregate instance corresponds to one event stream, and every command against it is checked with optimistic concurrency against an expected stream version, meaning only one command can successfully append to a given instance at a time. That single fact drives the design heuristic, articulated in depth in Vaughn Vernon's guidance on effective aggregate design: - model **true invariants**, business rules that must hold at every instant with no window where they can be false, inside a single aggregate; - anything that merely needs to **eventually be true**, and can tolerate being corrected shortly after a brief inconsistency, belongs in a separate aggregate coordinated through domain events and, where multiple steps are involved, a saga. ## Applied to order management Applied concretely to order management: a rule like 'an order line's quantity times its unit price must equal its recorded subtotal' is a true invariant, cheaply enforceable inside a single Order aggregate, since it only ever needs data already inside that aggregate. A rule like 'total inventory reserved across every order for a given SKU must never exceed warehouse stock' looks at first like it wants one combined Order-and-Inventory aggregate, but that would force every order in the entire system touching that SKU to serialize through one event stream. Instead, each SKU is modeled as its own Inventory aggregate holding its own reservation count, and Order coordinates with it via a saga that issues a `ReserveStock` command; for a brief window the system tolerates an order believing a reservation is pending while a concurrent order is racing for the same stock, resolved because the Inventory aggregate itself rejects a reservation that would push it negative, with the saga compensating on the Order side if that happens. ## The trade-off is symmetric The trade-off is symmetric. | Boundary | What you gain | What you pay | |---|---|---| | **A larger aggregate** | Its advantage is invariant simplicity: no explicit coordination code, no compensations, since one transaction guarantees the rule. | Its cost is contention: every command touching that aggregate instance serializes through one stream and one version counter, so throughput is capped by how many commands can pass through sequentially, producing a 'hot aggregate' problem, for example one enormously popular product's Inventory aggregate becoming a bottleneck during a flash sale, plus a growing event stream that makes replay and snapshotting progressively slower. | | **A smaller aggregate** | Its advantage is concurrency: independent instances scale and can be processed in parallel with no contention between unrelated entities, and streams stay short. | Its cost is coordination: any business rule spanning aggregates now needs an explicit saga, with the real engineering overhead of a state machine, compensating actions, idempotent command handling, and monitoring, plus a window of temporary cross-aggregate inconsistency the business must explicitly accept. | ## Getting the boundary wrong Getting the boundary wrong shows up differently in each direction. - **Modeling too coarsely**, for instance one Order aggregate holding an entire customer's full order history, produces frequent optimistic-concurrency conflicts and retried or failed commands under any meaningful concurrent load, visible in production as elevated error and retry rates concentrated on that one aggregate type. - **Modeling too finely** without building the necessary saga produces silent invariant violations instead, such as inventory going negative because nothing actually coordinates the cross-aggregate rule at all; this is a worse failure because it's data corruption rather than a visible, retryable error. ## The heuristics behind it This is the central design question addressed by Vernon's aggregate-design heuristics -- model true invariants, reference other aggregates only by identity rather than by object reference, and favor eventual consistency between aggregates by default -- and it's exactly why large e-commerce platforms typically model inventory per SKU, sometimes even per SKU per warehouse, as aggregates entirely separate from Order, using saga-based reservation workflows rather than one mega-aggregate that would serialize every order in the system through a single stream. ## Changing a boundary later In practice, teams rarely get the boundary perfectly right on the first design pass, and that has a cost specific to Event Sourcing: because events are immutable and the aggregate boundary is encoded in which stream an event lives on, splitting or merging aggregates after the fact is a genuine migration, not a quick refactor. It typically means writing a one-time transformation that reads the old, coarser stream and re-partitions its events into new, per-instance streams under a new identity scheme, then cutting every command handler and every projector over to the new boundary together, often behind a blue-green-style rollout so in-flight traffic isn't split across old and new boundaries mid-migration. That migration cost is exactly why it pays to load-test a candidate boundary against a realistic concurrency profile, for example simulating flash-sale-level concurrent orders against a single popular SKU, before committing to it in production, rather than discovering the hot-aggregate problem only after real traffic exposes it. ## Reference other aggregates by identity A related principle worth naming explicitly is that aggregates should reference each other only by identity, such as a plain SKU id stored on the Order aggregate, never by holding a live in-memory reference to another aggregate's full object graph. This isn't a stylistic preference: it's what keeps the two aggregates independently loadable, independently versioned, and independently scalable in the first place. An Order aggregate that held a direct reference to a full Inventory aggregate object would implicitly need both to be loaded and kept consistent together, quietly reintroducing the same coupling and contention that splitting them was meant to avoid.

  • What is a 'hot aggregate' and how does it show up as a production incident?
    It's an aggregate instance that receives many concurrent commands, such as one especially popular product's stock aggregate during a flash sale. Because only one command can win the optimistic-concurrency race against that instance at a time, the rest fail and retry, showing up as a spike in command retries and elevated latency specifically for that one entity while the rest of the system stays healthy.
  • How would you tell from production signals alone whether your aggregate boundaries are drawn too coarsely versus too finely?
    Too coarse typically shows up as frequent optimistic-concurrency conflicts and retries concentrated on one aggregate type even under fairly normal load. Too fine typically shows up as business-invariant violations slipping through, such as data that was supposed to be impossible actually occurring, because no saga was built to enforce the cross-aggregate rule, or as a growing pile of ad hoc code trying to patch gaps that a proper saga should own.

Like deciding who must sign off in the same room before an action is final versus who can just be told about it afterward and asked to adjust: rules that must never be broken even for an instant need everyone in the same room, one aggregate and one transaction; rules that only need to end up correct can be handled by a note passed around afterward, a saga.

saying these in an interview costs you the question

  • treats aggregate size as an arbitrary or stylistic choice rather than tied to true invariants
  • defaults to one aggregate per database table with no reasoning about invariants
  • doesn't recognize optimistic-concurrency conflicts as a signal that an aggregate is too large
  • assumes a cross-aggregate invariant can just be checked in application code with no saga or compensation strategy at all
  • never mentions that splitting aggregates requires the business to accept a window of temporary inconsistency

context