In Domain-Driven Design, when designing an aggregate's boundary, what heuristics determine how many entities/value objects it should include, and why do most experienced DDD practitioners favor small aggregates?
answer
- true invariant test
- boundary = transaction = lock
- small aggregate, favor eventual consistency elsewhere
- reject 'belongs to' OO intuition
- optimistic locking contention grows with aggregate size
basics
~20 sInclude only what must change together, at the same instant, to keep one business rule true. Keep aggregates small: fewer things bundled together means fewer collisions when different requests try to change the same aggregate at once.
solid answer
~40 sDraw the aggregate boundary around exactly the data that must be consistent within a single transaction — no more. The test is: "if this rule were violated, would it be a real business problem right now, or could the system tolerate it being fixed shortly after?" True invariants (e.g., an order's total must equal the sum of its lines) belong inside one aggregate; everything else (e.g., "an order shouldn't exceed the customer's credit limit") can be enforced via a domain service or eventual consistency across aggregates. Favor small aggregates — often a single entity plus value objects — because every aggregate is typically also a concurrency/locking unit: the bigger it is, the more unrelated use cases contend for the same lock/version, causing optimistic-locking failures and reduced throughput under load.
go deeper
Can describe that an aggregate groups related data but likely hasn't yet connected boundary size to transaction/locking cost.
Applies the true-invariant test to distinguish what belongs inside an aggregate versus what should be a separate aggregate coordinated via a service or event.
Actively designs for small aggregates, anticipates optimistic-locking contention, and can point to unbounded-collection anti-patterns from experience.
Sets team-wide conventions/review standards for aggregate sizing, weighs eventual-consistency operational cost against contention risk for the specific domain, and can justify exceptions.
## Why the boundary is the hard part Deciding what belongs inside an aggregate boundary is arguably the hardest and most consequential decision in tactical DDD, because the boundary you draw becomes, in most implementations, simultaneously a **consistency boundary**, a **transaction boundary**, and a **concurrency/locking boundary**. Get it too big and you create contention and unnecessary coupling; get it too small and you lose the ability to enforce a genuine business rule atomically. The primary heuristic is: **an aggregate should contain exactly the data that must be true together, right now, for the business rule to hold — and nothing else.** Vaughn Vernon, in his widely cited "Effective Aggregate Design" series, frames this as designing aggregates around **true invariants**, meaning rules that must never be false even for an instant as observed by any other transaction. - A canonical example: within an `Order`, the sum of `OrderLine` amounts must equal the Order's total — if that's ever inconsistent, even momentarily, downstream billing breaks. That rule justifies bundling `Order` and its `OrderLines` into one aggregate. - Contrast that with "an order total should not exceed the customer's available credit" — this looks similar but spans two aggregates (`Order` and `Customer/Account`), and DDD's guidance is explicitly not to merge them for this; instead you either enforce it via a **domain service** that reads both at the time of the request (accepting that credit might change between the check and the commit), or accept it as an **eventually-consistent** rule enforced shortly after the fact, with compensating action if violated. ## Drawing the line, one step at a time The step-by-step process for drawing a boundary usually looks like: - (1) identify a candidate **root entity** that has clear identity and a lifecycle (`Order`, `Cart`, `Shipment`); - (2) list the invariants that entity must uphold; - (3) for each invariant, ask which other data must be read and written atomically to guarantee it; - (4) pull only that data inside the boundary as **child entities** or **value objects**; - (5) for every other entity that seems related but isn't required by a true invariant, push it out to its own aggregate and reference it by ID. A common trap is following object-oriented "belongs to" intuition — because `OrderLines` "belong to" an `Order` conceptually and `Customer` "belongs to" an `Order` relationally in a diagram, it's tempting to include everything reachable by navigation. DDD deliberately rejects this: **aggregate boundaries follow transactional necessity, not conceptual ownership or foreign-key structure in a relational schema.** ## The case for keeping them small Why practitioners so strongly favor small aggregates comes down to two production-facing costs. 1. **First, concurrency**: most implementations use **optimistic concurrency control** (a version number incremented on every save) at the aggregate level, meaning two requests that touch the same aggregate instance concurrently will have one fail and need to retry. A large aggregate — say, a `Customer` that embeds every `Order`, every `Address`, every `Payment` method as child entities — turns unrelated operations (adding an address, placing an order) into contenders for the same lock, causing avoidable `OptimisticLockException` failures and retries under real traffic, even though the two operations have nothing to do with each other from a business-rule standpoint. 2. **Second, cost of loading/persisting**: an aggregate is typically loaded and saved as a whole graph; a large aggregate means every mutation — even a single-field change — pulls and rewrites a large object graph, which is wasteful in both memory and I/O, and gets worse as the "big" aggregate's collections grow unbounded (a `Customer` with ten years of `Orders` embedded is a classic anti-pattern: the aggregate never stops growing). ## What the small ones charge in return The trade-off on the other side is that small aggregates push more coordination work into the application layer and into eventual consistency. If `Order` and Customer's loyalty points are separate aggregates, then "placing an order awards loyalty points" cannot be a single atomic database transaction — it typically becomes an `OrderPlaced` **domain event** that a separate handler consumes to update the `Customer` aggregate afterward, in its own transaction. - This means there's a window, however brief, where the order exists but points haven't been awarded yet, and the system (and the team building it) must be comfortable reasoning about that window and handling failure/retry in the handler. - Teams unused to this discipline sometimes "solve" the discomfort by widening the aggregate back out — recreating the concurrency and load problems above — rather than accepting eventual consistency for genuinely non-critical cross-aggregate rules. ## The shape a mature e-commerce model settles on A well-known real-world illustration is the shopping cart / order split used by many e-commerce reference architectures: `ShoppingCart` is deliberately its own small, short-lived aggregate distinct from `Order`, and `Order` itself is kept lean (line items and totals) while `Payment`, `Shipment`, and `Customer` loyalty are separate aggregates coordinated through events — this is precisely the **"favor small aggregates, reference the rest by ID, coordinate via events"** pattern DDD practitioners recommend, and it's the same shape recommended in Vernon's rule of thumb: model true invariants in consistency boundaries, favor small aggregates.
- A junior engineer proposes embedding all of a Customer's past Orders as child entities inside the Customer aggregate so 'everything is in one place.' What's wrong with that?It creates an unbounded aggregate that keeps growing forever, meaning every load/save of Customer pulls and rewrites an ever-larger object graph, and it puts unrelated operations (editing a profile field vs. placing a new order) in contention for the same optimistic-lock version. Orders should be their own aggregate referencing CustomerId, with a separate query/read model used to list a customer's order history.
- How do you enforce a rule like 'an order total should not exceed the customer's credit limit' if Order and Customer are separate aggregates?You typically use a domain service invoked by the application layer that loads both aggregates, performs the check at the moment of the request, and rejects the operation if it fails — accepting that this check is a point-in-time read, not an atomic guarantee, since credit limit could change between check and commit. Some teams instead treat it as eventually consistent: allow the order, then flag or compensate if a later reconciliation finds it violated.
- What symptom in production logs often signals that an aggregate has been drawn too large?Frequent optimistic-locking or version-conflict exceptions and retries on that aggregate's save path, especially between operations that a domain expert would say are unrelated (e.g., updating a shipping address failing because of a concurrent unrelated order placement on the same Customer aggregate). That pattern indicates the boundary should be split.
Like deciding how many people must be in the same room to sign a contract at once — only include the ones whose signatures must be simultaneous; everyone else can sign a related but separate document later.
saying these in an interview costs you the question
- designs aggregate boundaries by following foreign-key/ownership shape in the database schema
- includes an unbounded growing collection (all historical orders) inside a parent aggregate
- cannot distinguish a true invariant from a business rule that can tolerate brief inconsistency
- assumes every related entity must be inside the same aggregate for the app to 'make sense'
- proposes widening an aggregate specifically to avoid dealing with eventual consistency