skip to content

In DDD, why is the rule 'one repository per aggregate root' - not one per entity, not one per database table - and what breaks when a team violates it?

level: middleimportance: must knowfreq 78%

answer

  1. aggregate root = only entry point
  2. repository per aggregate type, not per table
  3. invariants live at the root
  4. coarse load/save vs narrow update
  5. non-root repository = bypass risk

basics

~20 s

You only build a repository for the 'root' object of a cluster of related objects (like an Order, not its individual OrderLine items), because that root is the only safe entry point - going around it to touch the inner pieces directly can leave data in a broken, inconsistent state.

solid answer

~40 s

An aggregate is a cluster of objects (entities and value objects) treated as one consistency boundary, with a single designated aggregate root controlling all access and enforcing invariants across the cluster. Repositories exist to fetch and persist whole aggregates, so the rule is: exactly one repository per aggregate root, keyed by the root's identity. Giving a repository to an internal entity (e.g., an OrderLineRepository) lets code load or mutate that entity independently of its root, bypassing the invariants the root exists to protect - two line items could be changed in ways that violate a rule the Order aggregate is supposed to enforce as a whole. It also multiplies transactional surfaces: instead of one aggregate loaded/saved atomically, you now have partial updates spread across calls.

go deeper

for a junior

Should know the basic rule (one repository per aggregate root) and be able to name the aggregate root vs. child entities in a simple example.

for a middle

Should explain why bypassing the root via a child-entity repository breaks invariants, with a concrete example of an invariant that could silently break.

for a senior

Should be able to design aggregate boundaries and repository interfaces together, and reason about the load/save cost trade-off that comes from keeping the whole aggregate as the atomic unit.

for a principal

Should be able to diagnose and fix architectural erosion (accumulated non-root repositories) across a codebase, and set boundary-referencing conventions (reference by ID vs. embed) that keep aggregates - and therefore repositories - appropriately sized.

## Why the rule follows from what an aggregate is The 'one repository per aggregate' rule follows directly from what an aggregate is. In DDD, an **aggregate** is a cluster of associated objects — entities and value objects — that DDD treats as one unit for the purpose of data changes, with a single **aggregate root** as its only externally-addressable member. The root is responsible for enforcing every invariant that spans the cluster: an `Order` aggregate root might enforce - 'total must equal the sum of line items', - 'cannot add a line item to a cancelled order', - or 'shipping address is required before the order can transition to Placed'. Everything inside the boundary — `OrderLine` entities, a `Money` value object, a `ShippingAddress` value object — is only supposed to be reached and mutated through the root's methods, never directly. Given that, the natural unit of persistence access is the aggregate root, and repositories exist specifically to load and save aggregate roots (and by extension, everything nested inside them) as a single, atomic, invariant-respecting whole. There is exactly one repository per aggregate type because there is exactly one thing being loaded/saved: the root, with its internals coming along for the ride. ## Contrast with the per-table data-access pattern Contrast this with the more conventional per-table or per-entity data-access pattern common outside DDD, where each database table typically gets its own DAO/repository — an `OrderRepository`, an `OrderLineRepository`, a `ShippingAddressRepository` — each independently queryable and independently writable. That structure maps naturally onto a relational schema, but it **destroys the consistency boundary** DDD is trying to establish. If `OrderLineRepository` lets application code fetch and update a line item without going through the `Order` aggregate root, then nothing stops that code from changing a quantity in a way that leaves `order.total` out of sync, or from adding a line item to an order that business rules say should be closed. The invariant enforcement that the aggregate root exists to provide gets silently bypassed, and the bypass is invisible in the type system — callers innocently call `orderLineRepository.update(line)` with no compiler or runtime signal that they've skipped the root's validation. ## The trade-off The trade-off this rule accepts is **coarser-grained loading and saving in exchange for consistency**. - **Loading** an entire `Order` aggregate to change one line item's quantity means fetching the root and every child the aggregate boundary includes, even if only one field actually changes — more data moved per operation than a narrowly-targeted `UPDATE order_line SET quantity = ? WHERE id = ?` would require. - **Saving** typically means re-persisting the whole aggregate (or at least the parts that changed) in one transaction, which is why aggregate boundaries are deliberately kept small in well-designed DDD models. A large aggregate with dozens of children makes 'one repository, one atomic load/save' expensive, and pressure builds to either shrink the aggregate (referencing related aggregates by ID instead of embedding them) or relax the rule for performance, both of which are legitimate responses, but relaxing the rule is the one that reintroduces the consistency risk the rule was protecting against. ## Failure modes Failure modes from violating this rule are common and often subtle. 1. **The most direct** is exactly the bypass scenario above: a repository for a non-root entity lets code mutate that entity in isolation, and an invariant the root was supposed to guard quietly breaks — a classic version is two concurrent requests each independently updating different line items of the same order through separate line-item repositories, with neither request ever loading (or re-validating) the order total, so the total silently drifts out of sync with the sum of lines. 2. **A second** failure mode is transactional: if 'saving an order' now means three separate repository calls (order header, then each line item, then the shipping address) instead of one atomic save of the whole aggregate, a failure partway through — a dropped connection after the header save but before the line items — leaves the aggregate in a state that was never valid in the domain model's terms, a state the aggregate root's own invariants would have rejected had it been asked to validate the whole thing at once. 3. **A third**, subtler failure is architectural erosion: once one non-root repository exists 'just for this one query', it becomes precedent, and over time the aggregate boundary stops being enforced anywhere in code even though the documentation and diagrams still describe one. ## A concrete example A concrete example: an `Order` aggregate contains `OrderLine` entities and a `ShippingAddress` value object, all reachable only through the `Order` root. There is exactly one `OrderRepository`, with `findById(orderId): Order` and `save(order): void`. To change a line item's quantity, application code calls `orderRepository.findById(id)`, then `order.changeLineQuantity(lineId, newQty)` — a method on the aggregate root that recalculates the total and checks any invariant about minimum order value — and then `orderRepository.save(order)`, persisting the whole aggregate atomically. There is no `OrderLineRepository`. If the team later needs a fast, read-only listing of all line items across orders for a warehouse-picking screen, that need is met by a separate read-model/query service reading directly from the database (or a projection), explicitly not through an `OrderLine` 'repository' that doesn't exist — keeping the write-side aggregate boundary intact while still serving the read need.

  • If Order and OrderLine live in the same aggregate, is it ever acceptable to query OrderLine rows directly from the database, bypassing the Order repository?
    Yes, for read-only purposes outside the write model - e.g., a reporting query or a warehouse-picking screen can read OrderLine rows directly from a read replica or projection, since that doesn't risk mutating state and breaking an invariant. The rule against a separate OrderLineRepository is specifically about a write-capable repository that lets code load-and-mutate a child independently of its root; read-only projections for display/reporting are a different, legitimate need usually served by a separate read path.
  • How does the one-repository-per-aggregate rule interact with an aggregate that references another aggregate, like an Order referencing a Customer?
    Cross-aggregate references are normally held by identity only (e.g., a customerId on Order), not by embedding the other aggregate's object graph, so Order's repository never needs to load or save Customer - Customer has its own repository, loaded independently when needed. This is exactly what keeps aggregates small and their repositories' load/save operations cheap: each repository's atomic unit stays scoped to one aggregate's own internal cluster.
  • What's a reasonable response when an aggregate's repository is getting too expensive to load/save as a whole, because the aggregate has grown large?
    The usual fix is to shrink the aggregate boundary itself rather than break the one-repository rule - split off a part that doesn't strictly need transactional consistency with the rest into its own aggregate, referenced by ID, with its own repository. That preserves 'one repository per aggregate' while reducing how much data each repository moves per operation.

It's like a company only ever letting outside partners deal with the department head, never walking straight into a junior employee's office to hand them instructions - the head is the one accountable for making sure everything the department does stays coherent and within policy.

saying these in an interview costs you the question

  • Creates a repository for a child entity of an existing aggregate (e.g., a LineItemRepository alongside OrderRepository)
  • Lets application code mutate a nested entity directly via its own repository without going through the aggregate root
  • Treats 'one repository per table' as equivalent to 'one repository per aggregate'
  • Can't explain what invariant is protected by routing all writes through the root
  • Assumes cross-aggregate references must be loaded eagerly as part of the owning aggregate's repository

context