When mapping Domain-Driven Design aggregates onto microservice boundaries, why must a single aggregate always live entirely within one service, and what breaks if you split one across two services?
answer
- aggregate = local ACID transaction boundary
- root is sole write entry point
- cross-aggregate = saga, not distributed transaction
- reference other aggregates by ID only
basics
~20 sAn aggregate is a group of objects that must stay consistent together, like an order and its line items. If you split it across two services, you can't guarantee both parts update together, so the data can end up in an invalid, inconsistent state.
solid answer
~40 sAn aggregate is DDD's transactional consistency boundary: a cluster of entities/value objects (with one designated 'aggregate root') whose invariants must hold true after every commit, enforced by wrapping all changes to it in a single local ACID transaction. If you split an aggregate's state across two services, you can no longer commit those changes atomically — you're forced into a distributed transaction or a saga with compensating actions, and there's a window where the invariant is violated and another request could read or act on the inconsistent state. This is why the rule of thumb is 'one aggregate, one service' (or at minimum, one aggregate, one owning datastore) — the service boundary should never bisect an aggregate's data, even if two of its fields feel loosely related and tempting to shard for scaling reasons.
go deeper
Should be able to explain that an aggregate is a group of data that must stay consistent, and give a simple example of an invariant like 'total matches line items.'
Should know the aggregate root concept, that aggregates reference each other by ID not composition, and be able to state the 'one aggregate, one service' rule and why crossing it breaks atomicity.
Should be able to design a saga (choreography or orchestration) for a cross-aggregate workflow, and reason about when an oversized aggregate should be split into two smaller ones.
Should be able to weigh aggregate granularity against system-wide consistency and performance trade-offs, recognize aggregate-boundary drift over a system's evolution, and guide a team through re-splitting a mis-drawn aggregate/service boundary with minimal downtime.
## What an aggregate is In Domain-Driven Design, an **aggregate** is a cluster of associated entities and value objects treated as a single unit for the purpose of data changes, with one member designated the **aggregate root** — the only object external code is allowed to reference directly. The aggregate's job is to enforce **invariants**: business rules that must be true at the end of every transaction, such as: - an Order's total must equal the sum of its line items; - a bank Account's balance must never go negative. The mechanism that makes this work is transactional: every change to an aggregate is wrapped in one local **ACID transaction** against one datastore, so the invariant is checked and enforced atomically before the transaction commits. Nothing outside the aggregate boundary is allowed to reach in and mutate its internals directly; all changes go through the root, which is what lets the aggregate guarantee its own consistency. ## Why one aggregate belongs entirely to one service This is why the rule *one aggregate belongs entirely to one service* exists. A microservice boundary is, among other things, a **data-ownership boundary** — each service owns its own datastore and no other service is allowed to write to it directly. If an aggregate's state were split across two services (say, an Order's header fields in an Order service and its line items in a separate Line Item service), then enforcing *total equals sum of line items* would require a transaction spanning two independently-deployed, independently-failing services and their separate databases. Relational databases don't offer atomic transactions across two different services' schemas in a microservices world (no shared connection, no two-phase commit in practice), so you'd be forced to either: - **fake atomicity with a distributed transaction protocol** — slow, fragile, rarely used in production microservices; or - **accept eventual consistency via a saga** — a sequence of local transactions coordinated by events, with compensating actions to undo earlier steps if a later step fails. ## The trade-off The trade-off here cuts both ways. - **Keeping the whole aggregate in one service** preserves strong, immediate consistency and a simple transactional model — the invariant is either true or the commit didn't happen, full stop. - **The cost** is that if the aggregate is large or has fields that genuinely have different scaling or ownership needs (e.g. an `Order` aggregate that also embeds shipment-tracking state which changes far more frequently and is owned by logistics), you're stuck keeping unrelated concerns co-located, or you have to do the harder domain work of asking whether what you called *one aggregate* is really two aggregates in disguise, connected only by an ID reference rather than a hard object composition. DDD explicitly recommends designing aggregates small — reference other aggregates by ID, not by object composition — precisely so this decision (what must be transactionally consistent vs. what can be eventually consistent via an ID reference and a domain event) is made deliberately during modeling, not accidentally during a later *let's split this service for scaling* exercise. ## Failure modes Failure modes show up in two directions. 1. **First**, if a team draws the service boundary through the middle of what should have been one aggregate — usually because it looked like two tables that could scale independently — they end up building ad hoc distributed-transaction logic (retry loops, manual reconciliation jobs, *fix the data* scripts) to patch over consistency bugs that the aggregate boundary was supposed to prevent for free. This shows up in production as orders with totals that don't match their line items after a partial failure, or accounts that go negative because a debit and a credit landed in two services and only one committed. 2. **Second**, the opposite failure — treating everything reachable by object graph as one giant aggregate spanning unrelated concerns — produces an aggregate so large that most transactions lock or touch far more data than necessary, hurting throughput and forcing the service itself to become a monolith-in-miniature that nobody can safely evolve independently. ## Where it shows up A concrete, widely-used real-world pattern is the **Saga pattern** paired with well-drawn aggregate boundaries: an e-commerce `Order` aggregate lives entirely inside an Order service, and when placing an order also needs to reserve inventory (owned by an Inventory service, a different aggregate, different bounded context) and charge a payment (owned by a Payment service, another aggregate), the Order service doesn't try to wrap all three in one transaction. Instead it runs a choreography- or orchestration-based saga: 1. Place the order locally — one atomic commit against the `Order` aggregate. 2. Publish an `OrderPlaced` event. 3. Let Inventory and Payment react and complete their own local, atomic transactions against their own aggregates. 4. Use compensating events (e.g. `OrderCancelled`, triggering a compensating stock release) if a later step fails. Each individual step stays strongly consistent within its own aggregate/service; only the overall multi-service workflow is eventually consistent — which is the accepted, standard trade-off for cross-aggregate coordination in microservices.
- If an Order aggregate should never span two services, how do you handle the fact that an Order needs to reference a Customer, which is a completely different aggregate owned by a different service?The Order aggregate holds only a Customer ID (a reference), never the Customer's full object graph — this is DDD's standard rule that aggregates reference other aggregates by identity, not by composition. If the Order service needs Customer details for a specific operation, it either calls the Customer service's API synchronously or keeps a locally-cached, eventually-consistent read copy populated via domain events, rather than trying to own or transact against Customer data directly.
- What's the practical difference between choreography and orchestration when implementing a saga across aggregates in different services?In choreography, each service reacts to events published by the others with no central coordinator — the Order service publishes OrderPlaced, Inventory reacts and publishes StockReserved, Payment reacts to that, and so on, which keeps services decoupled but makes the overall workflow harder to see and debug as one unit. In orchestration, a dedicated saga orchestrator explicitly calls each service in sequence and tracks the workflow's state centrally, which is easier to reason about and monitor but introduces a coordinating component that some argue re-centralizes logic the microservices split was meant to distribute.
- A team notices their Order aggregate has grown to include shipment tracking events that update every few seconds from a carrier webhook, while order placement itself only changes a few times. What does this suggest about their aggregate design?It suggests shipment tracking is really a separate aggregate (and likely a separate bounded context/service, e.g. Shipping) that should be linked to the Order by ID reference rather than embedded, since the two have very different consistency needs and update frequencies. Keeping them fused forces every high-frequency tracking update to contend for locks on the same aggregate as the comparatively rare order-placement writes, hurting throughput and complicating the transactional model unnecessarily.
Like a single bank teller window that processes a withdrawal atomically — cash out and ledger update happen together or not at all. You wouldn't split 'take the cash out of the drawer' and 'update the ledger' across two different bank branches that only talk to each other by mail, because there'd be a window where the drawer is short and the ledger doesn't know it yet.
saying these in an interview costs you the question
- Suggests using two-phase commit / distributed transactions across services as the normal way to keep an aggregate consistent
- Doesn't know what an aggregate root is or why it's the only mutation entry point
- Thinks 'aggregate' just means 'database table' with no mention of invariants
- Proposes embedding a full referenced aggregate's object graph inside another aggregate instead of an ID reference
- Can't explain the difference between transactional (in-aggregate) and eventual (cross-aggregate) consistency