In Spring Data JDBC, what is an aggregate and an aggregate root, and how are they persisted?
answer
- cluster + root = consistency unit
- one repository per root only
- children reached through root
- save/load act on whole aggregate
- cross-aggregate = reference by id
basics
~20 sAn aggregate is a group of related objects saved and loaded as one unit. The aggregate root is the top object you access it through. Spring Data JDBC has one repository per root and saves the whole aggregate together.
solid answer
~40 sSpring Data JDBC follows Domain-Driven Design aggregates: an aggregate is a cluster of an entity plus the child entities it owns, treated as a single consistency unit. The aggregate root is the entity outside code references; children are reached only through it. You get exactly one repository (e.g. CrudRepository) per aggregate root, and every save or load acts on the whole aggregate. When you save the root, its owned children (mapped via @MappedCollection) are written too; when you load the root, they are all fetched eagerly. There is no repository for children. This keeps a clear ownership boundary: the root controls the lifecycle of everything it contains, so transactional consistency is scoped to one aggregate at a time.
code
java · 20 lines// Aggregate root
class Order {
@Id Long id;
String customerRef; // cross-aggregate ref: just an id, not an object
@MappedCollection(idColumn = "order_id")
Set<OrderItem> items = new HashSet<>();
}
// Owned child — NO repository of its own
class OrderItem {
String sku;
int qty;
}
// Exactly one repository: for the root
interface OrderRepository extends CrudRepository<Order, Long> {}
// save writes Order + all items; findById loads them all eagerly
orderRepository.save(order);
Order loaded = orderRepository.findById(1L).orElseThrow();go deeper
Know the definitions: aggregate = group saved/loaded together, root = the top entity, one repository per root.
Explain that children map to child tables and are written on the root's save; cross-aggregate links are ids.
Frame the aggregate as the transaction/consistency boundary and justify keeping aggregates small.
Discuss aggregate design trade-offs: boundaries, referencing by id vs embedding, and how this shapes the domain model and load/write cost.
**Spring Data JDBC** maps objects to relational tables using JDBC directly — deliberately simpler than JPA/Hibernate: no proxies, no persistence context, no lazy loading, no dirty tracking. Its central design idea is the **aggregate**, borrowed from Domain-Driven Design (DDD). **Aggregate** — a cluster of objects that belong together and must stay consistent as a unit. Example: an `Order` with its `OrderItem` lines forms one aggregate. **Aggregate root** — the single entity that is the entry point. Outside code references only the root; child entities are reached *through* it, never loaded or saved on their own. In the example `Order` is the root and `OrderItem` is an owned child. **One repository per aggregate root.** You declare a repository (e.g. `interface OrderRepository extends CrudRepository<Order, Long>`) only for roots. You do **not** create an `OrderItemRepository`. This enforces the rule that children have no independent lifecycle. **Whole-aggregate operations.** - `save(order)` writes the root row *and* all its owned children (each child collection maps to a child table via `@MappedCollection`). - `findById(id)` loads the root *and* eagerly fetches every child — there is no lazy loading, so what you get back is the complete aggregate. - `delete(order)` removes the root and cascades to its children. **Why aggregates matter.** The aggregate is the **consistency and transaction boundary**. Spring Data JDBC guarantees consistency only *within* one aggregate. References between *different* aggregates are modeled not as object references but as **plain id values** (e.g. store a `customerId` field, not a `Customer` object). This keeps aggregates small and independently loadable, and avoids accidentally dragging half the database into memory. **Gotchas / when to use.** Because the whole aggregate is loaded and rewritten together, keep aggregates small — a root with thousands of children means huge loads and saves. Model only truly-owned, lifecycle-bound data as children; everything else is a cross-aggregate reference by id. Use Spring Data JDBC when you want explicit, predictable SQL and a clean DDD ownership model; reach for JPA when you need lazy graphs, dirty tracking, or complex inheritance mapping.
- How do you model a reference from one aggregate to another?Not with an object reference — store the other aggregate's id as a plain field (e.g. a Long customerId or an AggregateReference). This keeps aggregates independently loadable and preserves the consistency boundary.
- Why is there no repository for child entities?Children have no independent lifecycle; they exist only as part of the root's aggregate. Giving them a repository would let code load/save them outside the root, breaking the ownership and consistency boundary DDD aggregates enforce.
saying these in an interview costs you the question
- Thinking every entity gets its own repository
- Modeling cross-aggregate links as object references instead of ids
- Assuming children can be loaded independently of the root
- Believing children are lazy-loaded like in JPA