skip to content

What happens to child rows when you save an existing aggregate in Spring Data JDBC, and what are the implications?

level: seniorimportance: should knowfreq 45%

answer

  1. no diff → delete all children, re-insert
  2. root = UPDATE, children = rewrite
  3. generated child ids churn each save
  4. write amplification → small aggregates
  5. triggers/audit fire every save

basics

~20 s

On updating a root, Spring Data JDBC deletes all its existing child rows and re-inserts the current ones, because it can't tell which children changed. This can churn auto-generated child ids and cause extra DELETE/INSERT SQL.

solid answer

~40 s

Because there is no dirty tracking, on an update Spring Data JDBC cannot compute a per-child diff. Its default strategy is delete-and-recreate: it DELETEs all child rows referencing the root, then INSERTs the aggregate's current children afresh. Implications: (1) child rows with database-generated @Ids get brand-new ids after each save, so nothing outside the aggregate should reference a child by id; (2) more SQL and write amplification than a targeted UPDATE, which matters for large child collections; (3) any DB triggers/audit on the child table fire on every save. This is a direct, intentional consequence of the no-dirty-tracking design and reinforces keeping aggregates small. The root row itself is a normal UPDATE; only the owned children are rewritten.

code

java · 8 lines
java
Order order = repo.findById(1L).orElseThrow(); // loads root + items
order.getItems().add(new OrderItem("SKU-9", 1)); // change ONE child
repo.save(order);
// Executed (conceptually):
//   UPDATE \"order\" SET ... WHERE id = 1;            -- root: plain update
//   DELETE FROM order_item WHERE order_id = 1;        -- ALL children removed
//   INSERT INTO order_item(order_id,...) VALUES (...); -- every current item re-inserted
// => any DB-generated order_item.id values are now different

go deeper

for a junior

Know that saving a root rewrites its children (delete + insert), not a smart update.

for a middle

Tie the rewrite to the absence of dirty tracking and note ids can change.

for a senior

Enumerate implications: id churn, write amplification, trigger/audit effects, and the small-aggregate rule.

for a principal

Use it as an aggregate-boundary design driver; decide when to split children into their own aggregate for targeted persistence and stable ids.

**The mechanism.** Spring Data JDBC has **no dirty tracking** — it holds no before-image of your objects, so on `save` of an *existing* aggregate it cannot know *which* children were added, removed, or modified. Its default resolution is the **delete-and-recreate** strategy for owned children: within the save's transaction it issues a `DELETE` of **all** child rows whose back-reference points to the root, then `INSERT`s the children currently held in the aggregate. The **root** row is handled as a normal `UPDATE` (Spring knows the root is not new because its `@Id` is set); it is specifically the **owned child collections** that are wholly rewritten. **Why it works this way.** It is the simplest correct behavior given no change detection: rather than risk a wrong diff, rewrite the children so the persisted state exactly matches the in-memory aggregate. This trades write efficiency for correctness and simplicity — consistent with the library's explicit philosophy. **Implications you must design around.** 1. **Child ids are not stable.** If children have **database-generated `@Id`s**, each save deletes the old rows and inserts new ones, so children receive **new ids**. Therefore **never let anything outside the aggregate reference a child by its id** — those references would dangle. (Owned children frequently have no `@Id` at all, which sidesteps this.) 2. **Write amplification.** A root with N children costs roughly N DELETEs (often one bulk delete) + N INSERTs on *every* save, even if only one child changed. For large or frequently-saved collections this is real overhead — another reason to **keep aggregates small**. 3. **Side effects fire every save.** Database **triggers, audit columns, or `updated_at` logic** on the child table run on each rewrite, which can distort audit history (everything looks freshly inserted). 4. **Referential integrity / cascades.** The child table's FK back to the root must permit this delete/insert cycle; other tables should not FK to child rows (see point 1). **Nuances.** - **Inserts vs. updates of the root** are decided by the root's `@Id` (null/zero → INSERT the whole aggregate; set → UPDATE root + rewrite children) or by `Persistable`/`@Version` (see the id-assignment topic). - Newer Spring Data versions optimize some cases, but the safe mental model — and the one to state in an interview — is **'children are deleted and re-inserted on update.'** Do not claim a fine-grained per-row diff exists. - **`@Version` optimistic locking** on the root still applies: a stale version fails the whole save. **When it bites / when it's fine.** Fine for small, cohesive aggregates saved occasionally. Painful for wide child lists updated hot-path frequently, or when external systems key off child ids or child-table audit trails — in those cases reconsider the aggregate boundary (split the children into their own aggregate referenced by id, so they get targeted persistence).

  • Why can't Spring Data JDBC just UPDATE the single child that changed?
    It has no dirty tracking / before-image, so it cannot compute which child changed. Delete-and-recreate guarantees the stored children match memory without needing a diff.
  • What breaks if another table has a foreign key to a child row's generated id?
    The FK dangles: each save deletes and re-inserts the child with a new id, so the external reference points at a row that no longer exists. Don't reference owned children from outside the aggregate.
  • How would you avoid the rewrite cost for a large, frequently-updated child collection?
    Reconsider the aggregate boundary — promote those children to their own aggregate root with its own repository and reference them by id, so they get targeted inserts/updates instead of full rewrite.

saying these in an interview costs you the question

  • Claiming Spring Data JDBC diffs children and updates only the changed one
  • Referencing owned children by their generated id from outside
  • Assuming child audit/updated_at columns reflect real change history
  • Ignoring write cost for wide child collections

context