skip to content

You remove a single element from a Hibernate-owned collection — an @ElementCollection or a unidirectional @OneToMany — and the SQL log shows a DELETE for every row of that collection followed by re-INSERTs of the survivors. Why does Hibernate do this, and how do you stop it?

level: seniorimportance: should knowfreq 34%

answer

  1. snapshot diff needs row addressability
  2. bag = no index + duplicates → recreate
  3. owned only: @ElementCollection, unidir @OneToMany, @ManyToMany owner
  4. mappedBy inverse collection is exempt
  5. fixes: Set, @OrderColumn, bidirectional + orphanRemoval

basics

~20 s

The collection is a bag: a List with no index column. Hibernate cannot tell which row an element corresponds to, so it recreates the whole collection. Fix it by mapping the collection as a Set, adding @OrderColumn, or making a one-to-many bidirectional with orphanRemoval.

solid answer

~60 s

Hibernate compares a collection against the snapshot taken when it loaded, then emits the statements needed to reconcile them. To emit a *targeted* statement it must be able to identify the row behind an element. In a bag — a `List` with no `@OrderColumn` — duplicates are permitted and there is no index, so there is no key linking element to row. The only correct reconciliation is to delete all rows for the owner and re-insert the current contents. It shows up on collections whose DML Hibernate owns: `@ElementCollection`, unidirectional `@OneToMany` with `@JoinColumn` or a join table, and the owning side of a `@ManyToMany`. Three fixes: - **`Set`** — elements identify rows by value or id, so removal is one `DELETE`. Usually the right answer. - **`@OrderColumn`** — the collection becomes an indexed list; removal is a `DELETE` plus renumbering `UPDATE`s. - **Bidirectional `@OneToMany(mappedBy = ...)` with `orphanRemoval`** — the collection is inverse and drives no DML; the child side handles it. The cost is not just statement count: recreating rows churns indexes, takes write locks on rows nothing changed, and inflates replication traffic.

code

java · 11 lines
java
// recreates on any removal
@ElementCollection
private List<String> tags = new ArrayList<>();

// fix 1: Set semantics -> targeted DELETE
@ElementCollection
private Set<String> tags = new HashSet<>();

// fix 2: entity children, bidirectional + orphanRemoval -> child side owns DML
@OneToMany(mappedBy = "post", cascade = CascadeType.ALL, orphanRemoval = true)
private List<Comment> comments = new ArrayList<>();

go deeper

for a junior

Recognise the symptom and know the one-line cause: a List with no index column cannot address its rows.

for a middle

Explain the snapshot diff and name the mappings affected plus the Set/@OrderColumn fixes.

for a senior

Quantify the damage — write amplification, index churn, lock footprint, CDC noise — and pick the fix that matches the write pattern.

for a principal

Discuss guarding against regressions with statement-count assertions and when a collection should not be a mapped association at all.

## The mechanism When Hibernate initialises a persistent collection it stores a **snapshot** of its contents. At flush it diffs the live collection against that snapshot and decides on statements. The question it must answer for each change is: *which database row does this element correspond to?* For a set, the answer comes from the element itself — the child entity's identifier, or the full value tuple for a value collection. For an indexed list, the `@OrderColumn` value answers it. For a **bag** there is no answer: the mapping explicitly allows two indistinguishable elements, and there is no index column. Hibernate therefore falls back to the only sound reconciliation available: *recreate*. It deletes every row belonging to the owner and re-inserts the collection's current contents. Internally this is the `needsRecreate` path on the collection persister: a bag whose contents changed at all (not merely appended) is not updatable in place. ## Which mappings are affected Only collections whose rows Hibernate is responsible for writing: - `@ElementCollection` — a value collection with its own table. - Unidirectional `@OneToMany` with `@JoinColumn`, or with an implicit/explicit join table. - The owning side of a `@ManyToMany` (the join-table rows). Not affected: a bidirectional `@OneToMany(mappedBy = "parent")`. That collection is the inverse side; it drives no DML at all. Removing an element there produces nothing by itself — the child's `@ManyToOne` owns the FK, and deletion comes from `orphanRemoval` or an explicit `remove`. This distinction is what separates a real answer from the folk rule "never use List in Hibernate". Also worth noting: a pure **append** to a bag can be handled with a single `INSERT` — Hibernate does not need to recreate when nothing was removed or reordered. The pathological case is removal or mutation in the middle. ## Why it hurts more than statement count - **Write amplification.** Removing one of 500 rows writes 500 statements (or one multi-row delete plus 499 inserts). At scale that is orders of magnitude more work. - **Index churn.** Every re-inserted row is a fresh index entry in every index on the table; the deleted ones become dead tuples the engine must clean up. - **New identifiers.** For an `@ElementCollection` with a surrogate key, or a join table with a synthetic id, the re-inserted rows are new rows. Anything outside the ORM that referenced them — an external report, a cached id, an audit trigger — sees a mass delete and a mass insert rather than a single change. - **Lock footprint.** Rows that did not change are still written, so they are still locked for the duration of the transaction, widening the window for contention with other writers. - **Replication and CDC noise.** Change-data-capture consumers see the whole collection change every time one element does. ## The fixes, in preference order **1. Map it as a `Set`.** For an `@ElementCollection` of values or a many-to-many, uniqueness is almost always the correct domain semantics anyway, and Hibernate can then target rows. Watch the `equals`/`hashCode` requirement for entity elements: a hash derived from a generated id changes at flush and breaks the set. **2. Add `@OrderColumn`** when order is genuinely part of the domain. The collection becomes an indexed list: removal is a `DELETE` plus renumbering `UPDATE`s for the elements after it. Cheaper than recreate for long collections with tail-biased changes, more expensive for head insertions. **3. Make the one-to-many bidirectional** with `mappedBy` on the parent, `@ManyToOne` on the child, and `orphanRemoval = true`. The inverse collection drives no DML and the child's removal is one `DELETE`. As a bonus this drops the join table or the extra FK-update statement that unidirectional `@JoinColumn` mappings produce. **4. Do not model it as a mapped collection at all.** For large or high-churn child sets, query children explicitly and mutate them with targeted operations or bulk statements. A collection mapping is a convenience for aggregates that are small enough to load whole. ## How to spot it Turn on SQL logging (`hibernate.show_sql` / a statement-count assertion in tests) and assert on statement counts for the mutation path. A test that fails when one removal produces N statements is the durable defence — a code review of annotations alone will not catch it, because the mapping looks perfectly reasonable.

  • Does simply appending an element to a bag also trigger the delete-all-and-reinsert?
    No. When nothing was removed or reordered, Hibernate can satisfy the diff with a plain INSERT for the new element, so append-only usage of a bag is cheap. The recreate path is taken when an element is removed or the contents change in a way Hibernate cannot express as additions, because then it cannot identify which rows to touch.
  • Why doesn't a bidirectional @OneToMany mapped as a List suffer this?
    Because with `mappedBy` the collection is the inverse side and never drives DML. The foreign key lives on the child and is written from the child's `@ManyToOne`, so Hibernate has no rows of its own to reconcile for the collection. Removals turn into a child-side DELETE via orphanRemoval or an explicit remove, both of which target a single row by its identifier.

saying these in an interview costs you the question

  • Concluding "never use List in Hibernate" without distinguishing owned from inverse collections
  • Blaming the database or a missing index for the extra statements
  • Believing @OrderBy fixes it — it sorts but adds no index
  • Assuming appending to a bag is equally expensive
  • Overlooking that re-inserted @ElementCollection rows get new surrogate keys

context