skip to content

A detached order object with its collection of line items arrives back from a client and the code calls merge() on the order. What determines whether the child rows are written, and what commonly goes wrong with that collection?

level: seniorimportance: should knowfreq 40%

answer

  1. cascade MERGE or nothing propagates
  2. 'object references an unsaved transient instance'
  3. orphanRemoval + short client list = deleted rows
  4. mutate collections in place, never reassign
  5. fix the inverse side on the merged copies

basics

~20 s

Children are merged only where the association declares CascadeType.MERGE or ALL; otherwise new children fail at flush as transient references. With orphanRemoval, any child missing from the incoming collection is deleted — so a partially populated collection silently destroys rows.

solid answer

~60 s

Merge walks the graph only along associations mapped with `CascadeType.MERGE` (or `ALL`). Along those edges each child is resolved and copied like the root, so a graph of N nodes can cost N SELECTs. Without the cascade, a new child reachable from a managed parent triggers `TransientObjectException` / 'object references an unsaved transient instance' at flush. The collection itself is the sharp edge: - With `orphanRemoval = true`, children present in the database but absent from the detached collection are **deleted**. A client that posts a trimmed or lazily-unfetched collection wipes rows. - If the detached collection was an uninitialised proxy, merging it can null out or drop the association state. - Bidirectional links must be fixed up on the managed side — the child's `order` reference must point at the merged parent, or the FK is written null. - Assign the result: `order = em.merge(order)` — the children you keep working with are the copies inside it. Safer pattern: load the aggregate by id and apply an explicit diff instead of merging client-supplied graphs.

code

java · 8 lines
java
@OneToMany(mappedBy = "order",
           cascade = CascadeType.ALL,
           orphanRemoval = true)
private List<OrderLine> lines = new ArrayList<>();

// merge copies the incoming list onto the managed collection:
// lines missing from `incoming` are deleted at flush
Order managed = em.merge(incomingOrder);

go deeper

for a junior

Know that cascade settings decide whether children are saved and recognise the 'unsaved transient instance' message.

for a middle

Explain the recursive copy, cascade MERGE, the persistent-collection wrapper and why collections must be mutated rather than replaced.

for a senior

Lead with the data-loss risk from orphanRemoval plus a partially populated client collection, the per-node SELECT cost, and the load-and-diff pattern you would put in its place.

for a principal

Argue about aggregate boundaries and API shape: whether entity graphs should be accepted from clients at all, versus explicit commands, and how that choice constrains cascade and orphan-removal mappings across the model.

## What merge does to a graph `merge` is defined recursively. For the root it resolves a managed target and copies state; for each association annotated with `CascadeType.MERGE` (implied by `ALL`) it repeats the process on the referenced entity or collection elements. Associations without that cascade are treated as plain references: whatever managed instance the copied reference resolves to is used, and if it resolves to nothing persistent you get an error at flush. Three distinct outcomes follow from that definition. **1. Missing cascade.** The order is merged, but a brand-new `OrderLine` hanging off it is not, so at flush Hibernate finds a managed entity pointing at a transient one and throws `object references an unsaved transient instance — save the transient instance before flushing` (`TransientObjectException` / `TransientPropertyValueException`). The fixes are to add `cascade = CascadeType.MERGE`/`ALL`, or to persist the child explicitly. This is the single most-quoted Hibernate error message and interviewers expect you to recognise it instantly. **2. Cost.** Every merged node may cost a SELECT to fetch the copy target. Merging an order with 200 lines can be 201 statements before a single write. Batch imports that merge whole graphs per record are a classic cause of 'the job got slower as data grew'. If the state is genuinely new, `persist` on the root with cascade PERSIST is far cheaper because it never needs to read. **3. Deletion by omission.** This is the dangerous one. `@OneToMany(orphanRemoval = true)` means 'a child no longer referenced by the parent has no reason to exist — delete it'. Merge copies the *incoming* collection contents onto the managed collection, so any child the client did not send is now unreferenced and is deleted at flush. If the client posted only the two lines it edited, the other eighteen are gone. Even worse, if the detached parent's collection was never initialised (the transaction that produced it closed before the collection was touched), the merged collection may end up empty or the association state may be discarded, deleting everything. ## Collection mechanics worth naming Hibernate replaces mapped collections with its own **persistent collection** wrappers (`PersistentBag`, `PersistentSet`, ...) that track additions and removals. Two habits break this: - **Replacing the collection instance** on a managed entity (`order.setLines(new ArrayList<>(incoming))`) throws away the tracked wrapper. Historically this produced `A collection with cascade="all-delete-orphan" was no longer referenced by the owning entity instance`. The rule is to mutate in place: `lines.clear(); lines.addAll(incoming);`. - **Forgetting the inverse side.** In a bidirectional one-to-many the child owns the foreign key. Merging a parent whose children do not point back at it writes null FKs or leaves orphans; that is why entities carry `addLine(line) { lines.add(line); line.setOrder(this); }` helpers, and why after a merge the fix-up must run against the *returned* managed instances, not the detached ones. ## Optimistic-lock interaction The detached root carries whatever version value it had when it left. Merge copies that value onto the managed instance, so the flush compares the client's version against the row — which is precisely how a stale client edit is rejected. Two consequences: re-merging the same detached instance twice fails the second time, because its version is now behind; and a graph where only children changed will not bump the root's version unless the mapping asks for it. ## The pattern most teams end up with Rather than merging client-supplied graphs, load the aggregate inside the transaction and apply an explicit diff: 1. `Order managed = em.find(Order.class, id);` (with the collection fetched). 2. Compare incoming line identifiers against the managed collection. 3. Update matched lines' allowed fields, add genuinely new ones, remove the ones the request explicitly deleted. 4. Let dirty checking emit the statements. That costs one deliberate query, makes deletion an intent rather than an accident, keeps unmapped or non-editable fields safe from overwrite, and turns 'the client sent a short list' from data loss into a validated decision. Merge on a graph is right when you truly own the whole detached aggregate — an import file, a serialised conversation state — and want insert-or-update semantics for all of it.

  • What exactly triggers 'object references an unsaved transient instance' during a merge?
    At flush Hibernate finds a managed entity holding a reference to an instance with no persistent identity, along an association that has no PERSIST/MERGE cascade. It refuses to write a foreign key it cannot resolve. Either add the cascade to that association or persist the referenced instance explicitly before flushing.
  • How does the version value on a detached root affect a merge?
    Merge copies the detached instance's version onto the managed copy, so the flush writes with that value in the WHERE clause and fails if the row moved on. That is what makes a stale client edit visible, and it also means the same detached instance cannot be merged twice — after the first merge its version is behind the row.

saying these in an interview costs you the question

  • Expecting children to be saved without a MERGE or ALL cascade on the association
  • Not realising orphanRemoval deletes children merely absent from the incoming collection
  • Reassigning a managed entity's collection to a new List instead of clearing and re-adding
  • Ignoring the inverse side, then blaming Hibernate for null foreign keys
  • Merging large graphs in a loop and being surprised by the query count

context