A developer calls merge() on a detached JPA entity, keeps mutating the variable they passed in, and the later changes never reach the database. Explain what merge actually did and how the persistence context's one-instance-per-identity rule forces that behaviour.
answer
- one managed instance per identity
- merge copies, never adopts
- always assign the return value
- em.contains(arg) == false
- merged reference != original reference
basics
~20 smerge copied the detached object's state onto a managed instance with the same id and returned that instance. The persistence context allows only one managed object per identity, so it cannot adopt your object. Later edits to the original are invisible.
solid answer
~50 sThe persistence context guarantees **one managed instance per entity identity**. If a row with that id is already represented in the context, adopting your detached object would create two managed instances of the same row — so JPA copies instead of adopting. So `merge` (a) resolves the target: the context's existing instance for that id, otherwise a row loaded by SELECT, otherwise a new instance for an INSERT; (b) copies your argument's basic fields and mapped associations onto it, cascading where `CascadeType.MERGE`/`ALL` is declared; (c) returns the managed target. Your argument never enters the context, so nothing dirty-checks it. The fix is simply to keep the return value: ```java order = em.merge(order); order.setStatus(SHIPPED); // now tracked ``` A useful diagnostic is `em.contains(order)` — false for the argument, true for the return value. It also explains why `merge` cannot be `void`, and why `==` comparisons between the pre- and post-merge references fail.
code
java · 6 linesOrder inContext = em.find(Order.class, 1L); // canonical instance for id 1
Order detached = deserializeFromRequest(); // also id 1, different object
Order merged = em.merge(detached);
assert merged == inContext; // same identity -> same managed instance
assert em.contains(detached) == false;go deeper
Recall the rule: use the object merge returns. Being able to say 'the argument stays detached' is enough at this level.
Explain why — one managed instance per identity — and show the diagnostic (em.contains, discarded return value) plus the null-id and stale-version symptoms.
Bring in the failure modes you have actually debugged: lost updates, reference-equality bugs in caches or collections, cascade gaps producing transient-instance errors, and the extra SELECT per merged node.
Argue about whether detached entity graphs should be part of your API surface at all; weigh insert-or-update convenience against client-supplied identifiers and accidental overwrites of fields the caller never saw.
## The invariant that drives everything A persistence context is a map from *entity identity* (entity type + primary key) to a single Java instance. This is what makes the first-level cache behave like an identity map: two `find` calls for the same id in the same context return `==` the same object. Hibernate relies on this to do dirty checking against exactly one snapshot per row and to avoid emitting two conflicting UPDATE statements for the same row. Now imagine `merge` attached your detached instance directly. If the context had already loaded that row — say some validation code called `find` earlier in the request — there would be two managed objects claiming the same identity, each with its own field values and its own snapshot. Hibernate would have no defensible answer for which wins. JPA therefore specifies merge as a **state copy onto the canonical instance**, not an attach. ## What merge does step by step 1. **Already managed?** If `em.contains(arg)` is true, merge returns the argument itself and does nothing else. 2. **Removed?** Merging an instance scheduled for removal is illegal and throws `IllegalArgumentException`. 3. **No identifier?** Treated as new: Hibernate creates a fresh instance, copies the state, persists it, and returns the new instance. The generated id appears on the copy, not on your object. 4. **Has an identifier?** Hibernate resolves the managed target — from the persistence context if present, otherwise by loading it (a SELECT, possibly served by the second-level cache). If no row exists, the outcome depends on the mapping and provider: classically, Hibernate treats it as a new instance to insert, which is why merge is often described as 'insert-or-update'. 5. **Copy.** Basic properties, embeddables and mapped associations are copied. Associations annotated with `CascadeType.MERGE` (or `ALL`) are merged recursively; associations without it are not, and referencing a transient object through such an association will blow up at flush with 'object references an unsaved transient instance'. 6. **Return** the managed target. ## Symptoms candidates should recognise - **Lost updates after merge** — the classic. Mutations go to the detached argument. - **`==` surprises** — `merged != original` in the common case, so code that later compares references, or holds the original in a collection/map, silently works on the wrong object. - **Identifier still null** — after merging a new instance, the caller's object has no id and the code that returns `order.getId()` to the client returns null. - **Version field not advanced on your copy** — the optimistic-lock version is incremented on the managed instance; the detached one keeps the stale value, so reusing it for a second merge later can fail. ## Diagnosing and avoiding it `em.contains(x)` is the direct test: it answers 'is this exact instance managed in this context'. In review, the smell is any statement of the form `em.merge(x);` where the return value is discarded, or `x` is used afterwards. Treat `merge` like `String.toUpperCase()` — it produces a new thing rather than modifying in place; discarding the result is almost always a bug. The deeper avoidance strategy is to not round-trip entities through the client at all. Load the entity by id inside the transaction, apply the fields the use case permits, and let dirty checking write the UPDATE. That eliminates both the merge SELECT and the whole class of 'which instance is the real one' confusion. Merge earns its place when you genuinely hold a detached graph — for instance state carried across a long conversation, or an object deserialised from an outbox/import file — and you want insert-or-update semantics for it. Finally, note the asymmetry with the legacy native alternative: Hibernate's old `Session.update()` re-attached your actual instance and therefore threw `NonUniqueObjectException` when the context already held that id. Merge's copy semantics are exactly what buys you the ability to call it unconditionally.
- What does em.contains() tell you here, and what does it not tell you?It reports whether that exact instance is managed in the current persistence context. It is an identity check, not a database check: it returns false for a detached object that certainly has a row, and true for a newly persisted instance whose INSERT has not been flushed yet.
- You merge a detached entity that has no matching row in the database. What happens?Hibernate cannot find a copy target, so it treats the instance as new and schedules an INSERT with the copied state, returning a fresh managed instance. If the id was assigned rather than generated, the row is inserted with that id, which is how merge earns its 'insert-or-update' reputation — and why it is dangerous for entities whose ids come from clients.
- Why did the legacy Session.update() throw NonUniqueObjectException where merge does not?update() re-attached the instance you handed it. If the persistence context already held a different instance for that identifier, attaching would break the one-instance-per-identity rule, so Hibernate refused with NonUniqueObjectException. merge sidesteps the conflict by copying state onto the instance that is already there.
saying these in an interview costs you the question
- Treating merge as an in-place attach and discarding its return value
- Claiming merge and update are interchangeable because 'both save'
- Expecting the generated identifier to appear on the object passed to merge
- Assuming the returned instance is always a different object (it is the same one if it was already managed)