skip to content

Detached Entities & Their Pitfalls

What goes wrong when entities outlive their session — lost updates from merging stale state, orphanRemoval collection traps, and the case against using entities as API payloads. Interviewers use detachment scenarios to test update-flow design.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

questions

5

A service loads an order object inside one transaction, hands it to a caller, and minutes later takes that same object back and passes it to EntityManager.merge in a new transaction. Explain how this can silently wipe out edits another user made in between, and what makes the database reject the stale write instead.

level: middleimportance: must knowfreq 65%

answer

  1. Detached = photo of the row at load time
  2. merge copies ALL fields, not a diff
  3. merge returns the managed copy, argument stays detached
  4. No version travelling → last writer wins
  5. UPDATE ... WHERE id=? AND version=?

basics

~20 s

merge copies the whole detached object onto the row, so columns you never touched are rewritten with the values they had at load time, overwriting concurrent edits. A version column carried with the object makes the UPDATE version-checked, so a stale copy fails instead of clobbering.

solid answer

~50 s

Once the loading transaction ends, the object is **detached**: a plain Java object holding a snapshot of the row as it looked then. `merge` does not diff it against the database. It looks up (or SELECTs) the current row into a managed copy, copies **every** mapped field from your detached object onto that copy, and dirty checking emits an UPDATE for the whole entity. Any column another transaction changed while your object was outside the persistence context is rewritten with your stale value: a lost update, and nothing in JPA prevents it by default. The guard is a `@Version` field that travels out to the client and back. Hibernate compares the version you carry with the loaded row and generates `UPDATE ... SET ..., version = version + 1 WHERE id = ? AND version = ?`. If someone else bumped it, no row matches and the write fails loudly. If the version never round-trips, you have no protection at all — last writer wins.

code

java · 10 lines
java
// Request 1
Order o = em.find(Order.class, 7L);
tx.commit(); em.close();          // o is now detached

// ... minutes pass; another transaction sets status = CANCELLED ...

// Request 2
o.setShippingAddress(newAddress); // only this was edited
Order managed = em2.merge(o);     // every column is rewritten from o
tx2.commit();                     // status silently reverts to NEW

go deeper

for a junior

Be able to say that a detached object holds old values, that merge writes the whole entity back, and that a version field is what makes a stale write fail.

for a middle

Explain the merge copy semantics step by step, show the version-qualified UPDATE, and note that the version must round-trip to the client to matter.

for a senior

Add the operational angle: partial payloads writing NULLs, where the conflict surfaces (merge time versus flush time), and why load-and-apply with an explicit field whitelist is the safer default.

for a principal

Frame it as an offline optimistic-concurrency protocol the application owns: choose the token, define whole-object versus field-level write semantics, and decide the conflict-resolution UX (reject, merge, or last-writer-wins per field) rather than inheriting whatever merge does.

## What detached means for a write While a transaction is open, an entity you loaded is **managed**: the persistence context holds it, keeps a snapshot of the values it was loaded with, and at flush compares the object against that snapshot to generate SQL. When the transaction ends and the persistence context closes, the object survives in memory but is no longer tracked. It is now **detached** — an ordinary object whose field values are a photograph of a row taken at load time. Nothing keeps that photograph current; the row can change under it any number of times. ## Why merge writes columns you never touched A very common misconception is that `merge` is smart: that it compares the detached object with the database and updates only the differing fields. It does not, and it cannot — it has no record of what you actually edited, only the final field values. What `merge` really does: (1) take the identifier from your detached object, (2) find the managed instance for that id in the persistence context, or SELECT the row if it is not there, (3) copy the state of your detached object onto that managed instance, field by field, (4) return the managed instance (note: not the object you passed — your argument stays detached). From step 4 onward, the managed copy is dirty relative to what was just read, so at flush Hibernate emits an UPDATE that, by default, sets every mapped column. So `merge` is a **whole-object overwrite with the values you are holding**. Two consequences follow. First, if the caller edited only `shippingAddress` but the object also carries a `status` loaded ten minutes ago, the flush writes the old `status` back over whatever another transaction set meanwhile. Second, if the object came back over the wire with fields missing (a partial JSON body deserialized into the entity class), those fields are `null` on the detached object and `merge` happily writes NULLs into columns that had data. ## The lost-update sequence 1. T1 reads order #7 (status = NEW, total = 100). Transaction ends; the object detaches. 2. T2 reads order #7, sets status = CANCELLED, commits. 3. T1's user edits the address on the stale object and calls `merge`. 4. Hibernate SELECTs order #7 (status = CANCELLED), copies the detached state over it (status becomes NEW again), flushes `UPDATE orders SET status='NEW', total=100, address='...' WHERE id=7`. The cancellation is gone, no error was raised, and no log line looks wrong. This is the canonical lost update, and it is a **staleness window** problem, not a database isolation problem — both transactions were short and correct on their own; the gap between them is application-held state. ## What actually stops it Add a version attribute (`@Version` on an int/long/timestamp field). Hibernate then treats the version as part of the write condition. During `merge`, the version you carry is compared with the version of the row that was just read; if yours is older, Hibernate raises an optimistic-lock failure rather than copying stale state. Even when the versions match at merge time, the generated statement is version-qualified, so a transaction that commits between the SELECT and the flush still causes a zero-row update and a failure at flush/commit. The part candidates miss: **the version has to survive the round trip**. It must be serialized to the client (a hidden form field, a field in the JSON, an ETag returned as `If-Match` on the next call) and set back onto the object you merge. A common bug is a mapping layer that builds a fresh entity from the request body and leaves `version` null or zero — that either resurrects an unversioned overwrite, or makes Hibernate treat the object as new. ## Safer patterns - **Load and apply.** Inside the new transaction, `find` the entity fresh and copy only the fields the request is allowed to change onto the managed instance. Dirty checking then writes only real modifications, and the version check still guards the concurrency window if you compare an incoming version token first. - **Explicit conditional update.** A JPQL/SQL `UPDATE ... WHERE id = :id AND version = :v` when you want a single narrow write with no entity load. - **DTO in, entity never out.** Keeps the detached graph from ever leaving the service and makes “which fields may this endpoint change?” an explicit list. The conceptual rule to state in an interview: any time an entity's state leaves the persistence context and comes back, you are doing an offline, application-managed concurrency protocol — and every such protocol needs a token (the version) plus a rule for whole-object versus field-level writes.

  • The client sends back only the two fields it edited and the rest of the JSON is absent, so those fields are null on the object you merge. What happens?
    merge copies null into every one of those mapped attributes, so the flush writes NULLs over columns that previously held data — and may violate NOT NULL constraints or, worse, succeed. merge has no concept of “absent” versus “explicitly null”. The fix is to load the managed entity and apply only the fields the request actually carried, or to use a DTO with explicit per-field presence.
  • Where must the version value live for the check to be meaningful?
    It has to leave the server with the data and come back with the write — a hidden field, a property in the response/request JSON, or an HTTP ETag echoed as If-Match. The version on the detached object is what Hibernate compares, so if the mapping layer drops it or resets it to zero, every request looks current and the protection is gone.
  • Does the object you passed to merge become managed?
    No. merge returns a different, managed instance; your argument stays detached. Continuing to modify the argument after the call changes nothing in the database, which is a frequent source of “my update didn't save” bugs. Always keep working with the returned reference.

Editing a printed copy of a shared document, then retyping your whole page back into the system: everything the other editor changed on that page vanishes, unless the page carries a revision number the system checks before accepting it.

saying these in an interview costs you the question

  • Saying merge only updates the fields that changed
  • Assuming a version column protects you even when the client never sends the version back
  • Claiming a higher database isolation level would prevent this — the two transactions never overlap in time
  • Believing the object passed to merge becomes managed and can be used afterwards
  • Treating this as a database concurrency bug rather than application-held stale state

context

open as a page

What concretely goes wrong when a JPA entity class is used directly as the request and response body of an HTTP API, and what does introducing a separate DTO type at that boundary actually buy you?

level: middleimportance: should knowfreq 50%

basics

~20 s

Inbound, deserialization can set fields the caller should never control (id, version, audit, ownership) and absent fields become nulls that merge writes over real data. Outbound, the object is detached and may drag lazy graphs or leak columns. A DTO makes the writable and readable field lists explicit.

open as a page

An entity mapped with @OneToMany(orphanRemoval = true) is detached, its child collection is replaced with a brand-new ArrayList rebuilt from an incoming payload, and the entity is merged back. What can go wrong, and why does Hibernate care which collection instance the entity holds?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Hibernate replaces your collection with its own tracking wrapper. Assigning a new list to a managed entity with orphanRemoval throws "a collection with cascade all-delete-orphan was no longer referenced". And on merge, any child missing from the incoming list is treated as an orphan and deleted — so a truncated payload silently deletes rows.

open as a page

You call EntityManager.merge with an entity whose @Version field is older than the value now stored in the row. Walk through what Hibernate does internally: where the two versions are compared, at which point the failure surfaces, and what happens instead if that version field is null or the row was deleted meanwhile.

level: seniorimportance: should knowfreq 38%

basics

~20 s

merge fetches the current row, compares your carried version with it, and fails immediately with an optimistic-lock error if yours is older. If versions match, the flush still issues a version-qualified UPDATE that fails on a zero-row result. A null version makes the object look new, so Hibernate tries an INSERT; a deleted row is re-created rather than reported.

open as a page

For an edit that spans several requests — a multi-step wizard, or a form a user keeps open for minutes — how would you decide between carrying a detached object graph across the steps, holding an extended persistence context, and re-reading fresh state each step while applying a recorded change set? What does each choice cost, and what do you owe the user when a conflict is detected?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

All three are ways to hold an application-level transaction across user think-time. Detached graphs are cheap but stale and destructive on merge; an extended persistence context keeps identity and dirty checking at the cost of server memory and pinned state; a change set re-read each step is the most robust. Whichever you pick, carry a version token and surface conflicts to the user rather than silently resolving them.

open as a page