skip to content

persist vs merge vs update

The write-side API distinctions interviewers never skip: persist's no-return contract, merge returning a managed copy of a detached object, and the legacy Session.update/saveOrUpdate. Also find vs getReference — the lazy reference trick that saves a SELECT.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

questions

5

In JPA, what is the difference between EntityManager.persist() and EntityManager.merge(), and what happens to the object instance you hand to each of them?

level: juniorimportance: must knowfreq 82%

answer

  1. persist = attach *this* object
  2. merge = copy onto a managed twin, returns it
  3. argument stays detached after merge
  4. merge may cost a SELECT per node
  5. IDENTITY forces the INSERT during persist

basics

~20 s

persist() makes a new object managed: that same instance is tracked and gets the generated id. merge() copies state from a detached instance onto a managed copy and returns it; the object you passed stays unmanaged.

solid answer

~50 s

**persist(entity)** is for *new* objects. It attaches the instance you passed to the persistence context, assigns an identifier (immediately for sequence/identity generators, at flush otherwise) and schedules an INSERT. That very instance is now managed, so later mutations are picked up by dirty checking. Passing an entity that already has a database identity is an error path (`EntityExistsException`, or a failure at flush). **merge(entity)** is for *detached* state. Hibernate finds the managed instance with that id in the persistence context, or loads it with a SELECT, copies your argument's field values onto it and **returns the managed copy**. Your argument remains detached, so mutating it after the call changes nothing. If the argument has no id, merge behaves like a persist onto a fresh copy — and again the new managed instance is the return value, not your object. Rule of thumb: `persist` to create inside the transaction, `merge` when an object crossed a transaction boundary and came back.

code

java · 9 lines
java
Order fresh = new Order("A-1");
em.persist(fresh);
fresh.setStatus(PAID);        // seen: fresh is managed

Order detached = loadFromPreviousTransaction();
detached.setStatus(SHIPPED);
Order managed = em.merge(detached);
detached.setStatus(CANCELLED); // lost: detached is not managed
managed.setStatus(CANCELLED);  // seen

go deeper

for a junior

Be able to state the three facts crisply: persist is for new objects and attaches your instance, merge is for detached objects and returns a managed copy, and you must use the returned reference.

for a middle

Add the mechanics: identifier assignment timing per generator, the SELECT merge may issue, and why dirty checking makes merge unnecessary on an already-managed entity.

for a senior

Talk about cost and safety — cascaded merges multiplying SELECTs, merge overwriting fields absent from a partially-populated detached object, and why loading-then-copying-allowed-fields is often the better request-handling shape.

for a principal

Frame it as an API boundary decision: whether detached entity graphs should cross the transaction boundary at all, versus commands/DTOs applied to freshly loaded aggregates, and what that choice costs in mapping complexity and accidental overwrites.

## The states involved A JPA **persistence context** is a short-lived map of entity instances the `EntityManager` currently tracks. Each instance sits in one state: - **transient / new** — built with `new`, no row, unknown to any context; - **managed / persistent** — in the context, changes are detected automatically and written at flush; - **detached** — has an identity and probably a row, but the context that loaded it is closed or the instance was evicted; - **removed** — scheduled for DELETE. `persist` and `merge` are the two transitions *into* managed state, and they differ in how they get there. ## persist `em.persist(order)` takes a transient instance and makes **that object** managed. Two consequences matter: 1. **Identity assignment.** With `GenerationType.SEQUENCE` or `TABLE`, Hibernate can obtain the id right away, so `order.getId()` is populated when persist returns, while the INSERT is deferred to flush. With `GenerationType.IDENTITY` the id only exists once the row is inserted, so Hibernate must execute the INSERT during the persist call itself — which is also why IDENTITY defeats JDBC batching of inserts. 2. **Same-object semantics.** Because your reference *is* the managed reference, subsequent setters are seen by dirty checking. `persist` returns `void` precisely because there is nothing new to hand back. `persist` on an instance that already carries a row is not a legal way to update it. Hibernate either throws `EntityExistsException` or, if it cannot tell yet, fails later with a constraint violation at flush. ## merge `Order managed = em.merge(detachedOrder);` performs a **state copy**, not an attach: 1. If the argument is already managed in this context, it is returned unchanged. 2. If it has an identifier, Hibernate looks for that id in the persistence context; on a miss it issues a `SELECT` to load the row (unless the mapping lets it skip it, e.g. a `@Version` value that proves the row exists is still not enough — the copy target must be loaded). 3. It copies the argument's basic and association state onto the managed instance, cascading `merge` where the mapping says so. 4. It returns the managed instance. The argument itself is **never** attached. This is the single most common source of "my update disappeared": the code writes `em.merge(order); order.setStatus(SHIPPED);` and the setter lands on a detached object nobody is watching. The correct shape is `order = em.merge(order);` or working with the returned reference. If the argument has a null identifier, merge treats it as new: it creates a fresh managed instance, copies the state, and schedules an INSERT. Note the generated id is set on the **copy**, so your original object still has `null` afterwards. ## Cost and when to prefer each `merge` can cost an extra SELECT per entity, and with cascading it can cost one per node of the graph. Inside a transaction where you have already loaded the entity, you do not need merge at all — load it with `find`, mutate it, and let dirty checking emit the UPDATE. Reaching for merge on an already-managed object is a smell that suggests the developer has not internalised automatic dirty checking. Also remember merge is defined to overwrite the target with *everything* the detached instance carries. If your detached object was built from a partial DTO with unset fields, merge happily writes those nulls over good data. That is a mapping/design decision, not a Hibernate bug — which is why many teams load the entity and copy only the fields the request is allowed to change, rather than merging client-supplied objects wholesale. ## Summary table | | persist | merge | |---|---|---| | Intended input | transient | detached (or new) | | Argument becomes managed | yes | no | | Return value | void | managed copy | | May issue a SELECT | no | yes | | Safe on an existing row | no | yes |

  • Why does merge return a value while persist returns void?
    persist attaches the very instance you passed, so there is nothing new to give back. merge cannot attach your instance — the persistence context may already hold a different object for that identifier, and only one instance per id is allowed — so it copies your state onto the managed instance and must return that instance for you to keep working with.
  • You loaded an entity with find(), changed a field, and called merge() before commit. What does merge do there?
    Essentially nothing useful. The instance is already managed in this context, so merge returns it unchanged and the UPDATE would have been emitted at flush anyway by dirty checking. The call is harmless but signals a misunderstanding, and in a hot path it adds needless work.
  • Does merge always issue a SELECT?
    Not always — if the identifier is already present in the persistence context (or the entity is resolvable from the second-level cache), Hibernate uses that instance as the copy target. But in the common case of a detached object arriving into a fresh persistence context, yes, it loads the row first, and cascaded merges multiply that cost.

persist hands the office your own notebook and they start tracking it. merge photocopies your notebook into the official binder and gives you the binder — scribbling in your copy afterwards changes nothing.

saying these in an interview costs you the question

  • Saying merge attaches the object you passed in
  • Believing persist returns the saved entity like a repository save method
  • Calling merge on an entity already loaded in the same transaction to 'make the update stick'
  • Assuming the generated id appears on your original object after merging a new instance
  • Thinking persist can be used to update an existing row

context

open as a page

A developer calls merge() on a detached JPA entity, keeps mutating the variable they passed in, and the later changes never reach the database. Explain what merge actually did and how the persistence context's one-instance-per-identity rule forces that behaviour.

level: middleimportance: must knowfreq 64%

basics

~20 s

merge copied the detached object's state onto a managed instance with the same id and returned that instance. The persistence context allows only one managed object per identity, so it cannot adopt your object. Later edits to the original are invisible.

open as a page

What is the difference between JPA's EntityManager.find() and EntityManager.getReference() when loading an entity by primary key, and when would you deliberately choose getReference()?

level: middleimportance: should knowfreq 52%

basics

~20 s

find() hits the database (or cache) and returns the loaded entity or null. getReference() returns a lazy proxy with only the id set, with no query until a non-id property is touched, and throws EntityNotFoundException at that point if the row is missing.

open as a page

Hibernate's native Session historically offered save(), update() and saveOrUpdate() alongside JPA's persist() and merge(). How do those legacy operations differ in behaviour, and what is their status in modern Hibernate?

level: middleimportance: should knowfreq 44%

basics

~20 s

save() inserts and returns the generated id, update() re-attaches a detached instance and forces an UPDATE, saveOrUpdate() picks between them by inspecting the id or version. All three are deprecated in Hibernate 6 and gone in Hibernate 7; use persist and merge.

open as a page

A detached order object with its collection of line items arrives back from a client and the code calls merge() on the order. What determines whether the child rows are written, and what commonly goes wrong with that collection?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Children are merged only where the association declares CascadeType.MERGE or ALL; otherwise new children fail at flush as transient references. With orphanRemoval, any child missing from the incoming collection is deleted — so a partially populated collection silently destroys rows.

open as a page