skip to content

Hibernate 6 added an upsert() operation to StatelessSession. What does it do, what does it require of the entity, and when would you use it instead of insert() or update()?

level: middleimportance: nice to knowfreq 20%

answer

  1. Insert-or-update in one statement
  2. Hibernate 6.2+, MERGE where supported
  3. Identifier must already be assigned
  4. Writes all columns — populate fully
  5. Not JPA merge(): no persistence context

basics

~20 s

upsert() writes a row without you knowing whether it already exists — one insert-or-update statement, typically SQL MERGE where the dialect supports it. It needs the identifier already assigned on the entity, and like all stateless writes it writes every mapped column.

solid answer

~50 s

`StatelessSession.upsert(entity)` (Hibernate ORM 6.2+) performs an **insert-or-update** in a single operation, so a job does not have to know whether the row already exists. Hibernate emits a SQL `MERGE` on dialects that support it, or the vendor's equivalent conflict-handling form. The key requirement is that the entity's **identifier is already assigned** — the operation is keyed by primary key, so it fits natural or externally supplied keys, not database-generated identity values you would only learn on insert. It replaces the classic ETL anti-pattern of `get()` then branch to `insert()` or `update()`, which costs an extra SELECT per row and still races with concurrent writers between the check and the write. Two caveats: it is not JPA's `merge()` — there is no persistence context, no state copying, no dirty checking — and, like every stateless write, it writes all mapped columns, so the object must be fully populated.

code

java · 8 lines
java
StatelessSession ss = sessionFactory.openStatelessSession();
Transaction tx = ss.beginTransaction();
for (SourceRow r : batch) {
    Product p = new Product(r.sku(), r.name(), r.price()); // sku is the assigned id
    ss.upsert(p);
}
tx.commit();
ss.close();

go deeper

for a junior

Know it is a one-call insert-or-update for stateless bulk writes, available in Hibernate 6.

for a middle

Add the requirements — assigned identifier, fully populated object, all columns written — and the motivation: removing a SELECT per row from ETL loops.

for a senior

Contrast it explicitly with JPA merge(), flag the stale second-level cache and missing callbacks, and say you would read the emitted SQL for your dialect before relying on it for versioned entities.

for a principal

Position it inside an idempotent-load design: upsert plus checkpointed chunking makes a job safely re-runnable, but weigh it against a staging-table-then-merge approach or a native bulk path for very high volumes.

## The problem it solves Bulk-load and synchronisation jobs constantly face the same question per row: *does this row already exist?* The naive implementation is read-then-branch: `get()` the row, then call `insert()` if it was null and `update()` otherwise. This has two defects. It doubles the statement count — a SELECT for every row on top of the write — which on a multi-million-row load is the dominant cost. And it is racy: between the SELECT and the write, another process can insert the same key, so your `insert` fails with a constraint violation or your `update` touches nothing. `StatelessSession.upsert(entity)`, added in Hibernate ORM 6.2, collapses that into one statement. Hibernate emits a SQL `MERGE` where the dialect supports it, or the vendor's conflict-handling equivalent, letting the database decide whether the row is inserted or updated. ## What it requires **An assigned identifier.** The operation is keyed on the primary key, and Hibernate must know that key before issuing the statement. That makes it a natural fit for entities whose ids come from the source data — natural keys, business keys, UUIDs generated in the application, ids supplied by the upstream system. It is not the tool for entities relying on a database identity column whose value only exists after an insert. **A fully populated object.** Like every stateless write, `upsert` has no loaded-state snapshot, so it writes **all** mapped columns from the object's current state. Passing a partially populated instance nulls out the rest of the row. If your job only knows some columns, either read the row first (giving back the extra SELECT you were trying to avoid) or use a targeted bulk HQL `update ... set` statement instead. ## What it is not It is emphatically **not** JPA's `EntityManager.merge()`, despite the surface similarity of "write whatever I give you". `merge()` operates inside a persistence context: it loads the current row if needed, copies your detached state onto the managed instance, returns that managed instance, and lets dirty checking decide the SQL at flush. `upsert()` has no persistence context, returns nothing, copies nothing, and issues one statement immediately. Conflating the two is a classic interview slip. It also inherits every other stateless property: no cascades (associated objects are not upserted), collections ignored, no lifecycle callbacks or interceptors, no second-level cache invalidation. If the entity is second-level cached, upserting it through a stateless session leaves stale cache entries behind. ## When to reach for it - **Idempotent ETL.** Re-running an import must not fail on rows already loaded. Upsert makes the job naturally re-runnable, which pairs well with checkpointed chunked reads. - **Synchronising from an upstream system of record** that supplies its own keys, where every batch is "make the local row match this". - **Backfills** that may partially overlap with data already present. ## When not to - When you know every row is new — a plain `insert()` is cheaper and will actually tell you (via a constraint violation) if your assumption is wrong, which is often valuable. - When only a few columns should change — the all-columns write can blindly overwrite fields another writer owns. - When the entity has cascaded associations, callbacks, or cache entries that must stay coherent; that is regular-`Session` territory. - On a Hibernate version older than 6.2, where the operation does not exist and you fall back to read-then-branch, a native `INSERT ... ON CONFLICT` / `MERGE` statement, or a load-into-staging-table-then-merge strategy. ## A note on versioned entities Because behaviour around version columns and conflict resolution varies by Hibernate version and dialect, verify what your combination emits before relying on it for entities that carry a version column — read the generated SQL rather than assuming. "I would check the emitted SQL for my dialect" is a stronger answer than a confident guess.

  • How is upsert() different from JPA's EntityManager.merge()?
    They share only the intuition. `merge()` works inside a persistence context: it may load the current row, copies your detached state onto a managed instance, returns that managed instance, and defers the actual SQL to flush where dirty checking narrows it. `upsert()` has no persistence context at all — it copies nothing, returns nothing, fires no callbacks, and issues one insert-or-update statement immediately with every mapped column.
  • Why does upsert() require the identifier to be assigned already?
    The statement resolves the insert-or-update decision by matching on the primary key, so Hibernate must have that key at statement-build time. Entities relying on a database-generated identity value only learn their key as a result of an insert, which is incompatible with the operation. In practice it targets natural keys, business keys, application-generated UUIDs, or ids supplied by the upstream source.

saying these in an interview costs you the question

  • Describing it as "the stateless version of JPA merge()" without noting there is no persistence context or state copying.
  • Expecting it to work with a database-generated identity column.
  • Passing a partially populated entity and being surprised that other columns are nulled.
  • Assuming it cascades to associated entities or fires @PreUpdate.
  • Claiming it exists in Hibernate 5.x.

context