skip to content

You run a bulk JPQL statement such as "update Product p set p.price = p.price * 1.1 where p.category = :c" via Query.executeUpdate(). What happens to Hibernate's second-level cache and to entities already loaded in the current persistence context?

level: middleimportance: must knowfreq 48%

answer

  1. One SQL statement, no entities loaded
  2. No keys known, so whole region evicted
  3. Update-timestamps bumped, query results stale
  4. Persistence context untouched: clear() or refresh()
  5. Native SQL: declare synchronized entity classes

basics

~20 s

The statement runs as one SQL UPDATE without loading entities. Hibernate cannot know which rows changed, so it evicts the whole cache region for the affected tables. Entities already loaded in the session are not refreshed and stay stale until you clear or refresh them.

solid answer

~50 s

Bulk JPQL DML is translated straight to SQL and executed in the database; no entities are loaded, no dirty checking runs, and no lifecycle callbacks fire. Because Hibernate only knows the *tables* touched (its query spaces), not the individual identifiers, it invalidates at region granularity: the entity regions for those tables are flushed, and the update-timestamps used by the query cache are bumped so any cached query results over those tables are treated as out of date. That protects the shared cache but not the persistence context. Entities already managed in the current session keep their pre-update field values and their old snapshot, so a subsequent dirty check can even write the stale values back. The disciplined pattern is to run bulk DML early, on a session with nothing relevant loaded, then `clear()` the persistence context (or `refresh()` the specific entities) before continuing.

code

java · 6 lines
java
int rows = em.createQuery(
    "update Product p set p.price = p.price * 1.1, p.version = p.version + 1 " +
    "where p.category = :c")
  .setParameter("c", category)
  .executeUpdate();
em.clear(); // otherwise managed Products still hold pre-update prices

go deeper

for a junior

Know that a bulk JPQL update runs as raw SQL, does not refresh objects already loaded in the session, and that you should clear the persistence context afterwards.

for a middle

Explain region-level eviction and why identifiers are unknown, plus the stale-snapshot danger where a later flush writes old values back.

for a senior

Add operational consequences: cache stampede after evicting a hot region, version-column handling, callbacks not firing, and the native-SQL synchronized-spaces requirement.

for a principal

Weigh bulk DML against entity-by-entity updates as a design choice: throughput and lock footprint versus cache churn, auditing, and lifecycle semantics; decide where in a job schedule it runs.

## What bulk DML is A bulk statement is a JPQL/HQL `update` or `delete` executed with `Query.executeUpdate()`. It is not a loop over entities: Hibernate translates it into a single SQL statement and sends it to the database. That is why it is fast, and also why it sidesteps almost everything the ORM normally does for you: no entity instances, no dirty checking, no cascade, no `@PreUpdate` callbacks, no optimistic version bump unless you write one into the statement yourself. ## Effect on the second-level cache Hibernate tracks, for every query, the set of tables it touches, known as query spaces. For bulk DML it knows the query spaces but not the primary keys of the affected rows, since the database applied the `where` clause and never reported which rows matched. With no key list, precise invalidation is impossible, so Hibernate takes the conservative route: - the entity cache regions mapped to those tables are evicted wholesale; - collection regions belonging to those entities are evicted; - the update-timestamps cache is stamped for those spaces, which makes every cached query result over them stale. This is correct but blunt. One `update Product ...` touching ten rows throws away every cached `Product`, and on a busy node the immediate consequence is a burst of database reads while the region refills. On a warm system a bulk statement over a heavily cached table is a small self-inflicted stampede, which is why such jobs are usually scheduled off-peak or split into narrower entities. ## Effect on the persistence context The first-level cache is untouched. If the session already holds `Product#7` with `price = 10`, that instance still says 10 after the bulk statement raised it to 11. Worse, its loaded snapshot also says 10, so: - a later `em.find(Product.class, 7L)` returns the stale managed instance without going to the database; - if you then change some unrelated field, the generated UPDATE writes the whole row from the stale in-memory state and can silently undo the bulk change. Hibernate flushes pending changes before executing the bulk statement (so your own dirty entities are not lost), but it does not and cannot re-read what the statement did. ## The safe pattern ``` em.createQuery("update Product p set p.price = p.price * 1.1 where p.category = :c") .setParameter("c", cat) .executeUpdate(); em.clear(); // drop stale managed copies ``` Either run bulk DML in a short session dedicated to it, or `clear()` afterwards, or `refresh()` the handful of entities you still need. In a versioned entity, remember to bump the version column inside the statement (`set p.version = p.version + 1`) so that concurrent optimistic checks still behave. ## Native SQL is worse A native `update`/`delete` executed through Hibernate carries no query-space information unless you supply it. Without `addSynchronizedEntityClass(...)` (or the equivalent on a named native query), Hibernate evicts nothing and the second-level cache keeps serving pre-update rows indefinitely. ## What interviewers are checking They want to hear that bulk DML bypasses the ORM layer, that cache invalidation happens at region granularity because identifiers are unknown, that the persistence context is a separate problem you must solve by hand, and that native SQL needs explicit table declarations. Candidates who answer only "Hibernate handles it" have not thought about the two caches separately.

  • Why does Hibernate evict a whole region instead of just the affected identifiers?
    Because the database applied the WHERE clause and does not report the matched primary keys back to Hibernate. Knowing only the tables involved, the only conservative choice is to drop every cached entry for those tables. Precise invalidation would require first selecting the ids, which defeats the purpose of bulk DML.
  • Does bulk DML update the @Version column or fire entity callbacks?
    No. It bypasses the entity lifecycle entirely, so no optimistic version increment, no @PreUpdate/@PostUpdate, no cascading to associations. If the entity is versioned you should increment the version explicitly inside the statement, otherwise concurrent sessions holding old versions will not detect that the row moved under them.
  • What is the equivalent risk with a native SQL update run through Hibernate?
    It is larger: without declaring the affected tables via addSynchronizedEntityClass or synchronized query spaces, Hibernate invalidates nothing at all, and the second-level cache keeps serving pre-update state until the entry expires. Always declare the spaces or evict the region manually.

saying these in an interview costs you the question

  • Claiming Hibernate refreshes the loaded entities after a bulk update
  • Believing only the matching rows are evicted from the cache
  • Assuming @Version is bumped and @PreUpdate fires for bulk statements
  • Thinking native SQL updates invalidate the cache the same way JPQL DML does
  • Running bulk DML mid-session and continuing to use previously loaded entities

context