When your code updates an entity through Hibernate and that entity is held in Hibernate's second-level cache, how does Hibernate stop the cached copy from going stale, and at what point in the transaction does that happen?
answer
- Hibernate wrote the SQL, so it knows the key
- Marked at flush, published after commit
- Rollback evicts, never writes
- Only covers Hibernate's own entity writes
- Triggers, cascades, DBA scripts are invisible
basics
~20 sHibernate issues the UPDATE itself, so it knows which cache key changed. At flush it marks the entry as in-flight; only after the database transaction commits does it replace or remove the entry. On rollback the entry is just dropped, never updated.
solid answer
~50 sBecause Hibernate performs the SQL, it can invalidate precisely. The write is two-phase and tied to transaction completion, not to flush. During flush, when the UPDATE or DELETE is sent, Hibernate marks the affected cache entry as being modified so it is no longer served as valid. The real put-or-remove happens in an after-transaction-completion callback: on commit the entry is refreshed with the newly written state or removed; on rollback it is simply evicted, so the cache can never hold a value the database rejected. Two consequences matter in interviews. First, uncommitted changes are never published through the shared cache, so other sessions cannot read dirty data from it. Second, this write-through path covers only writes Hibernate itself makes on managed entities. Bulk JPQL DML, native SQL, database triggers, ON DELETE CASCADE at the schema level, and other applications all bypass it, and those cases need explicit eviction or a decision not to cache that entity at all.
code
java · 6 linesem.getTransaction().begin();
Product p = em.find(Product.class, 42L); // may come from L2
p.setPrice(new BigDecimal("19.99")); // dirty checking
em.getTransaction().commit();
// flush -> UPDATE product ... ; after commit -> L2 entry for Product#42
// is refreshed or removedgo deeper
Know that Hibernate updates or clears the cached copy itself when you change an entity, and that this happens when the transaction commits, not before.
Explain the two-phase timing (marked at flush, published or dropped after commit), the rollback behaviour, and that only Hibernate-issued writes on managed entities are covered.
Add the blind spots: bulk DML, native SQL without synchronized spaces, triggers, database cascades, and other writers on the schema, plus what you do about each.
Frame it as a consistency contract: the shared cache gives eventual consistency bounded by commit propagation, so it must never be the mechanism a correctness-critical read depends on.
## Why the problem exists The second-level (L2) cache belongs to the `SessionFactory` (the `EntityManagerFactory`), so it outlives any single session and is shared by every session in the JVM. It stores entity state keyed by identifier. The instant a row changes, any cached copy of that row is wrong. Hibernate's defence for writes it performs itself is called write-through invalidation: the cache is updated or cleared as part of the same unit of work that changed the row. ## What Hibernate actually does, step by step 1. You mutate a managed entity, or call `persist`/`remove`. 2. At flush time Hibernate compares the entity against the loaded snapshot, decides an UPDATE/INSERT/DELETE is needed, and sends the SQL. 3. Before or while that SQL runs, Hibernate touches the cache entry to mark it as under change. Depending on the configured concurrency strategy this is a soft lock or an outright removal; what matters conceptually is that from this moment nobody is served the old value as if it were fresh. 4. Hibernate registers an after-completion callback with the transaction. 5. On **commit**, the callback either writes the new state into the cache (an update-in-place, cheap because the state is already in hand) or removes the key so the next reader reloads from the database. 6. On **rollback**, the callback removes the key. The cache is left cold rather than wrong. The key timing fact: nothing valid is published to the shared cache at flush. Flush only synchronises the persistence context with the database inside the transaction; the L2 cache is transaction-visible only after commit. That is what keeps other sessions from reading uncommitted data through the cache. ## Inserts and deletes Inserts may be added to the cache after commit (for strategies that support it) or simply left absent so the first reader loads them. Deletes always remove the key after commit; a cached entry for a row that no longer exists is one of the nastier stale states because code then works with a phantom object. ## What write-through does not cover This mechanism is only as good as Hibernate's visibility into the writes: - **Bulk JPQL/HQL DML** (`update`/`delete` executed with `executeUpdate()`) does not load entities, so Hibernate cannot know which keys changed; it invalidates the affected regions wholesale instead. - **Native SQL** run through Hibernate invalidates nothing unless you declare which tables it touches (synchronized query spaces). - **Anything outside the application**: DBA scripts, ETL and batch loaders, another service on the same schema, database triggers, and referential actions such as `ON DELETE CASCADE`. Hibernate never sees these rows change. For those cases the options are: route all writes through the application, evict explicitly after the external job runs, bound staleness with a time-to-live in the cache provider, or simply do not cache that entity. ## Clusters With more than one node, the after-commit step must also tell the other nodes. In an invalidation topology the node broadcasts "forget this key"; in a replicated topology it ships the new value. Asynchronous propagation leaves a short window where another node still answers with the pre-commit value, which is the fundamental reason L2 caching is an eventual-consistency feature, not a correctness feature. ## How to talk about it Say the mechanism (Hibernate performs the SQL, therefore it knows the key), say the timing (marked at flush, published or dropped after commit), and say the boundary (only Hibernate's own writes on managed entities). Candidates who stop at "Hibernate updates the cache automatically" miss both the timing and the large class of writes it cannot see.
- If the transaction rolls back after the UPDATE was already flushed, what is in the cache?Nothing for that key. Hibernate's after-completion callback removes the entry on rollback rather than restoring the old value, because it cannot be sure what the database ended up with. The next reader takes a miss and reloads from the database, which is correct but slightly slower.
- A database trigger rewrites a column whenever the row is updated. What does the cache hold?Potentially the value Hibernate wrote, not the value the trigger produced, because Hibernate caches the state it believes it persisted. The usual fixes are to mark the column as generated so Hibernate re-reads it, to exclude that entity from the cache, or to refresh the entity after write.
- Does write-through mean another session can read your changes before you commit?No. The cache entry is invalidated or soft-locked during flush and only republished after commit, so other sessions either miss and hit the database (where the change is still uncommitted and invisible under normal isolation) or read the pre-change value. The cache never publishes dirty state.
saying these in an interview costs you the question
- Saying the cache is updated at flush time, so other sessions see the new value immediately
- Claiming Hibernate keeps the cache correct for all writes, including native SQL and external jobs
- Believing a rollback restores the previous cached value
- Thinking the second-level cache participates in database transaction isolation the way the database does
- Confusing the shared cache with the per-session persistence context