skip to content

Dirty Checking & Write-Behind

How Hibernate notices your changes without a save() call — snapshot comparison at flush — and the write-behind queue that batches work. Interviewers ask because accidental UPDATEs from mutated managed entities are a classic production surprise.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

questions

5

Inside a transaction you load an entity with EntityManager.find, call a setter on it, and never call persist, merge, or any save method — yet an UPDATE reaches the database when the transaction commits. Explain the mechanism that makes that happen.

level: juniorimportance: must knowfreq 76%

answer

  1. Managed = tracked in the persistence context
  2. Snapshot copied at load; setters never touch it
  3. Flush compares object vs snapshot property by property
  4. Different → dirty → UPDATE queued
  5. Detached entity: same setter, no write

basics

~20 s

A loaded entity is managed by the persistence context, which kept a snapshot of the values it was loaded with. At flush (commit, by default) Hibernate compares the object with that snapshot, sees the changed field, and generates the UPDATE automatically. This is dirty checking; no save call is involved.

solid answer

~50 s

When `find` returns an entity, it becomes **managed**: the persistence context stores it by identity and keeps a **loaded-state snapshot** — a copy of the values read from the row. Your setter changes the object but not the snapshot. At flush time, Hibernate walks the managed entities and compares each one against its snapshot property by property. Anything that differs is **dirty**, and Hibernate queues an update action for it. Flush happens automatically before the transaction commits (and, with the default flush mode, before queries that might read affected tables), so the UPDATE is emitted without any explicit call. Two consequences worth stating. First, calling a save/merge method on an already-managed entity is redundant — the change is tracked whether you call it or not. Second, and less comfortable: **any** mutation of a managed entity is a database write, including one made accidentally deep inside business logic or in a read-oriented code path. Managed objects are not scratch space.

code

java · 4 lines
java
tx.begin();
Order order = em.find(Order.class, 7L);   // managed; snapshot taken
order.setStatus(Status.PAID);              // object changed, snapshot unchanged
tx.commit();                               // flush compares -> UPDATE orders set ... where id=7

go deeper

for a junior

State the mechanism plainly: loaded entities are managed, a snapshot was taken at load, flush compares and emits the UPDATE — no save call needed.

for a middle

Add when flush happens, that comparison is per property, and that only managed entities are tracked so detached mutations are silently lost.

for a senior

Emphasise the cost side and the accidental-write failure mode, and how read-only loading or projections keep read paths from generating writes.

for a principal

Turn it into a convention: which layers may hold managed entities, immutability of loaded objects in read paths, and how that boundary is enforced so unintended writes cannot occur.

## The three ingredients **1. Managed state.** When an entity is loaded (`find`, a query, or navigation from another entity) or persisted, the persistence context — Hibernate's `Session`, exposed as JPA's `EntityManager` — keeps a reference to it in a map keyed by entity type plus identifier. That map is the first-level cache, and membership in it is what “managed” means. A managed entity is under observation for the lifetime of the persistence context. **2. The loaded-state snapshot.** At the moment of loading, Hibernate also copies the entity's property values into an array stored alongside the entity in its internal bookkeeping (the `EntityEntry`). This snapshot is what the row looked like when read. Your code never sees it, and your setters do not touch it — which is exactly why it is useful. **3. Flush.** Flush is the act of synchronising the persistence context with the database: Hibernate figures out which SQL statements are needed and sends them. It happens automatically before the transaction commits, and — under the default automatic flush mode — before executing a query whose result could be affected by pending changes. ## What actually happens at flush For each managed entity, Hibernate compares the current property values against the snapshot. Comparison is per property and type-aware: a string is compared by value, a number by value, an association by the referenced identifier, and so on. If nothing differs, the entity is clean and no statement is produced — loading a thousand rows and changing none produces no writes. If something differs, the entity is dirty and Hibernate creates an update action for it, which is executed when the queued actions are flushed to JDBC. The statement is then a normal `UPDATE ... SET ... WHERE id = ?`, with every mapped column in the SET list by default. So the sequence for the question asked is: `find` → entity managed, snapshot taken → setter changes the object → commit triggers flush → comparison finds the changed property → UPDATE emitted → transaction commits. ## Why JPA is designed this way The programming model aims to let you manipulate objects and have persistence follow. If every field change required an explicit save call, the API would be a thin wrapper over SQL and the object model would be permanently at risk of being out of sync with what was actually written. Automatic detection also lets Hibernate batch and order the resulting statements: it decides *when* to write, so it can group updates, reuse prepared statements, and reduce round trips. ## The consequences that get asked about next **Calling save on a managed entity is a no-op in effect.** `merge` on an entity that is already managed returns the same instance; the change would have been written regardless. Code that calls it “to be safe” creates the false impression that saving is what persists the change. **Accidental writes are real.** Because *any* mutation counts, code that touches a managed entity for non-persistence reasons — normalising a string before rendering, setting a transient-looking field, applying a default — issues an UPDATE. Common examples: a method that lowercases an email “just for comparison”, a mapper that writes back into the entity it read from, or a lifecycle hook that stamps a timestamp on every read path. The symptom is UPDATE statements in read-only flows and unnecessary row versions. **The scope is the persistence context, not the object.** Once the context closes, the entity is detached and mutations are no longer tracked; the same setter then has no effect on the database at all. The identical line of code either writes or does not depending on where it runs — which is why understanding managed versus detached matters as much as understanding dirty checking. **Detection is not free.** Comparing every managed entity against its snapshot at every flush costs CPU proportional to entities loaded times properties mapped, and holding snapshots costs roughly twice the memory of the entity data. That is the reason for read-only loading, projections for report-style queries, and keeping persistence contexts small in batch code. ## How to demonstrate it in an interview Say it in one sentence — “the persistence context keeps a snapshot from load time and compares at flush; different means dirty means UPDATE” — then add the sharp edge: it applies to every mutation, intended or not, and only while the entity is managed.

  • If dirty checking writes changes automatically, when do you still need merge?
    Only when the object is not managed by the current persistence context — typically a detached instance coming back from another transaction or from outside the application. For an entity you just loaded in the same transaction, merge adds nothing: it returns the same managed instance and the change would have been written anyway.
  • You see UPDATE statements in a code path that only reads data. What is the likely cause?
    Something is mutating a managed entity along the way — a normalisation step, a mapper writing back into the entity it read, a default being applied, or a mutable value whose comparison flips. Because every mutation of a managed entity is a write, read paths must treat loaded entities as immutable, or load them read-only so no snapshot comparison occurs.

A librarian photocopies a page when you check the book out and compares your copy against the original when you return it — any difference gets recorded, whether or not you meant to write in it.

saying these in an interview costs you the question

  • Believing a save or merge call is required for the change to be persisted
  • Thinking the UPDATE is sent at the moment the setter runs
  • Assuming only fields you “intended” to change are written
  • Not realising that the same setter on a detached instance does nothing
  • Calling merge on an already-managed entity and thinking it is what saved the change

context

open as a page

How does Hibernate determine which loaded objects changed, why does the UPDATE statement not appear in the SQL log at the moment you call the setter, and what does this detection cost when one transaction has loaded tens of thousands of rows?

level: middleimportance: must knowfreq 55%

basics

~20 s

Hibernate keeps a copy of each entity's values from load time and compares them property by property at flush. Statements are not sent when you call a setter: they go into an action queue and are executed at flush, which allows batching. Cost is proportional to loaded entities times mapped properties, in CPU and in snapshot memory.

open as a page

By default the UPDATE Hibernate generates for a changed entity sets every mapped column, not just the ones that were modified. Why is it built that way, and in what situations does switching that entity to Hibernate's @DynamicUpdate pay for itself?

level: middleimportance: should knowfreq 38%

basics

~20 s

Hibernate precompiles one UPDATE per entity so the statement can be cached and batched — hence all columns. @DynamicUpdate builds the SQL per flush from the changed columns only, which helps with wide tables, large LOB or JSON columns, column-level triggers, and versionless optimistic locking, at the cost of per-flush SQL generation and weaker statement reuse.

open as a page

Your logs show Hibernate issuing an UPDATE for the same entity in transactions where the application never modifies it. What kinds of mapping or type problems cause these repeated no-op updates, and how do you pin down which column is responsible?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Some property compares unequal to its loaded snapshot even though nothing meaningful changed — typically a type whose write-then-read round trip is not value-preserving: BigDecimal scale, date/time precision or class mismatch, a custom or JSON type without correct equality and deep copy, trimmed CHAR padding, or a mutable value replaced on read. Temporarily enabling dynamic updates names the column in the SQL.

open as a page

Hibernate can be configured with build-time bytecode enhancement so entities track their own modifications instead of being compared against snapshots at flush. What changes at runtime when you enable it, and when is that trade actually worth making?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Enhancement instruments entity classes so each field write records the attribute name in a tracker on the object. At flush Hibernate asks the entity which attributes are dirty instead of walking snapshots, making flush proportional to what changed rather than what was loaded. Worth it only for large persistence contexts or wide entities; mutating a value object in place can go undetected.

open as a page