skip to content

How does Hibernate determine which loaded objects changed, why does the UPDATE statement not appear in the SQL log at the moment you call the setter, and what does this detection cost when one transaction has loaded tens of thousands of rows?

level: middleimportance: must knowfreq 55%

answer

  1. Loaded-state array kept beside the entity
  2. Per-property, type-aware comparison at flush
  3. Setter queues nothing; action queue drained at flush
  4. Deferral → coalescing + JDBC batching + statement reuse
  5. Cost = entities × properties, every flush

basics

~20 s

Hibernate keeps a copy of each entity's values from load time and compares them property by property at flush. Statements are not sent when you call a setter: they go into an action queue and are executed at flush, which allows batching. Cost is proportional to loaded entities times mapped properties, in CPU and in snapshot memory.

solid answer

~60 s

**Detection.** Loading an entity stores a loaded-state array — the snapshot — next to it in the persistence context. At flush, Hibernate iterates the managed entities and compares each property with its snapshot value using the mapped type's equality rules. Differences mark the entity dirty. **Write-behind.** Nothing is sent when the setter runs. Dirty entities produce update actions appended to an internal action queue; the queue is executed at flush, so the SQL appears late, grouped, and in Hibernate's chosen order. Deferring is what enables JDBC batching, statement reuse, and coalescing several mutations of one entity into a single UPDATE. **Cost.** Flush is O(managed entities × properties), and every managed entity carries roughly a second copy of its data. Load 50,000 rows and every flush re-scans all of them — with automatic flush mode, that can be per query. Mitigations: read-only loading (no snapshot, no dirty check), scalar or DTO projections for reporting queries, `StatelessSession` for bulk work, `clear()` between batches, and keeping the persistence context small in the first place.

code

java · 5 lines
java
List<Order> rows = em.createQuery("select o from Order o where o.status = :s", Order.class)
        .setParameter("s", Status.SHIPPED)
        .setHint("org.hibernate.readOnly", true)   // no snapshot kept
        .getResultList();
// these entities are never dirty-checked and cannot be updated

go deeper

for a junior

Say that Hibernate keeps a copy from load time, compares it at flush, and sends the queued SQL then rather than at the setter.

for a middle

Explain the action queue and what deferral enables (coalescing, batching, statement reuse), and give the entities-times-properties cost model.

for a senior

Diagnose real flush-cost problems: automatic flush inside loops, oversized contexts, and the fixes — read-only, projections, clear() batching, StatelessSession.

for a principal

Set loading policy as a design rule: which paths may materialise entities at all, where set-based statements replace object loading, and how context size is bounded in long-running jobs.

## The snapshot When Hibernate hydrates an entity from a result set, it builds the object and also keeps an array of the values it read — the **loaded state**. It lives in the persistence context's bookkeeping entry for that entity, not in the object itself, so application code cannot accidentally modify it. This is the reference point for every later decision about that entity: what changed, and — for features that need old values, such as versionless optimistic locking or certain audit hooks — what the previous values were. ## The comparison At flush, Hibernate walks the managed entities and, for each, compares the current property values against the snapshot. Comparison is delegated to the mapped **type** for each property, not to blanket `equals` on the entity: a basic type compares by value, an association compares by the referenced identifier, an embeddable compares component-wise, a collection is handled by its own persistent-collection snapshot. Any property that reports a difference makes the entity dirty. Nothing dirty means no statement — a transaction that only reads produces no writes, even with thousands of managed entities. ## Write-behind and the action queue A setter never sends SQL. When flush decides an entity is dirty it creates an update action and appends it to the persistence context's action queue; inserts, deletes and collection operations queue up the same way. The queue is drained at flush, which is what people mean by **transactional write-behind**: the database is told about your changes as late as is consistent with correctness rather than statement by statement. Deferral buys concrete things. Multiple mutations to the same entity within a transaction collapse into one UPDATE, because only the end state is compared. Statements of the same shape can be added to a JDBC batch and sent in one round trip instead of many. Prepared statements are reused because the SQL for an entity's update is precompiled and constant (which is also why the default UPDATE lists all columns). And Hibernate retains freedom to order the work sensibly rather than in the arbitrary order your code happened to touch objects. The practical implication for debugging: the line number where SQL appears in the log has almost nothing to do with the line that changed the data. Do not reason about “when did this write happen” from statement order alone; reason about flush boundaries. ## What it costs Two costs, both proportional to what you loaded: - **Memory.** Every managed entity keeps its snapshot, so the persistence context holds roughly twice the field data of the entities in it, plus per-entity bookkeeping. A query that materialises 100,000 rows to compute a sum is paying for 200,000 rows' worth of state. - **CPU at flush.** Each flush re-scans every managed entity, property by property, whether or not anything changed. One flush over 50,000 entities with 20 properties is a million comparisons — and with automatic flush mode, a loop that runs a query per iteration triggers a flush per iteration, turning a linear job quadratic. “The method got slower as the data grew, and profiling points at flush” is the classic symptom. ## How to keep it cheap 1. **Do not load entities you will not modify.** Select the columns you need — a constructor expression into a DTO, or scalar projections — and no entity enters the persistence context at all. This is the single biggest win for report-shaped queries. 2. **Load read-only when you do need objects.** Hibernate's read-only mode (`Session.setReadOnly`, or the `org.hibernate.readOnly` query hint) skips keeping the snapshot, so those entities are never dirty-checked and cannot be updated. Memory and flush cost both drop. 3. **Use a `StatelessSession` for bulk work.** It has no persistence context, no snapshot, no dirty checking and no first-level cache — you issue explicit inserts and updates. Right for migrations and bulk pipelines, wrong for normal domain logic. 4. **Bound the context in batch loops.** Flush and `clear()` every N items so the scanned set stays small and memory is released. Without this, each successive flush is more expensive than the last. 5. **Prefer set-based statements for wholesale changes.** A bulk JPQL `UPDATE` touching a million rows beats loading a million entities — at the cost of bypassing the persistence context, which then holds stale copies and must be cleared. 6. **Consider bytecode enhancement** if the context genuinely must be large: instrumented setters record which attributes changed, so flush no longer walks snapshots. ## The sentence to land Dirty checking is a convenience whose cost scales with *what you loaded*, not with *what you changed*. Keeping persistence contexts small is therefore not tidiness — it is the main lever on flush cost.

  • Why does the default generated UPDATE list every mapped column rather than only the changed ones?
    Because the SQL for an entity's update is generated once and reused, which allows prepared-statement caching on both the driver and the database and lets consecutive updates of the same entity type join one JDBC batch. Emitting only changed columns would require a different statement per combination of changed fields, defeating both. Hibernate offers dynamic update generation when the trade is worth reversing.
  • Why can a loop that runs a query per iteration make flush cost blow up?
    Under the default automatic flush mode, Hibernate flushes before a query whose results could be affected by pending changes, so each iteration re-scans every managed entity accumulated so far. The work grows quadratically with the number of loaded entities. Keeping the context small, loading read-only, or restructuring to a single query removes the effect.
  • When is StatelessSession the right tool rather than tuning the persistence context?
    For bulk pipelines — imports, migrations, mass recalculations — where you do not need identity, dirty checking, cascades or the first-level cache, and where explicit inserts and updates are clearer. You give up the object-graph conveniences entirely, so it is a poor fit for ordinary domain logic; there the answer is to load less and clear regularly.

saying these in an interview costs you the question

  • Thinking Hibernate compares the object with the database row at flush rather than with an in-memory snapshot
  • Assuming a transaction that only reads still generates updates
  • Believing the SQL log's ordering reflects where the mutation happened in code
  • Ignoring that flush cost scales with entities loaded, not entities changed
  • Reaching for a second-level cache to fix what is really an oversized persistence context

context