skip to content

By default the UPDATE Hibernate generates for a changed entity sets every mapped column, not just the ones that were modified. Why is it built that way, and in what situations does switching that entity to Hibernate's @DynamicUpdate pay for itself?

level: middleimportance: should knowfreq 38%

answer

  1. Static SQL generated once per entity → cache + batch
  2. Dynamic = SQL built per flush from dirty properties
  3. Wins: wide tables, LOB/JSON, triggers/CDC noise
  4. Required by OptimisticLockType.DIRTY/ALL
  5. Costs: per-flush generation, cache misses, split batches

basics

~20 s

Hibernate precompiles one UPDATE per entity so the statement can be cached and batched — hence all columns. @DynamicUpdate builds the SQL per flush from the changed columns only, which helps with wide tables, large LOB or JSON columns, column-level triggers, and versionless optimistic locking, at the cost of per-flush SQL generation and weaker statement reuse.

solid answer

~60 s

The default UPDATE is **static**: generated once when the persistence unit starts, listing every mapped column, so it can be reused for every instance of that entity. That reuse is what enables driver and database prepared-statement caching and lets many updates of the same entity type go out in one JDBC batch. `@DynamicUpdate` tells Hibernate to build the statement at flush time from the properties dirty checking found changed. Cases where that wins: - **Wide entities** where a typical change touches a couple of columns out of fifty — less SQL, fewer bind parameters, less redo/WAL. - **Large columns** (LOB, TEXT, JSON) that would otherwise be rewritten untouched on every update. - **Column-level triggers or audit/CDC** that fire on any assignment, producing noise when unchanged columns are rewritten. - **Versionless optimistic locking** (`OptimisticLockType.DIRTY`), which requires dynamic updates because the WHERE clause is built from the changed columns' old values. Costs: SQL is composed per flush, statement caching hits less often, and batching is only possible across updates that changed the same column set — so it is a per-entity decision, measured, not a global default.

code

java · 12 lines
java
@Entity
@DynamicUpdate
class Document {
    @Id Long id;
    String title;
    @Lob String body;      // rewritten on every update without @DynamicUpdate
}

@Entity
@DynamicUpdate
@OptimisticLocking(type = OptimisticLockType.DIRTY)  // needs dynamic updates
class LegacyRow { /* no version column available */ }

go deeper

for a junior

Know that the default UPDATE writes all mapped columns because the SQL is generated once, and that @DynamicUpdate narrows it to what changed.

for a middle

Explain the reuse-and-batching rationale for the default and name concrete wins for dynamic updates: wide rows, large columns, trigger and audit noise.

for a senior

Weigh the costs — statement-cache misses, split batches, per-flush generation — make it a per-entity call backed by measurement, and connect it to versionless optimistic locking and to diagnosing phantom updates.

for a principal

Treat it as one knob in a write-path budget: statement shapes, batch sizes, WAL/redo volume and downstream CDC noise, decided per entity against the workload rather than as a house style.

## Why the default writes every column Hibernate generates the CRUD SQL for each entity once, at bootstrap. For updates, that means a single statement shape per entity: `update t set c1=?, c2=?, ..., cn=? where id=?` (plus `and version=?` when versioned). One shape has several benefits: - **Prepared-statement reuse.** The same SQL string is reused across every update of that entity, so both the JDBC driver's statement cache and the database's plan cache hit rather than compiling a new statement per distinct set of changed columns. - **Batching.** JDBC batching requires statements of identical shape. With a static UPDATE, a hundred dirty entities of the same type flush as one batch; with per-change shapes they scatter into many batches, or none. - **Predictability.** The plan for the statement is stable, and the bootstrap cost of SQL generation is paid once. The price is writing columns whose values did not change. For narrow entities on ordinary types that is usually irrelevant — the row is rewritten anyway by most storage engines, so the cost is mostly bind parameters and network bytes. ## What @DynamicUpdate changes Annotating the entity with `@DynamicUpdate` makes Hibernate compose the UPDATE at flush from exactly the properties that dirty checking reported as changed. Change one column of forty and the statement sets one column. Where that matters: **Wide tables.** With 50–200 columns, rewriting everything on every small edit means bigger statements, more bind parameters, and more work in the engine's write path. **Large values.** A CLOB/BLOB/JSON/TEXT column of hundreds of kilobytes that was not touched still gets sent and rewritten under a static update: network traffic, potential out-of-line storage churn, and larger write-ahead log volume. Excluding it is a real saving. **Trigger and change-data noise.** Triggers, audit tables and CDC pipelines that key off “which columns appeared in the statement” see every column on every update, producing audit rows and downstream events for changes that did not happen. Dynamic updates make the statement reflect intent. **Reduced write footprint under concurrency.** This is the nuanced one. A dynamic update does **not** make lost updates impossible — two transactions writing the same column still race, and that is what a version attribute is for. What it does remove is the *needless* overwrite: a transaction that only edited a note no longer rewrites a status another transaction changed since it read. In systems without versioning it materially reduces silent clobbering; it is not a substitute for a proper concurrency token. **Versionless optimistic locking.** Hibernate's `@OptimisticLocking(type = OptimisticLockType.DIRTY)` (and `ALL`) builds the WHERE clause from the previously loaded values of the changed (or all) columns, so the update only succeeds if the row still looks as it did at load. Because the WHERE clause varies with what changed, this strategy requires dynamic updates. It is the escape hatch for legacy tables where adding a version column is not possible — with the caveat that it cannot detect a change made and undone, and it makes statement shapes explode. ## The costs, stated honestly - **SQL generated per flush.** String building and parameter binding for each dirty entity, instead of reusing one prepared shape. - **Weaker caching.** Distinct column combinations mean distinct SQL strings, so driver and database statement caches see many shapes; on databases where hard parsing is expensive, this can outweigh the saved bytes. - **Weaker batching.** Only updates that changed the same columns can share a batch, which in a bulk job is often the case (the same field is being updated everywhere) but in mixed workloads is not. ## How to decide It is a **per-entity** annotation, and that is how it should be used. Reach for it when the entity is wide, holds large columns, or feeds triggers/CDC — or when versionless optimistic locking is required. Leave it off for small entities and for entities updated in bulk with uniform column sets, where batching dominates. As with all such switches, measure: statement counts and shapes, batch sizes, and bytes written before and after. A useful side effect worth mentioning: turning it on temporarily is a great **diagnostic**. If an entity is producing mystery updates, a dynamic update names the offending column in the SQL log immediately, which a static all-columns statement never does.

  • Does @DynamicUpdate remove the need for a @Version column?
    No. It narrows the set of columns you overwrite, so you stop clobbering fields you never touched, but two transactions editing the same field still race with last-writer-wins. Only a concurrency token — a version attribute, or versionless dirty-column checking, which itself needs dynamic updates — causes the second write to fail rather than silently win.
  • Why might enabling it globally make a bulk job slower?
    Bulk jobs rely on JDBC batching, which requires identical statement shapes. If the changed column sets vary across entities, dynamic generation splits one large batch into many small ones and multiplies distinct SQL strings, hurting driver and database statement caches. The per-flush SQL construction adds CPU on top. Enable it on the entities that benefit, not everywhere.

saying these in an interview costs you the question

  • Recommending @DynamicUpdate as a global default without measuring
  • Claiming it prevents lost updates or replaces optimistic locking
  • Not knowing that versionless dirty-column locking depends on it
  • Believing the default statement is rebuilt per update
  • Assuming batching is unaffected by varying column sets

context