skip to content

What extra work and memory does Hibernate spend on a query that returns managed entity objects, compared with the same query returning a DTO projection?

level: middleimportance: must knowfreq 58%

answer

  1. entity = instance + EntityEntry + loaded-state snapshot
  2. snapshot exists so dirty checking works
  3. flush cost ~ entities × properties, per flush
  4. all mapped columns vs only the selected ones
  5. DTO: no context entry, no snapshot, garbage immediately

basics

~20 s

For each managed entity Hibernate builds the instance, keeps it in the persistence context, and stores a second copy of its loaded state for dirty checking. Every flush then compares all of them property by property. A DTO row costs one small object and none of that.

solid answer

~60 s

Loading an entity costs far more than loading its values. - **All mapped columns are selected**, not just the ones the caller needs, so wider rows and more network bytes. - **The persistence context keeps the instance**, keyed by entity type plus identifier, until the session closes or you clear it. - **A loaded-state snapshot is stored alongside it** — an array of the hydrated property values, roughly doubling per-row memory. That snapshot is what makes automatic dirty checking possible. - **Every flush walks all managed entities** and compares current values against the snapshot, property by property, plus collection dirty checks and cascade traversal. Cost is proportional to entities × properties, and flush can happen more than once per transaction. - Managed entities also carry lazy proxies you can accidentally trigger later, and may be written into the second-level cache. A DTO row is a single immutable object: no snapshot, no persistence-context entry, no flush participation. On a read-mostly endpoint returning thousands of rows that is often the difference between comfortable and heap-bound.

code

java · 10 lines
java
// 20k managed entities + 20k state snapshots retained in the persistence context
List<User> users = em.createQuery("from User u", User.class).getResultList();

em.persist(new AuditLog("listed users"));
// commit -> flush -> Hibernate dirty-checks all 20k entities before the single INSERT

// projection: nothing is tracked, flush has nothing extra to compare
List<UserSummary> rows = em.createQuery(
    "select new com.app.UserSummary(u.id, u.name) from User u", UserSummary.class)
  .getResultList();

go deeper

for a junior

Say that entities are tracked by the persistence context and DTOs are not, and that entities load every column.

for a middle

Name the concrete costs: loaded-state snapshot, entity entries, dirty check at flush proportional to entity count, all mapped columns selected.

for a senior

Tie it to symptoms — heap retained by the session, flush time scaling with result size — and to when entities are still the right choice.

for a principal

Position it as a per-read-path budget decision, including where dirty checking is a feature worth paying for and where it is pure overhead.

## The three costs of a managed read ### 1. Hydration and column width An entity query selects every non-lazy mapped column. If the table has 30 columns and the response needs three, 27 columns are read from disk, serialised by the driver, converted to Java types, and thrown away. Wide `varchar`/`text`/`bytea` columns dominate here, and a narrow projection can additionally let the database serve the query from an index alone rather than touching the heap/row store. ### 2. The persistence context and the state snapshot Every entity Hibernate returns is put into the persistence context (the first-level cache): a map from `EntityKey` (entity name + identifier) to the instance, plus an `EntityEntry` per instance holding status, version, and — the expensive part — the **loaded state**, an `Object[]` snapshot of the property values as read from the database. That snapshot is not an optimisation you can skip; it is the mechanism behind automatic dirty checking. Hibernate has no change interceptors on your POJO fields by default, so at flush time it discovers modifications by comparing current field values against the snapshot. The price is roughly a second copy of every loaded row in memory, plus the entry object itself. Ten thousand entities means ten thousand instances, ten thousand snapshots, and map entries for all of them — all strongly reachable until the session closes or you call `clear()`, so they survive minor garbage collections and go straight into the old generation under load. ### 3. Flush-time dirty checking When the session flushes — before a query that touches affected tables, and at commit — Hibernate iterates every managed entity and compares each property against its snapshot, using the mapped types' `isEqual` logic. It also checks every managed collection for structural changes, and traverses cascades to find newly reachable transient instances. The work is O(entities × properties) *per flush*, and a transaction may flush repeatedly. A request that loaded 20 000 entities purely to render JSON pays that bill even though it modified nothing. There is a secondary hazard: the more managed objects float around, the easier it is to trigger lazy loads row by row and turn one query into hundreds. ## What a DTO read skips A constructor-expression or tuple result is not an entity, so: - only the listed columns are selected; - nothing is registered in the persistence context, so heap use is one small object per row and it becomes garbage as soon as the response is written; - there is no snapshot and therefore no dirty checking — flush cost is unaffected by how many rows you read; - there is no lazy loading, so the query result is the whole truth and no surprise statements can fire later; - second-level cache and collection caches are not exercised for these rows. Roughly, a DTO read turns "objects the ORM must supervise" into "values you happen to have". ## Cases where the entity read is still correct This is not a blanket rule. Read entities when: - you intend to modify them — dirty checking is the feature you are paying for, and a DTO cannot be updated; - the rows are few (a detail page, a lookup by id), where the overhead is noise; - they are served from the second-level cache, so the entity read may not hit the database at all while a DTO query would; - domain behaviour lives on the entity and the calculation needs the real object. The decision is per read path, not per application. ## How to see it in practice Two symptoms point at managed-read overhead. First, heap: a heap dump under load shows large retained sets under the session's persistence context, with an `Object[]` snapshot per entity. Second, time spent in flush that grows with result size rather than with the number of modified rows — a request that changes one row but takes longer as an unrelated list query grows is dirty-checking a bloated context. The reflex fix on a read-only endpoint is to select only what the response needs into a DTO. If the code must keep loading entities in a long loop, bounding the context with periodic `clear()` is the containment measure, but on a pure read path removing the entities altogether is simpler and strictly cheaper.

  • Why does Hibernate keep a snapshot at all instead of tracking changes as they happen?
    Plain mapped POJOs have no change notification, so the only portable way to detect a modified property is to compare the current value with the value that was loaded. Hibernate can avoid the snapshot with bytecode enhancement, which instruments setters to flag dirty attributes, but the default mapping model relies on the snapshot comparison at flush.
  • Reading fewer entities is not always possible. What bounds the cost when a job must iterate many entities?
    Process in chunks and call `flush()` then `clear()` at the end of each chunk so the persistence context does not grow without limit; detached entities then become collectable and later flushes only dirty-check the current chunk. Keep the chunk size aligned with whatever write batching you use, and be aware that clearing detaches everything, so hold identifiers rather than object references across chunks.

saying these in an interview costs you the question

  • Claiming an entity read and a DTO read cost the same because "it is one SQL query either way"
  • Not knowing that Hibernate keeps a second copy of the loaded state per entity
  • Thinking dirty checking only inspects entities you touched
  • Assuming the persistence context releases entities after the query returns
  • Answering "just add the second-level cache" for a read path whose cost is context size, not database round trips

context