skip to content

For a read-only list screen, why would you project query results into DTOs rather than loading the mapped entities and reading their fields?

level: middleimportance: must knowfreq 50%

answer

  1. entity = all columns + context entry + snapshot copy
  2. dirty check cost scales with context size, not with edits
  3. DTO = unmanaged, no snapshot, no flush
  4. setter on managed entity → surprise UPDATE
  5. commands load entities, queries project

basics

~20 s

Loading entities selects every mapped column, puts each object in the persistence context, and stores a snapshot copy for dirty checking that Hibernate must compare at flush. A DTO query selects only the needed columns, allocates one small object per row, and is never dirty-checked or flushed.

solid answer

~60 s

Three costs come with a managed entity that a read model does not need: 1. **Wide selects.** `select o from Order o` reads every mapped column, including large text columns and every eagerly mapped to-one — for a list showing four fields. 2. **Persistence-context footprint.** Every returned entity is put in the context, *plus* a loaded-state snapshot array Hibernate keeps to detect changes. Roughly two copies of every row you read. Ten thousand rows for a report is a memory event. 3. **Dirty checking on flush.** At each flush, Hibernate walks the managed entities and compares each property against its snapshot. Work proportional to what you loaded, even though nothing was modified. On top of that, entities in a read path invite accidents: navigating an association triggers extra queries, a stray setter turns a read into an UPDATE, and cascades apply. A constructor-expression or Tuple projection returns unmanaged objects: only the requested columns are selected, no snapshot exists, flush ignores them, and the DTO's fields state exactly what the screen consumes — which makes the read contract explicit instead of "whatever the entity happens to expose".

code

sql · 9 lines
sql
-- entity query: every mapped column, plus eager to-ones
select o.id, o.created_at, o.status, o.total, o.notes, o.payload,
       o.customer_id, o.shipping_address_id, o.version
from orders o where o.status = ?;

-- projection: only what the grid renders
select o.id, o.created_at, c.name, o.total
from orders o join customers c on c.id = o.customer_id
where o.status = ?;

go deeper

for a junior

State the two headline reasons — fewer columns selected, and the objects are not managed so nothing is dirty-checked.

for a middle

Name the snapshot copy and the flush-time comparison, the eager-association over-fetch, and the accidental-UPDATE risk.

for a senior

Draw the command/query line, weigh read-only sessions and lazy basic attributes as partial mitigations, and show how you would measure the difference.

for a principal

Argue the read-model boundary as an architectural rule: who owns projections, how they version with the UI, and when a separate read path or store becomes justified.

## The hidden bill on a managed entity When a query returns entities, Hibernate does considerably more than materialise objects. **It selects every mapped column.** The SQL for `select o from Order o` lists all basic properties plus the foreign keys of to-one associations, and if any association is mapped `FetchType.EAGER` it joins or issues extra selects for those too. A grid that shows id, date, customer name and total may be dragging a `text` description column and a serialized payload across the wire for every row. **It stores each entity in the persistence context**, keyed by identity, so the session can guarantee that one row equals one object. **It stores a loaded-state snapshot** alongside — an array of the property values as they were read. This exists solely so that at flush time Hibernate can compare current values against the snapshot and generate UPDATEs for what changed. So a read of N rows with M columns costs roughly 2×N×M values retained until the session ends or is cleared. **It pays that comparison at every flush.** Automatic dirty checking iterates the managed entities and compares property by property. Nothing was modified, but the work is done anyway, and it grows with the size of the context, not with the size of your change. ## What a DTO projection changes ```sql select o.id, o.created_at, c.name, o.total from orders o join customers c on c.id = o.customer_id where o.status = ? ``` One narrow row per result, one small object per row, no context entry, no snapshot, no flush participation, no cascade, no interaction with the second-level entity cache. The result is a value: immutable if you use a record, safe to hand to a serializer, safe to cache in your own layer, and impossible to accidentally mutate into an UPDATE. ## The correctness arguments, not just the performance ones **Accidental writes.** A managed entity is a write handle. Any code path that calls a setter — a mapper normalising a value, a formatter trimming a string — produces an UPDATE at the next flush. Read models cannot do this by construction. **Uncontrolled query fan-out.** Rendering a list from entities tempts navigation into associations, and each navigation on a lazy association is another query per row. A projection has to name every value up front, which forces the join to be explicit and the row count to be predictable — the N+1 problem is designed out rather than watched for. **Explicit contract.** The DTO's field list *is* the read contract. When a UI field is dropped, the projection changes and the SQL narrows. With entities, nobody can tell which columns a screen actually uses, so nothing can ever be safely made lazy or removed. **Stability under mapping change.** Adding a heavy column to an entity silently slows every entity-returning query in the system. Projections are unaffected unless you ask for the new column. ## When entities are still right This is not "never load entities". Load them when you intend to **change** something: that is what the persistence context, dirty checking and cascades are for, and hand-writing UPDATE statements to avoid them is a much worse trade. Load them for small results where the difference is immaterial and the code is simpler. Load them when domain behaviour — invariants, methods on the aggregate — has to run. The rule of thumb that survives review: **command paths load entities; query paths project.** The bigger the result set and the wider the entity, the more decisive the rule becomes. ## Middle grounds, and their limits - **Read-only sessions/queries.** Hibernate can mark results read-only (for example `org.hibernate.readOnly`), which skips the snapshot and dirty checking. That removes the flush cost but not the wide select, and the objects still occupy the context. - **Lazy basic attributes.** Bytecode enhancement can make a heavy column lazy, narrowing the select — powerful, but it needs enhancement configured and applies globally to that attribute. - **Clearing the session** between batches bounds memory but does nothing for the columns you read. A projection addresses all of it at once and needs no configuration, which is why it is the default for read models. ## The counterargument to answer in an interview Someone will say projections duplicate the model and multiply classes. The honest answer: yes, a DTO per read use case is more classes, and that is the point — each states exactly one screen's needs and can change with that screen without touching the domain model. The failure mode to avoid is one shared "fat DTO" that unions every screen's fields, which reproduces the entity's problems without its benefits.

  • Marking the query read-only removes dirty checking. Does that make loading entities as good as projecting?
    It removes one of the three costs. Read-only results skip the loaded-state snapshot and are not dirty-checked at flush, which is a genuine saving on large reads. But the SQL still selects every mapped column, the objects still sit in the persistence context, and lazy associations can still be navigated into per-row queries. A projection removes all of that and additionally documents which fields the screen consumes.
  • Doesn't a DTO per screen just duplicate the entity and multiply classes?
    It adds classes, and that is the intended trade: each DTO states one use case's read contract and can evolve with that screen without touching the domain model. The anti-pattern is one shared fat DTO unioning every screen's fields, which recreates the entity's over-fetching without its behaviour. Records keep the per-use-case classes cheap enough that the duplication argument rarely wins.
  • How would you actually demonstrate the difference rather than assert it?
    Turn on SQL logging or Hibernate statistics and compare the two implementations of the same screen: the column list in the emitted SQL, the number of statements executed, and the entity count in the session. Add a heap snapshot or a simple allocation measurement for a large result to show the snapshot copies. Numbers from the real table settle the argument far better than reasoning about it.

Loading entities for a report is like checking books out of the library to read one line from each: the librarian records every loan and must check each book back in, even though you changed nothing. A projection photocopies the line.

saying these in an interview costs you the question

  • Claiming entity loading and DTO projection issue the same SQL
  • Not knowing Hibernate keeps a separate loaded-state snapshot for dirty checking
  • Saying you should never load entities, including on write paths
  • Treating a single shared fat DTO as equivalent to per-use-case projections
  • Assuming a read-only flag removes the over-fetching of columns as well

context