The same JPQL query, executed twice inside one persistence context, returns a different number of rows the second time, yet the entities that were already loaded still show their old field values. Explain both halves of that behaviour.
answer
- Identity map is by id, not by query
- Queries always hit the database
- Phantom rows appear per isolation level
- Managed instance wins, read row discarded
- Scalars bypass the identity map
basics
~20 sQueries always execute SQL — the first-level cache is keyed by id, not by query — so rows committed by others appear or disappear as the isolation level allows. But rows whose entities are already managed are resolved to the existing instances, and the freshly read column values are discarded.
solid answer
~50 sTwo different mechanisms are at work. Result-set membership comes from the database. Hibernate has no query cache by default, and the persistence context is an identity map keyed by entity type and id, which cannot answer an arbitrary query. So the SELECT runs again and returns whatever your transaction is allowed to see: under READ COMMITTED that includes rows another transaction committed a moment ago, which is exactly a phantom read. Under a snapshot-based repeatable level the second execution would see the same set. Entity state comes from the persistence context. As Hibernate turns each row into an object it checks the identity map first; on a hit it returns the managed instance and throws away the columns it just read, because two instances per row would break identity and because the managed one may hold unflushed changes. The result is fresh membership with possibly stale contents. Diagnosing staleness by re-running a query therefore does not work; use refresh(), clear(), or a new persistence context.
code
java · 7 linesList<Order> first = em.createQuery("from Order where status = :s", Order.class)
.setParameter("s", Status.NEW).getResultList(); // 3 rows
// another session: insert a new NEW order, and update order 42's total
List<Order> second = em.createQuery("from Order where status = :s", Order.class)
.setParameter("s", Status.NEW).getResultList(); // 4 rows
// the new order is current; order 42 still shows the total loaded the first time
em.refresh(second.get(0)); // explicit re-read of one entitygo deeper
Know that queries always go to the database while find() by id may not, and that refresh() is how you re-read an entity.
Explain the identity map's scope, why a query cannot be answered from it, and why the managed instance wins over freshly read columns.
Connect membership to the isolation level, contents to the persistence context, and describe the mixed-vintage result lists and count-versus-collection confusion that follow, plus the levers to fix them.
Set policy on persistence-context lifetime and on where read consistency is required, choosing projections for reporting paths and reserving stable-snapshot transactions for the few places that genuinely need them.
## Half one: why the row count changes The first-level cache is often described as a cache, which invites the wrong expectation. It is an identity map from entity type plus primary key to a managed instance. It can answer "give me Order 42". It cannot answer "give me all orders created today", because it has no index over arbitrary predicates and no knowledge of rows it has never loaded. So every JPQL, Criteria or native query executes SQL. What the second execution sees is decided entirely by the database, which is the boundary this topic bridges to: at READ COMMITTED each statement takes a fresh view of committed data, so rows inserted and committed by another transaction between the two executions appear, and rows deleted vanish. That is the textbook phantom read, and it is visible through an ORM exactly as it is through plain JDBC. On engines and levels that give a transaction a stable snapshot, both executions see the same set. Nothing the ORM does changes that. One ORM-specific twist: with the default flush mode, a query triggers an automatic flush when the pending changes touch the tables the query reads. So the second execution can also differ because of your own unflushed inserts, not only because of other transactions. ## Half two: why the field values do not change Hibernate hydrates each returned row and then asks the persistence context whether that identifier is already managed. If it is, the existing instance wins and the newly read values are discarded. That rule is not an optimisation, it is a correctness requirement. The unit of work guarantees a single instance per row, so code can compare references and Hibernate has one place to hold the loaded-state snapshot used for dirty checking. If a query overwrote instances, a user's unflushed edits could be silently reverted by an unrelated query executed in the same request, and the snapshot would no longer describe the state the row had when it was read, corrupting the UPDATE that dirty checking generates. ## The combined effect and how it misleads You get a result list whose membership is current and whose contents may be minutes old. Typical confusions that follow: - A count query disagrees with the size of an in-memory collection, because the count is fresh while the collection was materialised earlier. - A developer re-runs a finder to reload an entity and concludes the database is wrong when the object still shows old values. - A newly appeared row is fully current — it was never in the identity map — sitting next to a stale sibling, so one list mixes two vintages. - Aggregate or projection queries (selecting scalars rather than entities) always show current data, because scalars are not entities and never pass through the identity map, so they can contradict the entities beside them in the same request. ## Getting the behaviour you actually want - To refresh known entities, call refresh() on each, or clear() the context and re-query. refresh() discards unflushed changes to that instance, which is the point. - To guarantee a stable set across several queries, put them in one transaction and rely on the database level that provides it, accepting the retry semantics that come with it. - To avoid mixed vintages entirely, keep the persistence context short: one per unit of work, closed at the end. Long-lived contexts are the root cause in nearly every incident of this shape. - For read-only reporting where consistency across queries matters more than object identity, consider selecting projections rather than entities: scalars bypass the identity map, so what you read is what the database returned. ## The framing to give in an interview Say it in one sentence: membership of a query result is decided by the database and the isolation level; the state of an entity in that result is decided by the persistence context, which prefers the instance it already has. Then name refresh, clear and short contexts as the levers. That shows you know where the ORM boundary sits without re-teaching isolation theory.
- Why doesn't Hibernate simply overwrite managed entities with the values a query just read?Because the persistence context must keep one instance per row and that instance may hold unflushed modifications. Overwriting would silently discard the application's changes and invalidate the loaded-state snapshot that dirty checking uses to generate the UPDATE, producing wrong SQL. refresh() exists as the explicit opt-in when re-reading is what you want.
- How would you make several queries in one request see a consistent set of rows?Run them inside a single transaction at an isolation level that gives the transaction a stable view, and accept that the engine may then fail the transaction with a serialization error that you must retry. Alternatively, reduce the requirement: fetch what you need in one query, or accept per-statement freshness where the business tolerates it.
saying these in an interview costs you the question
- Saying the first-level cache serves query results
- Expecting a re-run query to refresh already-loaded entities
- Blaming the ORM for phantom rows that the isolation level permits
- Assuming a count query and an in-memory list must agree inside one persistence context
- Using clear() casually without realising it detaches everything and discards pending changes