Explain the order in which Hibernate looks for an entity when you call em.find(Customer.class, 42L), and why a JPQL query such as "select c from Customer c where c.city = :city" still executes SQL even when every matching row is already in the second-level cache.
answer
- find: context, then shared cache, then database
- Context hit wins even if the shared cache is newer
- Caches are id-keyed maps: no predicate evaluation
- No completeness guarantee: cache is a partial subset
- Queries populate and deduplicate, never get served from
basics
~20 sfind checks the persistence context, then the second-level cache, then the database. A JPQL query with a predicate cannot be answered by either, because both are keyed by identifier and neither can evaluate a WHERE clause — so the SQL runs and only the returned rows are resolved through the caches.
solid answer
~50 s`find` is an identifier lookup, so it can walk a chain: (1) persistence context — hit means the exact managed instance is returned with no SQL; (2) second-level cache if the entity is cacheable — hit means Hibernate hydrates an instance from the cached state and puts it in the context; (3) database SELECT, followed by publishing to both. A JPQL query is not a key lookup. The caches are maps from identifier to state; nothing in them can evaluate `city = ?`, and the cache has no idea whether it holds *all* the matching rows. So Hibernate always sends SQL. The caches still participate afterwards: for each returned row, if that identifier is already managed, the fresh row data is discarded and the existing instance returned; otherwise an instance is hydrated and, if cacheable, published to the shared cache. Two consequences: cached entities do not reduce query count, and a stale managed instance can shadow newer values returned by the query.
code
java · 9 lines// Key lookup: may be answered entirely from cache
Customer c = em.find(Customer.class, 42L);
// Predicate query: always emits SQL
List<Customer> byCity = em.createQuery(
"select c from Customer c where c.city = :city", Customer.class)
.setParameter("city", "Berlin")
.getResultList();
// rows already managed are returned as the existing instancesgo deeper
Recall the order — persistence context, then second-level cache, then database — and that queries with conditions always hit the database.
Explain why: both caches are keyed by identifier and hold only a partial subset, so a predicate cannot be evaluated or trusted; then describe how results are deduplicated against the context.
Draw the consequences: query-dominated workloads gain little from entity caching, stale managed instances shadow query results, and access paths can sometimes be reshaped into identifier lookups.
Treat access-path design as the lever: decide which reads are key lookups worth caching, which must stay fresh, and whether the extra staleness and invalidation cost is justified for the measured hit rate.
## The identifier path `em.find(Customer.class, 42L)` and `getReference`, association navigation and natural-id lookups all boil down to "give me the entity with this key". That is exactly the shape both caches are built for, so Hibernate can try them in order: 1. **Persistence context (first level).** If (Customer, 42) is present, return that instance. No SQL, and importantly no consultation of the shared cache — identity within the unit of work wins over freshness. 2. **Second-level cache**, if the entity is cacheable and the session's cache retrieve mode allows it. On a hit, Hibernate hydrates a new instance from the dehydrated state, registers it in the persistence context, and returns it. No SQL. 3. **Database.** A SELECT by primary key. The result is hydrated, registered in the persistence context, and published to the shared cache if the entity is cacheable. (If the identifier is absent, Hibernate can also record that fact so repeated lookups of a missing key do not re-query, depending on provider behaviour.) ## Why a filtered query cannot use that path A JPQL query with a `where` clause asks a different question: "which entities satisfy this predicate?" Two independent reasons make the caches useless for answering it: - **Wrong key.** Both caches are keyed by identifier. There is no index by `city`, no query planner, no way to evaluate an arbitrary predicate over dehydrated arrays. - **No completeness guarantee.** Even if Hibernate scanned every cached entry, it could not know whether the cache holds *all* customers in that city. A cache is a partial, evictable subset. Answering from a partial set would silently return incomplete results, which is far worse than a slow query. So the SQL is issued, always. What varies is what happens to the rows that come back. ## Post-processing of query results For each row the query returns: - Hibernate extracts the identifier and checks the persistence context. If an instance for that key is already managed, **the freshly read column values are discarded** and the existing instance is returned in the result list. This preserves entity identity, and it is the reason a query can appear not to see a change someone else committed: your session already holds the row. - Otherwise Hibernate hydrates a new instance, registers it, and if the entity is cacheable publishes its state to the second-level cache. So queries *populate* the caches and *deduplicate* against them, but are never *served* from them. ## What this means in practice - Enabling entity caching on a table whose access is dominated by filtered list queries produces close to no reduction in SQL. Teams routinely measure this and conclude "caching doesn't work"; the real conclusion is "my access path is not a key lookup". - Where you can convert an access path to a key lookup, the caches start paying: fetch identifiers with a lean query and then load entities by id, navigate associations, or use natural-id lookup for a business key such as an ISBN or SKU. - A stale managed instance shadowing fresh query rows is a genuine bug source in long transactions. If a code path must see current values, refresh the entity explicitly or use a fresh persistence context. - Caching the *results* of a query is a distinct opt-in mechanism, stored as a list of identifiers with a validity timestamp per table, and it only helps for repeated identical queries with the same parameters. ## How to answer State the three-step chain for `find`, then give both reasons a predicate query cannot use it (keyed by id; no completeness guarantee), then describe what queries do to the caches afterwards (populate and deduplicate, never served from). Adding the practical consequence — that query-heavy workloads see little benefit from entity caching — is what separates a memorised answer from an understood one.
- You loaded Customer 42 early in a long transaction; another transaction then changed its city. A later JPQL query in your session returns that row. Which city do you see?The old one. The query does fetch the current row from the database, but on hydration Hibernate finds identifier 42 already managed in your persistence context and discards the freshly read values to preserve entity identity. Only refresh on that entity, or a new persistence context, shows the committed change.
- How can you restructure a read so that entity caching actually helps?Turn it into an identifier lookup. Run a lean query that returns only the identifiers (or navigate an association, or use a natural-id lookup for a business key), then load each entity by id so the persistence context and second-level cache can serve it. The predicate still costs one query, but the entity loads become cache hits.
saying these in an interview costs you the question
- Believing filtered JPQL queries are answered from the entity cache once caching is enabled
- Thinking Hibernate scans cached entries to evaluate a WHERE clause
- Assuming query results always reflect the latest database values even for already-managed entities
- Claiming the second-level cache is consulted before the persistence context
- Confusing entity caching with caching the results of a query