Hibernate has a query cache that is separate from its entity cache. What does it hold, and what two things must be done before a given JPQL query actually uses it?
answer
- use_query_cache=true + setCacheable per query
- Stores ids (or scalars) + timestamp, not rows
- Useless without entity L2 → N selects on hit
- No global “cache all queries” switch
- Same parameters must repeat to get hits
basics
~20 sIt caches the results of a query keyed by the query text and its parameters. You must set hibernate.cache.use_query_cache=true globally and mark each query cacheable individually — setCacheable(true) or the org.hibernate.cacheable hint. It also requires the second-level cache.
solid answer
~50 sThe entity second-level cache is a map from entity id to state, so it cannot answer “which products cost under 50?”. The **query cache** fills that gap: it stores, per query, the **identifiers** the query returned (or scalar values for projections), keyed by the query string plus its bind parameters. Two steps are needed and both are mandatory: 1. **Globally**: `hibernate.cache.use_query_cache=true`. This also creates the query-results region and the update-timestamps region. 2. **Per query**: `org.hibernate.query.Query#setCacheable(true)`, or the JPA hint `org.hibernate.cacheable` set to true. There is no “cache all queries” mode, deliberately — caching a query is a per-query decision. It also depends on the second-level cache being enabled, because the query cache stores ids and needs the entity region to rebuild rows without extra SELECTs. Enable it for stable, frequently repeated queries over rarely written tables.
code
properties · 3 lineshibernate.cache.use_second_level_cache=true
hibernate.cache.use_query_cache=true
hibernate.generate_statistics=truego deeper
Recall the two mandatory steps and that it caches results of specific queries, keyed by query text and parameters.
Explain that it stores identifiers, why the entity region must also be enabled, and which query shapes actually get hits.
Discuss per-query regions, cache modes, and how to read query-cache hit/miss/put statistics to prove value.
Frame per-query opt-in as deliberate policy and weigh the query cache against an application-level cache with an explicit invalidation contract.
## Why a second cache exists Hibernate's second-level cache is keyed by entity identifier. That makes it excellent at `find(Product.class, 42L)`, at initialising a lazy proxy, and at following a many-to-one — and useless for a query, because “all products in category 7 ordered by name” is not a key it knows. A JPQL or Criteria query therefore goes to the database every time, even when every returned row is already sitting in the entity cache. The **query cache** is the separate mechanism for that case. Conceptually it is a map: *(query text + parameters + paging + a few context bits)* → *(list of identifiers, plus the timestamp at which the query ran)*. ## Turning it on **Step 1, global.** `hibernate.cache.use_query_cache=true`. This is off by default. Setting it creates two internal regions: the results region (`default-query-results-region`) and the update-timestamps region (`default-update-timestamps-region`), the latter being how Hibernate knows when a cached result has been invalidated by a write. **Step 2, per query.** Nothing is cached just because the flag is on. Each query opts in: ```java List<Product> list = em.createQuery("select p from Product p where p.category = :c", Product.class) .setParameter("c", category) .setHint("org.hibernate.cacheable", true) .getResultList(); ``` or, on the Hibernate API, `query.unwrap(org.hibernate.query.Query.class).setCacheable(true)`. There is no blanket “cache every query” switch, and that is intentional: whether a query result is worth caching depends on how often it repeats with identical parameters and how often its tables change. **Implicit step 3.** The query cache is only sensible alongside the entity second-level cache. Because it stores identifiers rather than rows, a cache hit still has to materialise entities. If the entities are cached in L2 they come from memory; if not, Hibernate issues a SELECT per identifier — turning one query into N round trips and making the “optimisation” dramatically slower. ## What is stored, precisely - For queries returning entities: the **identifiers**, not the entity state. State comes from the entity region or the database. - For scalar projections (`select p.name, p.price from Product p`): the actual values, since there is no entity to hydrate. - Alongside either: the **timestamp** of when the query executed, which is what makes invalidation possible. ## Regions and per-query control By default all cacheable query results share one region. `setCacheRegion("catalog-queries")` puts a query's results in a named region so it can be sized, monitored and evicted independently — worth doing for a query whose results are large or whose churn would otherwise evict everything else. Cache behaviour per query can also be steered with `CacheMode` / the JPA `cacheRetrieveMode` and `cacheStoreMode` hints, for example to force a refresh of a cached result. ## When it is worth it Good candidates share three traits: the query repeats **often with the same parameter values**; the underlying tables are written **rarely**; and the entities involved are themselves cached in L2. Reference lookups (“all active countries”, “fee schedule for tier X”, navigation menus) fit well. Poor candidates: queries whose parameters are effectively unique per call (a user id, a timestamp, a search string) — every call writes a new entry that is never read again, so the region churns; and queries over tables that are written continuously, because any write to a queried table drops the cached results for that table. ## Verifying With `hibernate.generate_statistics=true`, `Statistics` exposes query-cache hit, miss and put counts, and per-query counts by query string. The signature of a misapplied query cache is puts climbing in step with executions while hits stay near zero — which means you are paying serialization and memory for nothing.
- Why is there no global setting that makes every query cacheable?Because cacheability is a property of a query's usage pattern, not of the application. A query whose parameters differ on every call — a user id, a search term, a timestamp — would fill the region with entries that are never read again, evicting the useful ones and adding serialization cost to every execution. Forcing an explicit per-query decision keeps that cost where someone has consciously judged it worthwhile.
- You enable the query cache but leave the entity second-level cache disabled. What happens?Hibernate can still cache the returned identifiers, but on a hit it has no cached entity state, so it issues a SELECT per identifier to rebuild the rows. One query that previously cost a single round trip becomes N of them. That is why the entity types a cacheable query returns should themselves be marked cacheable, or the query cache should not be used for them.
The entity cache is a filing cabinet of folders by customer number; the query cache is a sticky note listing which folder numbers matched last time you searched — handy only if the folders are still in the cabinet.
saying these in an interview costs you the question
- Thinking the global property alone caches queries
- Believing the query cache stores full result rows for entity queries
- Expecting it to work without the entity second-level cache
- Confusing it with the HQL-to-SQL query plan cache
- Marking user-specific or timestamp-parameterised queries cacheable