Hibernate maintains an internal region commonly called the update-timestamps cache. Explain the role it plays in deciding whether a previously cached query result may be returned, and what its granularity means in practice.
answer
- Timestamps region: table → last-modified time
- Result carries its execution timestamp + query spaces
- Any space written later → stale → re-execute
- Invalidation is per table, not per row
- Never size-limit or expire the timestamps region
basics
~20 sIt maps each table (query space) to the timestamp of its last write. A cached query result carries the timestamp it was produced at; if any table it touched was written later, the result is treated as stale and re-executed. Granularity is per table, not per row.
solid answer
~50 sEvery cached query result is stored with the timestamp at which it was produced, plus the set of **query spaces** (tables) it read. Separately, Hibernate keeps an update-timestamps region mapping each table to the time it was last modified. On lookup, Hibernate compares: if any of the result's tables has a last-modified timestamp **later than** the result's timestamp, the entry is stale and the query is re-executed against the database. Hibernate marks the tables as modified at flush time and finalises them after the transaction completes, so results produced from uncommitted state are not served. The critical property is **granularity**: invalidation is per table, not per row or per parameter. One insert into `order_line` invalidates *every* cached query that reads `order_line`, however narrowly parameterised. That is what makes the query cache excellent over rarely written tables and useless over busy ones. Operationally the region must never be size-limited or expired — losing a timestamp entry undermines the staleness check.
code
java · 4 linessession.createNativeQuery("update product set price = price * 1.1 where category_id = :c")
.addSynchronizedEntityClass(Product.class) // registers the query space
.setParameter("c", categoryId)
.executeUpdate();go deeper
Know that Hibernate records when each table was last written and discards cached results older than that.
Explain the timestamp comparison, that query spaces are whole tables, and that any write to a queried table invalidates all its cached results.
Cover flush-time versus post-commit marking, native queries needing declared query spaces, and why the region must be unbounded and never expire.
Position the coarse-invalidation trade-off as the deciding factor for adoption, including cross-node and cross-application writers that the mechanism cannot see.
## The problem it solves Caching a query result raises an obvious question: how do you know it is still true? Row-level tracking would require knowing which rows a query *would* now match, which is exactly the work you are trying to avoid. Hibernate takes a coarse but cheap approach: track modification times **per table** and treat any result older than the last write to a table it read as suspect. ## The mechanism Two pieces of bookkeeping: 1. **Per cached result**: the list of ids or values, plus a **timestamp** of when the query was executed, plus the query spaces (table names) it read. Hibernate derives the spaces from the mapping of the entities involved. 2. **The update-timestamps region** (`default-update-timestamps-region`): a map from table name to the timestamp of its most recent modification. **On write**: when Hibernate flushes an insert, update or delete — including bulk HQL `update`/`delete`, which declare their affected spaces — it marks the affected tables in the timestamps region. It uses a pre-invalidate/invalidate pair around the transaction so that in-flight, uncommitted changes do not let a stale result look fresh. **On read**: build the key, fetch the entry, then check every space in the entry. If any table's last-modified timestamp is later than the entry's timestamp, discard and re-execute. Otherwise return the ids and hydrate. The timestamps also use a clock that is coordinated for the region rather than naive wall-clock comparison, which matters in clustered providers where node clocks drift. ## Why granularity dominates the design Because invalidation is per table: - A single insert into a table wipes the usefulness of **every** cached query over it, even ones whose parameters could not possibly match the new row. - Consequently the value of the query cache is dictated by **write rate on the queried tables**, not by how well the query itself is parameterised. Perfect parameter cardinality plus a busy table equals near-zero hit rate. - Queries joining several tables are invalidated by writes to **any** of them, so multi-table queries are more fragile than single-table ones — a join over a stable reference table and a volatile transaction table is invalidated at the volatile table's rate. This is the single most important thing to say about the query cache in an interview: it is a **rarely-written-tables** feature. ## Operational rules **Never bound or expire the timestamps region.** It is small — one entry per table — and it is a correctness structure, not a performance one. Configure it without size limits or TTL. If an entry is evicted, Hibernate loses the record that a table was recently written and can serve results it should have discarded. Provider configurations that apply a blanket default expiry to all caches are a genuine hazard here; declare this region explicitly. **Watch what bypasses it.** Writes Hibernate does not perform — another service, a batch job, direct SQL, replication — never touch the timestamps region, so cached query results survive them indefinitely. Bulk HQL statements *do* register their spaces, which is why they are safer than native SQL for the same job; a native query can declare its affected spaces explicitly (`addSynchronizedEntityClass` / synchronized query spaces) so that Hibernate invalidates correctly — forgetting that is a classic stale-results bug. **In a cluster.** With a local (non-clustered) provider, each node has its own timestamps region, so node A's write does not invalidate node B's cached query results. A clustered provider such as Infinispan propagates them. Do not assume coherence you have not configured. ## Diagnosing With statistics enabled, a query cache showing high miss counts despite well-chosen parameters usually means the underlying tables are being written constantly — confirm by looking at write rates on those tables rather than at the cache. Conversely, cached results that never seem to refresh after external updates point at writes that never registered in the timestamps region. ## The mental summary The query cache trades precision for cost: keys are extremely fine-grained (query text, every parameter, paging), while invalidation is extremely coarse (whole tables). Adopt it where writes are rare enough that the coarse side never fires; avoid it everywhere else.
- Why must the update-timestamps region never be configured with an eviction policy or a TTL?It is a correctness structure holding one entry per table — the record of when that table was last written. If an entry is evicted or expires, Hibernate no longer knows a recent write happened and can consider an older cached query result still valid, serving stale data. The region is tiny, so there is no reason to bound it; declare it explicitly rather than letting a provider-wide default expiry apply to it.
- Your service caches a query over a table that another application also writes. What are your options?Hibernate never learns about the foreign writes, so the cached results stay valid indefinitely. Options are: stop caching that query; put a TTL on the query results region short enough to bound the staleness the domain tolerates; or have the other writer publish an event that triggers an explicit eviction of the region. Only the last is genuinely correct, and if it cannot be guaranteed, the honest choice is not to cache.
- A cacheable query joins a stable reference table and a busy transactional table. What hit rate do you expect?Close to zero. Invalidation is per query space, and the result declares both tables, so every write to the busy table invalidates the entry regardless of how stable the other table is. Multi-table cacheable queries are invalidated at the rate of their most volatile participant, which is why the query cache suits narrow queries over rarely written tables.
It is like a whiteboard listing when each filing cabinet was last touched: your photocopied search results are trusted only if every cabinet they came from was last touched before the copy was made.
saying these in an interview costs you the question
- Thinking invalidation is per row or per cached parameter set
- Applying a global TTL or eviction to the timestamps region
- Assuming external or native SQL writes invalidate cached results automatically
- Expecting a local provider's timestamps to propagate across nodes
- Blaming the entity cache when cached query results are missing constantly