How do you decide whether a given entity belongs in Hibernate's second-level cache, and in what situations does enabling that cache make a system worse rather than faster?
answer
- Read/write ratio, cross-session reuse, footprint, staleness budget
- Entity cache is id-keyed: query screens see no hits
- Write-heavy = pure overhead plus cluster chatter
- Balances, stock, prices, permissions: do not cache
- Default off, enable on measured hit ratio, keep a TTL
basics
~20 sCache entities that are read far more than written, small, reused across sessions, and tolerant of brief staleness. It hurts for write-heavy or huge datasets, for data that must be exactly fresh, and for workloads that read by query rather than by identifier, where the cache adds cost and no hits.
solid answer
~50 sDecide per entity from four numbers: read-to-write ratio, reuse across sessions, entry size times cardinality, and tolerable staleness. Good candidates are reference data, configuration, catalogue rows, and anything looked up by identifier or natural id many times per second. It actively hurts when: writes are frequent, so every commit pays invalidation plus a cluster message and the entry rarely survives to be read; the dataset is large, so it evicts itself and inflates heap and garbage-collection pauses; correctness needs the latest value, such as balances, stock levels or authorization decisions; or the access path is JPQL over columns, since the entity cache is keyed by identifier and those queries still hit the database. There is also an engineering cost: an extra source of truth to reason about during incidents, plus stampedes after a broad eviction, deploy, or restart. My default is off, then enable per entity with evidence — measured hit ratio and a database-load delta — and remove it when the ratio is low.
go deeper
Say that caching suits data read often and changed rarely, such as reference tables, and is a poor fit for data that changes constantly.
Add the mechanics behind that rule: invalidation cost per write, identifier-keyed lookups, memory footprint, and the staleness window.
Diagnose with numbers — hit ratio, puts versus misses, database load delta — and describe stampede and restart behaviour plus which reads must bypass the cache.
Present a decision procedure and the alternatives you would try first (access-pattern fixes, indexes, an explicit projection cache), then define policy: default off, evidence to enable, review to keep.
## The decision inputs For each entity, four properties decide the answer: 1. **Read-to-write ratio.** Every write invalidates, and in a cluster also sends a message. Below roughly ten reads per write the bookkeeping starts eating the benefit; near parity it is pure overhead. 2. **Reuse across sessions.** The shared cache only helps when *different* sessions ask for the *same* identifiers. Repeated access inside one request is already served by the persistence context for free. Data with a long tail of rarely repeated keys (per-user events, log rows, order history) gets a near-zero hit ratio. 3. **Footprint.** Entry size multiplied by cardinality must fit comfortably in the region. A large region competes with the application for heap, lengthens garbage-collection pauses, and, if undersized, thrashes: entries are evicted before they are re-read, so you pay to populate and never collect. 4. **Staleness tolerance.** Shared caching is eventually consistent. Ask what happens if a read is a few hundred milliseconds, or in a partition several seconds, out of date. For a product description, nothing. For an account balance, a stock level at checkout, a price at payment, a feature flag guarding a kill switch or a permission check, that is a defect. ## Where it actively hurts - **Write-heavy entities.** Cost on every commit, benefit almost never realised. In an invalidation cluster you have added a network message to every write in exchange for nothing. - **Query-shaped access.** The entity cache is keyed by identifier. A screen whose reads are all `select p from Product p where ...` still goes to the database; the entity cache only helps if the identifier path is taken (find/getReference, association navigation, natural-id lookup) or if results are re-hydrated through cached ids. Teams frequently enable caching and measure zero improvement for exactly this reason. - **Correctness-sensitive reads.** Once a value can be stale, every consumer must be audited for whether that is acceptable. That audit is rarely done, which turns the cache into a latent bug. - **Debuggability.** "It works on one node", "it fixed itself after a restart" and irreproducible support tickets are the standard symptoms. Incident response now has a second source of truth to check. - **Stampedes.** A broad eviction, a deploy, a rolling restart, or a bulk DML over a hot region sends the full read load at the database at once. Systems sized around the cache hit ratio have no headroom for that, so the cache silently becomes a capacity dependency. ## Cheaper alternatives to consider first Often the real problem is not "the database is too slow" but an access pattern: N+1 selects, missing indexes, over-fetching, or a read that could be a join. Fixing those is usually a larger win with none of the consistency cost. Above the ORM, an explicit application-level cache of a small immutable projection (loaded once, refreshed on a schedule) is easier to reason about than entity caching, because its staleness is deliberate and visible. Below the ORM, the database's own buffer cache already keeps hot pages in memory, so the gain from the second-level cache is the ORM round trip and hydration, not the disk read people imagine. ## How to operate it Default off. Turn it on entity by entity with a hypothesis and a measurement: hit ratio, put/miss counts, and database load before and after. Set a time-to-live even when invalidation is reliable, so lost messages self-heal. Keep the regions per entity so you can evict narrowly. Re-check the hit ratio periodically — workloads drift, and a region with a low hit ratio is pure cost that should be removed. ## What a strong answer sounds like Not a list of settings but a decision procedure: the four inputs, the classes of entity that fail them, the observation that the entity cache is identifier-keyed so query-heavy screens see no benefit, the operational risks (stampede, staleness audit, incident complexity), and the discipline of enabling it with evidence and removing it when the evidence disappears.
- A team enabled second-level caching on their main entity and saw no improvement. What is the most likely explanation?Their read path is JPQL filtered by columns, not lookups by identifier. The entity cache is keyed by primary key, so those queries still execute against the database. Either the access pattern needs to change to identifier or natural-id lookups, or the improvement must come from somewhere else entirely, such as indexing or fixing N+1 fetching.
- What single metric would make you remove caching from an entity?A sustained low hit ratio, meaning the region is mostly puts and misses. That says the data is either written too often or rarely re-read across sessions, so you are paying memory, invalidation and cluster traffic for nothing. I would also remove it if a staleness-related defect appeared for data nobody had classified as staleness-tolerant.
saying these in an interview costs you the question
- Treating the second-level cache as a general speed switch to enable globally
- Assuming cached entities also make JPQL queries faster
- Caching mutable, correctness-critical data such as balances or stock levels
- Ignoring heap and garbage-collection cost of a large region
- Having no measurement of hit ratio, so nobody can tell whether the cache pays for itself