You are deciding which parts of a domain model should get a shared second-level cache and how much heap to budget for it. How do you make that call, and what would make you leave the second-level cache switched off entirely?
answer
- Prove the DB is the bottleneck first
- Axes: read:write, access path, size, staleness tolerance
- L2 serves id lookups — not filtered queries
- Bulk HQL and outside writers bypass the cache
- No per-region metrics + kill switch → do not adopt
basics
~20 sCache small, hot, rarely-written, id-accessed data whose staleness the domain tolerates. Skip write-heavy, huge, or externally-written tables. Leave L2 off entirely when the database is not the bottleneck, or when a correctness-critical read cannot tolerate any staleness window.
solid answer
~60 sI score each candidate entity on four axes: 1. **Read:write ratio** — caching a write-heavy table means paying region maintenance on every write for a hit rate near zero. 2. **Access path** — L2 serves lookups by id, proxy initialisation and many-to-one navigation. If the traffic is list queries with filters, an entity cache barely helps. 3. **Size** — the hot working set must fit comfortably; a region that evicts constantly is pure overhead. Estimate entry size × residency against heap and GC headroom. 4. **Staleness tolerance** — who else writes this table? Batch jobs, other services, direct SQL, and Hibernate bulk HQL updates all bypass the cache. On N nodes with a local provider, there are N staleness windows. That usually leaves a small set: reference and configuration data, slowly changing catalogues, heavy many-to-one targets. I leave L2 off when the database is not the bottleneck, when correctness demands a read straight from the source, or when the team cannot yet operate and monitor it — an unmeasured cache is a liability, not an optimisation.
code
java · 6 lines// Bulk HQL runs in the database; no entities are loaded, so L2 is not maintained
em.createQuery("update Product p set p.price = p.price * 1.1 where p.categoryId = :c")
.setParameter("c", categoryId)
.executeUpdate();
em.getEntityManagerFactory().getCache().evict(Product.class);go deeper
Say that caching suits small, rarely changing reference data read by id, and that frequently written data is a poor fit.
Structure the answer around read:write ratio, access path, size and staleness, and name bulk HQL as a cache-bypassing write path.
Add measurement: prove the database is the bottleneck first, then judge each region by hit ratio and eviction rate, with a configuration-level kill switch.
Frame it as a distributed-state commitment with operational cost, articulate the classes of data where stale reads are unacceptable, and be willing to conclude that no cache is the right architecture.
## Start from the bottleneck, not from the feature A second-level cache is a distributed-systems commitment sold as a configuration flag. Before enabling it, establish that database access is actually the constraint: high query counts on the same ids, measurable database CPU or latency, connection-pool saturation. If the profile says time goes into serialization, N+1 query patterns, or missing indexes, caching hides the symptom while the underlying defect persists — and an N+1 pattern served from cache is still N+1 round trips through your own code. ## The four axes **Read:write ratio.** Every write to a cached entity must maintain the region: invalidate it, or under a read-write strategy take a soft lock and publish new state. On a table where writes rival reads, that is added latency on the write path with near-zero return. Reference data with essentially no writes sits at the opposite end and is the archetypal good candidate. **Access path.** The entity cache is a map from identifier to dehydrated state. It serves `find()` by id, initialising a lazy proxy, and navigating a `@ManyToOne`. It does **not** serve arbitrary queries; those go to the database unless you also adopt the query cache, which brings its own invalidation coarseness. So an entity read almost exclusively through filtered list queries is a weak candidate even if it is read-heavy. **Size and shape.** Entries are dehydrated state, but wide rows, large text columns and `@Lob` fields make them heavy. The rule of thumb: measure entry size, multiply by the residency you want, compare against heap and GC headroom, and check whether the *hot* subset is small even if the table is not. If nothing but the whole table would give a decent hit ratio and the whole table does not fit, the answer is no. **Staleness tolerance.** This is the axis that ends careers, not benchmarks. Ask who writes the data: - Hibernate in this JVM — the cache is maintained; fine. - Hibernate on another node with a local provider — each node diverges until expiry. N nodes, N windows. - Bulk HQL `update`/`delete` — executes in the database and does not load entities; without explicit region invalidation the cache goes stale. - Batch jobs, other services, DBA scripts, replication — completely invisible to Hibernate. Then ask what a stale read costs. A stale currency name is a shrug. A stale credit limit, entitlement, feature flag or price is an incident. Cache the first class of data; read the second from the database, and if that is too slow, fix the query or the model. ## Budgeting heap Give each region an explicit maximum in the provider's configuration and derive the total from what the JVM can carry alongside its normal working set. Two failure modes bracket the range: regions so small they thrash (constant puts and evictions, no hits) and regions so large they lengthen GC pauses and reduce headroom for request-time allocation. If large regions are genuinely justified, offheap or a dedicated grid moves the bytes out of the collector's path at the cost of serialization per access. Either way, cap it — an unbounded region is a slow-motion out-of-memory failure. ## The monitoring precondition Adopt L2 only with per-region visibility: hit ratio, put count, eviction rate, element count, plus the provider's own metrics. Without those you cannot answer the only question that matters six months later — is this region earning its memory? — and you cannot tell a cache problem from a database problem during an incident. Keep a **kill switch**: `jakarta.persistence.sharedCache.mode=NONE` disables participation by configuration alone, so a suspected staleness bug can be bisected in one deploy without unwinding the provider wiring. ## When to leave it off - The database is not the bottleneck, or the workload is write-dominated. - Reads must be authoritative — money, entitlements, anything with a legal or safety consequence. - The data is written from outside Hibernate and you cannot reliably invalidate. - The team has no per-region monitoring and no operational appetite for another stateful component. - A simpler fix exists: a covering index, killing an N+1 with a fetch join, a read replica, or an explicit application-level cache with a clear invalidation contract owned by the domain rather than the ORM. The honest senior framing is that L2 is a narrow tool with a sharp edge. It is superb for small, hot, stable, id-accessed reference data and a liability everywhere else — and “we cache less than we could” is a defensible position, while “we served a stale entitlement” is not.
- How would you tell, after six months in production, whether a given cache region was worth keeping?Read its per-region statistics: hit ratio, put count, eviction rate and element count. A region with many puts and few hits is paying write and memory cost for nothing; one with a high hit ratio and low eviction is earning its heap. Then test the counterfactual by disabling participation in a canary or with the shared-cache mode set to NONE and comparing database load and latency — if nothing moves, remove the region and reclaim the memory.
- A team proposes caching an entity that a nightly batch job also updates with direct SQL. What do you advise?Hibernate cannot see those writes, so the region will serve pre-batch state until it expires — potentially a whole day. The options are to have the batch (or a hook after it) explicitly evict the region, to set a TTL shorter than the staleness the domain tolerates, or not to cache the entity. I would prefer explicit eviction with a bounded TTL as a safety net, and if neither can be guaranteed, leave it uncached.
Caching is like printing a reference sheet and pinning it above every desk: wonderful for the tax table that changes yearly, disastrous for today's account balances.
saying these in an interview costs you the question
- Treating L2 as a general speed switch to enable everywhere
- Ignoring that bulk HQL updates and external writers bypass the cache
- Assuming a local cache is coherent across nodes
- Sizing regions by guesswork with no eviction or hit metrics
- Reaching for caching before fixing N+1 queries or missing indexes