skip to content

Someone proposes turning on Hibernate's second-level cache for every entity to fix slow reads, with no size or expiry configured per region. How do you evaluate that proposal, and what would you actually cache?

level: principalimportance: nice to knowfreq 30%

answer

  1. L2 = by-id entity cache, not a query cache
  2. cached N+1 is still N lookups
  3. unbounded region → table-sized memory
  4. read-only / read-write / nonstrict strategies
  5. reference data + size cap + TTL + hit ratio

basics

~20 s

Caching makes existing lookups cheaper; it does not remove queries a bad fetch plan issues. Cache only small, read-mostly, by-id-accessed reference data, with an explicit size cap and expiry per region, and measure the hit ratio. Fix fetch plans first.

solid answer

~60 s

I push back on three grounds. **It targets the wrong cost.** The second-level cache is a by-id entity cache. A screen firing 200 lazy selects still fires 200 lookups — cheaper ones. The fix for fan-out is the fetch plan, not a cache. And most slow list screens are driven by JPQL, which the entity cache does not serve at all; only the separate, far more fragile query cache does, and it is invalidated by any write to the tables involved. **Unbounded regions are a memory leak with a good reputation.** No maximum size and no TTL means the region grows to the size of the table. On heap that is GC pressure then OutOfMemoryError; off-heap or distributed it is cost and network. **Write-heavy entities get worse, not better.** Every update invalidates or re-writes the entry, and the concurrency strategy matters: `read-only` for immutable data, `read-write` for mutable data with soft locks, `nonstrict-read-write` only where a stale read is genuinely acceptable. So: cache small, read-mostly, high-reuse, id-addressed reference data; bound every region with a size and expiry; never cache anything mutated by external writers that bypass Hibernate; and gate the decision on a measured hit ratio.

code

java · 9 lines
java
@Entity
@Cacheable
@Cache(usage = CacheConcurrencyStrategy.READ_ONLY, region = "reference.country")
@Immutable
class Country {
    @Id private String iso2;
    private String name;
}
// region "reference.country": max 300 entries, TTL 24h — set in the cache provider config

go deeper

for a junior

Know that the second-level cache stores entities by id across sessions and that it is not a fix for queries returning too much data.

for a middle

Distinguish the entity cache from the query cache, name the concurrency strategies, and explain why regions need a size and expiry.

for a senior

Reason about invalidation cost on write-heavy entities, staleness from external writers, clustering effects, and validate with hit-ratio statistics before and after.

for a principal

Own the sequencing and the operational contract: fix fetch plans first, cache only a justified short list of reference data, bound and name every region, define the acceptable staleness per entity, and retire regions that fail their hit-ratio budget.

## What the second-level cache actually is The second-level cache is a `SessionFactory`-scoped store of entity **state keyed by identifier** (plus optional collection and query caches). On a cache hit for `em.find(Entity.class, id)`, Hibernate rebuilds the entity from cached state without touching the database. Three consequences follow immediately, and they decide the whole discussion: 1. **It serves lookups by id.** It does not serve `select ... where status = ?`. A JPQL query always goes to the database unless the separate **query cache** is enabled. 2. **The query cache is a different thing with different economics.** It stores the *identifiers* a query returned, keyed by the query text and parameters, and it is invalidated whenever any table in the query's space is written. On a write-active table it is invalidated constantly, so hit ratios collapse and you pay maintenance for nothing. It also still needs the entity cache to resolve those ids, or it issues n lookups. 3. **Collection caching stores ids too**, so a cached collection still resolves each element — usefully cheap only when those elements are themselves cached. ## Why "cache everything" is the wrong instinct **It does not remove work; it makes work cheaper.** A 1 + N fan-out with a warm cache is still N lookups: N map hits, N entity rehydrations, N objects. Faster, yes; but the correct fix — a fetch join, an entity graph, batch fetching, a projection — makes N *zero*. Caching a bad fetch plan freezes it in place and makes the eventual fix harder to justify. **Unbounded regions are unbounded.** Without a maximum entry count or size and without a TTL/TTI, a region converges on the whole table. On-heap that is direct GC pressure and eventually an OutOfMemoryError; the failure looks like a memory leak because it is one. Off-heap or distributed it becomes a cost and latency problem instead. **Writes invert the economics.** Every update must invalidate or update the entry, and under `read-write` semantics Hibernate takes a soft lock around the change. For a hot, frequently-updated entity the cache adds coordination on the write path and delivers few hits on the read path. **Staleness has a blast radius.** Anything that changes rows without going through this `SessionFactory` — another service, a migration, a DBA fix, a bulk `update` HQL statement — leaves the cache serving old data with no signal. That is a correctness risk, not a performance one, and it must be reasoned about per entity. **Clustering multiplies the questions.** In more than one node, entries are replicated or distributed, and invalidation crosses the network. A cache that is a clear win on a single node can be a net loss across six. ## What genuinely belongs in it The profile of a good candidate: - **Small and bounded** — currencies, countries, tax rates, product categories, feature definitions, permission types. - **Read-mostly or immutable** — changes are rare and deployment-shaped rather than user-shaped. Immutable entities can use the `read-only` strategy, which is the cheapest and safest. - **Accessed by identifier** — because that is what the cache serves. Association navigation to a to-one is exactly this pattern, which is why reference data behind `@ManyToOne` is the sweet spot. - **High reuse across sessions** — many requests want the same few rows. - **Tolerant of bounded staleness**, with an explicit expiry expressing what "bounded" means. What does not belong: user-generated content, orders, events, anything append-heavy or hot-updated, and anything large enough that caching it is really "keep the table in RAM" — a decision to make deliberately, if at all. ## Bounding every region Whatever provider is in use, each region gets an explicit maximum size (entries or bytes) and an explicit expiry (time-to-live, and time-to-idle if reuse is bursty). Two reasons: memory becomes a number you chose rather than a consequence of table growth, and expiry caps the staleness window for anything modified outside Hibernate. Regions should be named per entity or per group so they can be sized differently and observed separately. ## Deciding with evidence - Establish the **hit ratio** per region from Hibernate's statistics. A region below roughly 80–90% hits is usually not paying for its memory and invalidation cost; a region at 99% on a tiny table is doing exactly its job. - Watch **eviction/expiry rates**. High eviction means the region is too small or the data is not reusable. - Watch **puts versus hits**. A region with many puts and few hits is caching write-through traffic, which is pure overhead. - Remove regions that fail these tests. An unused cache is not free. ## The order of operations Fix the fetch plan, bound the queries, project what the screen needs, batch the writes. Re-measure. Then, for whatever remains and fits the reference-data profile, add a **named, bounded, expiring** region and prove the hit ratio. Caching last is not conservatism — it is that caching only makes remaining work cheaper, and everything above removes work outright. ## How to answer the proposer Agree on the goal, disagree on the instrument: ask for the statement count and row count behind the slow screens first. In my experience most of the wins are in the fetch plan, and the caching that survives is a short list of reference tables — which is a much easier thing to operate than a cache on every entity.

  • Why does enabling the second-level cache rarely speed up a slow list screen?
    Because list screens are driven by JPQL or Criteria queries, and the entity cache only answers lookups by identifier — the query itself still goes to the database. The query cache could serve it, but it stores only the returned identifiers and is invalidated by any write to the tables involved, so on an active table its hit ratio is poor while its maintenance cost is not. The real lever for a slow list is the fetch plan and the amount of data returned.
  • An entity is also updated by a separate reporting job writing directly with SQL. What does that mean for caching it?
    It means Hibernate has no way to know the row changed, so cached entries can serve stale data indefinitely. Either exclude that entity from the cache, or accept a bounded staleness window enforced by an explicit TTL, or arrange for the external writer to publish invalidations. The decision has to be made per entity against how wrong a stale read is allowed to be — this is a correctness question, not a tuning one.

Caching a bad fetch plan is like buying a faster van for a courier who makes 200 separate trips: each trip gets quicker, but the fix was always to load the parcels into one trip.

saying these in an interview costs you the question

  • Expecting the second-level cache to fix an N+1 or a slow JPQL list query
  • Leaving regions with no maximum size or expiry and calling it 'letting it warm up'
  • Caching hot, write-heavy entities and ignoring invalidation churn
  • Ignoring writers outside Hibernate when reasoning about staleness
  • Enabling caching without ever measuring the hit ratio, then keeping it because 'it can only help'

context