skip to content

A Hibernate application has enabled second-level caching on a parent entity's @OneToMany association, yet loading each parent still emits one SELECT per child row. How would you diagnose and fix that?

level: seniorimportance: should knowfreq 35%

answer

  1. Collection region = ids only; children still need a region
  2. Per-region statistics: hits vs puts vs misses
  3. ENABLE_SELECTIVE = uncached unless annotated
  4. Region name typo = unconfigured region
  5. @BatchSize turns N selects into id IN (...)

basics

~20 s

Almost always the child entity is not cached: the collection region stores ids, so each id resolves with a primary-key select. Confirm with per-region cache statistics, then add caching to the child entity, or add batch fetching, or stop caching the collection.

solid answer

~50 s

Start by proving where the miss is. Turn on Hibernate `Statistics` and read the **per-region** hit/miss/put counters: a healthy collection region with a cold or absent element-entity region is the signature of this symptom, because the collection region stores only element identifiers. Checklist: 1. Is the element entity annotated for caching at all? With `shared-cache-mode=ENABLE_SELECTIVE` (the usual default) an entity without `@Cacheable`/`@Cache` is simply not cached. 2. Are the region names you configured in the provider the ones Hibernate actually uses (`com.acme.Order.lineItems`, `com.acme.LineItem`)? A typo silently creates an unconfigured, tiny or unwritten region. 3. Is the element region evicting almost immediately — too small a max-entry count or too short a TTL? 4. Are the children being written constantly, so entries are invalidated before they are read? Fixes, in order: cache the element entity; add `@BatchSize` so residual misses load as `id in (...)`; or drop the collection cache and use a join fetch or batch fetch on the read path.

code

java · 7 lines
java
Statistics stats = sessionFactory.getStatistics();
for (String region : stats.getSecondLevelCacheRegionNames()) {
    CacheRegionStatistics s = stats.getCacheRegionStatistics(region);
    System.out.printf("%s hits=%d misses=%d puts=%d elements=%d%n",
        region, s.getHitCount(), s.getMissCount(),
        s.getPutCount(), s.getElementCountInMemory());
}

go deeper

for a junior

Recall the root cause — the collection region holds ids, so the child entity must also be cached — and name the annotation that fixes it.

for a middle

Add the measurement step (per-region statistics plus SQL counts) and the shared-cache-mode/region-name checks.

for a senior

Run the full checklist including region sizing, write churn and read-path shape, and pick between element caching, @BatchSize and removing the collection cache, then verify with before/after numbers under a realistic read/write mix.

for a principal

Frame it as whether this data belongs in a distributed cache at all: cost of two coordinated regions and invalidation traffic versus a simpler fetch plan or a purpose-built read model.

## The symptom and its usual cause One SELECT per child, on a collection that is supposed to be cached, is the classic outcome of caching *only* the collection. A collection region entry contains the **element identifiers**, not the elements' state. On a collection-cache hit Hibernate has the id list and must materialize each element: persistence context → the element entity's cache region → database `select ... from child where id = ?`. When the element entity is not in the second-level cache, that last branch fires once per id. Note what this replaced: an *uncached* collection would have loaded with a single statement, `select ... from child where fk = ?`. So the intervention made things worse — the diagnosis has to be honest about that, and "remove the collection cache" is a legitimate fix. ## Step 1: measure, do not guess Enable `hibernate.generate_statistics` and read `SessionFactory#getStatistics`, specifically the **per-region** views: second-level cache hit count, miss count, put count, and element count for each region name. Interpret: - Collection region: high hits, element region: high misses or zero puts → the elements are not cached. This is the case in the question. - Collection region: high misses too → the collection is not being cached either (wrong annotation, wrong region config, or the read path is a query rather than navigation). - Element region: high puts, high misses, low hits → thrashing; entries are evicted or invalidated before they are read. Pair that with SQL logging (`show_sql`, or better, a datasource proxy that counts statements per request) so you can tie the statement pattern to the region counters. A slow-query dashboard alone will not show this — the individual key lookups are each fast; it is their *count* that hurts. ## Step 2: work down the checklist **Is the element entity cacheable at all?** Under `jakarta.persistence.sharedCache.mode=ENABLE_SELECTIVE` — the effective default in most setups — nothing is cached unless annotated. Adding `@Cache(usage = ...)` (or `@Cacheable`) to the child class is the direct fix. **Do the region names match the provider config?** Hibernate's defaults are the fully-qualified entity name for an entity region and entity-name-plus-field for a collection region. Provider configuration (max entries, TTL, on-heap vs off-heap) is keyed by those exact names; a mismatch leaves the region on some tiny default. Print the region names from `Statistics#getSecondLevelCacheRegionNames()` rather than trusting the config file. **Is the region big enough and long-lived enough?** A region sized for 1,000 entries in front of 200,000 hot children evicts constantly. The statistics' put-versus-hit ratio exposes this immediately. **Is the association write-heavy?** Every structural change to a collection invalidates and rewrites that owner's entry, and every child update invalidates the child's entity-region entry. Under heavy writes the entries never live long enough to serve a read. **Is the read path actually navigation?** The collection cache is only consulted when the association is initialized through the owner. A JPQL query that selects children directly never touches it, so if the code was "optimized" into a repository-style query, the cache is bypassed by construction. **Is the transaction/session setup sane?** Second-level reads happen through a session; code that opens a session per child, or that reads outside a transaction with a provider requiring one, can defeat the strategy. ## Step 3: choose a fix deliberately 1. **Cache the element entity as well.** The minimal change that makes the existing design coherent — the id list resolves into the entity region, and the collection navigation becomes SQL-free. Requires the children to be read-mostly and bounded in count. 2. **Add `@BatchSize` to the element entity or the collection.** Residual misses then load as `where id in (?, ?, …)` instead of one statement each — an order-of-magnitude reduction in round trips even with a partially cold cache. Cheap insurance to keep even when entity caching is on. 3. **Drop the collection cache and fetch properly on the read path.** A `join fetch` in the query that loads the parents, or `@BatchSize` on the collection so N parents' collections load in a few `where fk in (…)` statements, is often faster and far simpler than two coordinated cache regions — especially for large or churning associations. 4. **Re-scope what is cached.** If only a handful of parents are hot, the fix may be to cache their children and leave the long tail alone, which the provider's size limit does naturally once the region is sized correctly. ## Verifying the fix Re-run the same load with statistics on and compare: statement count per request, collection-region hits, element-region hits versus puts. A correct fix shows the element region's hit ratio climbing and the per-child statement count collapsing. Then check the write path did not regress — measure invalidation churn under a realistic write mix, not just a read benchmark, because the cost of collection caching is paid by writers.

  • Statistics show high hit counts on the collection region and near-zero puts on the child entity region. What does that tell you?
    The collection is being cached and read successfully, but the child entity is never being written into the second-level cache — it is not marked cacheable, or its region name does not match anything the provider configured. Every id in the cached list therefore resolves with a primary-key select, which is exactly the observed statement pattern.
  • Would removing the collection cache entirely be an acceptable fix here?
    Yes, and sometimes the best one. Without the collection cache Hibernate loads the association with a single foreign-key query per owner, which beats N primary-key lookups. If the association is large or frequently modified, dropping the cache and relying on join fetch or @BatchSize gives better latency with far less cache machinery to keep consistent.

saying these in an interview costs you the question

  • Blaming the cache provider before checking whether the child entity is cacheable at all.
  • Assuming the collection region stores child rows and so the misses must be a provider bug.
  • Turning on the query cache as a fix — it is a different region and does not serve association navigation.
  • Concluding from a read-only benchmark that entity caching solved it, without measuring write-side invalidation churn.
  • Adding EAGER fetching as the fix, trading N selects for cartesian products and unconditional loading.

context