skip to content

Collection & Natural-Id Caching

The lesser-known cache regions — per-collection caches storing id lists, and the natural-id cache that maps business keys to primary keys. Interviewers use these to probe whether your L2 knowledge goes beyond @Cacheable on an entity.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

questions

5

In Hibernate, you put the org.hibernate.annotations.@Cache annotation on an entity's @OneToMany collection field. What does the resulting second-level cache region actually store, and what else has to be cached for that to save database work?

level: middleimportance: must knowfreq 50%

answer

  1. Region stores ids, not rows
  2. Owner id is the key; role names the region
  3. Collection cached + element uncached = N SELECTs
  4. com.acme.Order.lineItems region name
  5. @ElementCollection caches values, not ids

basics

~20 s

Only the element identifiers — an id list keyed by the owner's id, not the children's column data. Hibernate then resolves each id through the element entity's own cache region, so that entity must be cached too or you still hit the database.

solid answer

~50 s

A collection region entry is keyed by the owning entity's identifier plus the collection role, and its value is just the **identifiers** of the elements (plus index or map key for indexed collections). Child column data is never stored there — it lives in the element entity's own region. So caching a collection is a two-step resolution: hit the collection region to get `[7, 9, 12]`, then load each of those ids. If `LineItem` itself is not annotated with `@Cache`, those loads go to the database **one primary-key SELECT per element** — you have traded a single `where order_id = ?` query for N round trips, which is usually slower than not caching at all. That is why collection caching is only useful together with element-entity caching. The default region name is the fully-qualified entity name plus the field, e.g. `com.acme.Order.lineItems`, separate from the entity region `com.acme.Order`.

code

java · 16 lines
java
@Entity
public class Order {
    @Id Long id;

    @OneToMany(mappedBy = "order")
    @Cache(usage = CacheConcurrencyStrategy.READ_WRITE)
    List<LineItem> lineItems;
}

@Entity
@Cache(usage = CacheConcurrencyStrategy.READ_WRITE)
public class LineItem {
    @Id Long id;
    @ManyToOne Order order;
    BigDecimal price;
}

go deeper

for a junior

Recall that the collection needs its own @Cache annotation and that the region holds identifiers, so the element entity must be cached too.

for a middle

Explain the two-step resolution (collection region gives ids, entity region gives state) and be able to name the region-per-role layout and the default region name.

for a senior

Add the failure mode: a collection-cache hit against a cold entity region turns one query into N key lookups; discuss @BatchSize mitigation and per-collection concurrency-strategy choice.

for a principal

Frame it as cache topology — normalized id lists versus duplicated state, per-region sizing and eviction, and when the association should not be in the cache at all.

## The second-level cache is several regions, not one bucket Hibernate's second-level cache (L2) is scoped to the `SessionFactory` and is physically split into independent *regions*, each with its own key shape: - **Entity regions** — one per entity type, key = the entity identifier, value = a *disassembled* (hydrated) copy of the row's scalar state plus foreign-key values. - **Collection regions** — one per *collection role* (an owning entity type plus a field), key = the owner's identifier, value = the collection's contents in cache form. - **Natural-id regions** — natural key to identifier mappings. - **The query-results region** — only used when the query cache is explicitly enabled. These are configured separately. Annotating `Order` with `@Cache` does nothing at all for `order.getLineItems()`; the collection has to be annotated on its own field. ## What lands in a collection region entry For an association whose elements are entities (`@OneToMany`, `@ManyToMany`), the cached value is **only the identifiers of the elements**. A cached `Order#42.lineItems` is essentially `[7, 9, 12]`. None of the line-item columns are there. Variations by collection type: - A `List` with `@OrderColumn` also stores the index positions. - A `Map` association stores the map keys alongside the element ids. - An `@ElementCollection` of basic types or embeddables has no ids to store, so the region holds the **values themselves** (the strings, the embeddable state). The reason for the id-only layout is normalization of the cache: child state exists in exactly one place — the child's entity region — so an update to a line item does not have to be chased into every collection entry that mentions it. ## The consequence: two-step resolution Navigating `order.getLineItems()` on a cached collection goes: 1. Check the persistence context (L1). If the collection is already initialized in this session, nothing else happens. 2. Miss L1 → read the collection region with key `Order#42`. Hit gives the id list. 3. For each id, resolve the element: L1 → `LineItem` entity region → database `select ... from line_item where id = ?`. Step 3 is where naive setups go wrong. If `LineItem` has no `@Cache`, a collection-cache *hit* produces **N single-row SELECTs** instead of the one `select ... from line_item where order_id = ?` that an uncached collection would have issued. That is a self-inflicted N+1: you added caching and made the workload worse. The rule of thumb is: cache the collection **and** the element entity, or cache neither. (You can soften a partial miss with `@BatchSize` on the element entity, which turns the per-id loads into `where id in (?, ?, ?)` batches, but it is a mitigation, not the fix.) ## Configuration mechanics ```java @OneToMany(mappedBy = "order") @Cache(usage = CacheConcurrencyStrategy.READ_WRITE) private List<LineItem> lineItems; ``` - `@Cache` here is Hibernate's own annotation; JPA's `@Cacheable` applies to entities only, so there is no standard JPA way to mark a collection cacheable. - The `usage` (concurrency strategy) is chosen **per collection**, independently of the owner or the element entity. - The default region name is `<fully-qualified entity>.<field>` — `com.acme.Order.lineItems` — and can be overridden with `@Cache(region = "...")`. Knowing the naming matters for statistics, for eviction calls, and for provider-side sizing config, where you configure each region by name. - With `READ_ONLY` on a collection, mutating it (adding or removing an element) fails at flush time rather than silently going stale — appropriate only for genuinely immutable associations such as reference-data children. ## What collection caching does not do - It does **not** cache query results. A JPQL `select li from LineItem li where li.order.id = :id` never consults the collection region even though it returns the same rows; only association navigation (or `Hibernate.initialize`) uses it. - It does **not** cache the elements' state, so a collection region hit with a cold entity region is not a free read. - It is not automatically enabled by `shared-cache-mode`/`@Cacheable` on the entities; the field annotation is required. ## Sizing intuition Because entries hold ids, memory per entry is roughly `element count × identifier size` plus overhead — cheap for a 5-element association, non-trivial for a parent with 100k children, where the entry becomes a large array that is fully rewritten every time the association changes.

  • What gets stored if the collection is an @ElementCollection of basic values, such as a Set<String> of tags?
    There are no element identifiers, so the region stores the values themselves — the actual strings, or the disassembled state of an embeddable. That makes element collections self-sufficient in the cache: a region hit needs no second lookup and no element entity mapping. It also means the entry grows with the payload size, not just with an id count.
  • What are the default region names for the entity and the collection, and why does that matter operationally?
    The entity region defaults to the fully-qualified class name (`com.acme.Order`) and the collection region to the class name plus the field (`com.acme.Order.lineItems`). Cache providers are configured and sized per region name, and Hibernate's `Statistics` reports hits and misses per region, so the names are what you use to set eviction policy, read hit ratios, and evict programmatically.

The collection region is a library index card listing call numbers, not the books. Useless on its own unless the books are also on the nearby shelf — otherwise you still walk to the archive once per call number.

saying these in an interview costs you the question

  • Saying the collection region stores the child entities themselves.
  • Thinking @Cache or @Cacheable on the parent entity automatically caches its associations.
  • Believing a cached collection makes a JPQL query over the child table hit the cache.
  • Caching the collection but not the element entity, and expecting a speedup rather than an N+1.
  • Assuming JPA's @Cacheable can mark a collection cacheable — it applies to entities only.

context

open as a page

An order entity in Hibernate has a second-level-cached @OneToMany collection of 500 line items. One line item is added to that collection. What happens to that collection's cache entry, and what does the answer imply for write-heavy associations?

level: middleimportance: should knowfreq 40%

basics

~20 s

The whole entry for that one owner is invalidated and rewritten with the new id list — there is no per-element delta. Other owners' entries are untouched. Frequently modified collections therefore churn their cache entry on every write and rarely pay off.

open as a page

A Hibernate entity marks a business key such as a book's ISBN with @NaturalId and code looks rows up with session.bySimpleNaturalId(Book.class).load(isbn). What does adding Hibernate's @NaturalIdCache annotation change, and what SQL runs with and without it?

level: middleimportance: should knowfreq 35%

basics

~20 s

@NaturalIdCache stores the natural-key-to-primary-key mapping in its own second-level region, so the lookup skips the select id where isbn = ? resolution query. The entity state still comes from the entity cache or a load by primary key.

open as a page

A Hibernate application has enabled second-level caching on a parent entity's @OneToMany association, yet loading each parent still emits one SELECT per child row. How would you diagnose and fix that?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Almost always the child entity is not cached: the collection region stores ids, so each id resolves with a primary-key select. Confirm with per-region cache statistics, then add caching to the child entity, or add batch fetching, or stop caching the collection.

open as a page

How do you decide which associations in a Hibernate application deserve a second-level collection cache region and which should be left uncached?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Cache read-mostly, bounded associations that are hot on a navigation path, and only when the element entity is cached too. Skip large or frequently modified collections: each write rewrites the whole id list and, in a cluster, broadcasts invalidation.

open as a page