skip to content

An order entity in Hibernate has a second-level-cached @OneToMany collection of 500 line items. One line item is added to that collection. What happens to that collection's cache entry, and what does the answer imply for write-heavy associations?

level: middleimportance: should knowfreq 40%

answer

  1. One entry per owner id, replaced whole
  2. No per-element delta — 501 ids rewritten
  3. Child field change ≠ collection invalidation
  4. Owning-side-only FK change leaves stale list
  5. READ_WRITE soft lock; writes churn the entry

basics

~20 s

The whole entry for that one owner is invalidated and rewritten with the new id list — there is no per-element delta. Other owners' entries are untouched. Frequently modified collections therefore churn their cache entry on every write and rarely pay off.

solid answer

~50 s

Collection cache entries are **all-or-nothing per owner**. Adding, removing, or reordering one element makes Hibernate treat the collection as dirty and schedule a collection action at flush; on commit the entry keyed by `Order#42` is invalidated and replaced with the full new id list. There is no incremental "append id 501" operation, and the region as a whole is not flushed — sibling owners keep their entries. Two practical consequences: - For a **write-heavy** association the entry is destroyed and rebuilt constantly, so you pay cache-write and (in a cluster) invalidation-message cost for a value almost nobody reads before it dies. Collection caching suits read-mostly associations. - Under `READ_WRITE` the entry is soft-locked during the transaction so concurrent readers fall through to the database rather than see a stale list; `NONSTRICT_READ_WRITE` simply invalidates and briefly tolerates staleness. Also note: changing a *child's own fields* does not touch the collection entry at all — the id list is unchanged, and the child's state lives in its own entity region.

code

java · 10 lines
java
public void addLineItem(LineItem item) {
    lineItems.add(item);   // makes the PersistentBag dirty -> cache entry invalidated
    item.setOrder(this);   // sets the FK that is actually written
}

// After a bulk statement, nothing was dirty -> evict by hand:
em.createQuery("update LineItem li set li.order = :o where li.id in :ids")
  .executeUpdate();
sessionFactory.getCache()
              .evictCollectionData("com.acme.Order.lineItems", orderId);

go deeper

for a junior

Say that the whole entry for that one order is thrown away and rebuilt, and that other orders' entries are unaffected.

for a middle

Explain snapshot-based dirty checking on the persistent collection, the entry-per-owner key, and why write-heavy collections are poor cache candidates.

for a senior

Bring in the bidirectional owning-side pitfall and bulk-statement bypass, the soft lock under READ_WRITE, and surgical eviction by owner id.

for a principal

Reason about read/write ratio, cluster invalidation traffic and memory per entry to decide whether the association belongs in the cache at all, versus batch fetch or join fetch on the read path.

## Granularity: one entry per owner, replaced whole A collection region is keyed by the **owning entity's identifier**. `Order#42.lineItems` and `Order#77.lineItems` are separate entries in the same region. When the collection of order 42 is modified, Hibernate acts on **that entry only** — it does not flush the region, and it does not touch other orders. Within the entry, however, there is no partial update. The cached value is a single serialized structure (an array of element ids, plus indexes or map keys where relevant). Hibernate's flush produces one of three collection actions for a dirty collection — *recreate*, *update*, or *remove* — and the cache effect of all of them is the same: the old entry is invalidated and, if the strategy allows, a fresh entry with the complete new id list is put in place at transaction end. Adding one element to a 500-element list rewrites all 501 ids. ## How dirtiness is detected Hibernate wraps mapped collections in *persistent collection* proxies (`PersistentBag`, `PersistentSet`, `PersistentList`, …). An initialized persistent collection keeps a snapshot of its loaded contents; at flush time the current contents are compared against that snapshot, and any structural difference marks the collection dirty. That dirty flag is what triggers both the SQL (insert/delete rows in the join or child table) and the cache invalidation. This gives the single most important caveat of collection caching in a **bidirectional** association. If you only change the owning side — `lineItem.setOrder(otherOrder)` — and never touch either collection in a session that has them loaded, then from Hibernate's point of view no collection was dirty. The foreign key changes in the database, but the cached id lists for the old and the new owner are not invalidated, and they are now **stale**. The discipline is to keep both sides of the association in sync in your domain code (a helper `addLineItem` / `removeLineItem` that sets the parent and updates the list), which is good practice anyway and happens to keep the collection cache honest. Similarly, a bulk statement — a JPQL `update`/`delete` or native SQL that changes foreign keys — bypasses the persistence context entirely, so no collection is ever marked dirty and no cache entry is invalidated. Such statements need explicit eviction (`sessionFactory.getCache().evictCollectionData(...)`). ## What does *not* invalidate the entry Mutating a child's own columns — `lineItem.setPrice(...)` — leaves the id list identical. The collection entry stays valid, and correctly so: the child's state is versioned in the `LineItem` entity region, which *is* invalidated by that update. This separation is the payoff of the id-only layout described by the region design: child updates never have to hunt through collection entries. ## Concurrency strategy shapes the write path - **READ_WRITE** — the cache provider soft-locks the entry for the duration of the transaction. Concurrent readers see the lock, treat it as a miss, and read from the database; after commit the entry is updated. This is the safe default for a mutable cached collection. - **NONSTRICT_READ_WRITE** — the entry is invalidated around the update with no lock, leaving a small window in which a reader can repopulate it with pre-commit data. Acceptable only when a brief stale read is harmless. - **TRANSACTIONAL** — the provider participates in JTA so cache and database commit together; needs a fully transactional provider. - **READ_ONLY** — mutating the collection at all is an error: Hibernate/the provider refuses the write rather than let the entry drift. Use it only for immutable reference associations. ## The write-heavy trade-off Put the costs together for a hot, frequently modified association: 1. Every write invalidates and rewrites a whole id array. 2. In a clustered/replicated cache, every write also emits an invalidation or replication message to every node. 3. Readers arriving during the soft lock go to the database anyway. If writes are frequent relative to reads, the entry rarely survives long enough to serve a read, and you have added latency, memory pressure and network chatter for nothing — often measurably slower than no collection caching. Collection caching earns its keep on **read-mostly, bounded** associations: product-to-attributes, country-to-regions, tenant-to-feature-flags. For hot mutable associations, prefer batch fetching or a join fetch on the read path. ## Operational note Because invalidation is keyed per owner, you can evict surgically: `getCache().evictCollectionData("com.acme.Order.lineItems", 42L)` drops one owner's entry, while the region-wide overload drops all of them. Prefer the surgical form after out-of-band writes; the region-wide form is a blunt instrument that costs you every warm entry.

  • Does updating a line item's price invalidate the cached collection entry of its order?
    No. The collection entry holds only element identifiers, and those are unchanged, so the entry stays valid. The update invalidates the `LineItem` entity region entry instead, which is where the price actually lives. That split is precisely why Hibernate stores ids rather than child state in collection regions.
  • You reassign a child to a different parent by setting only the child's @ManyToOne field. What happens to the two parents' cached collections?
    Neither collection was structurally modified in the session, so neither is dirty and neither cache entry is invalidated — both id lists are now stale, one missing an element and one containing an id that no longer belongs to it. The fix is to maintain both sides of the association in a helper method, or to evict the affected collection entries explicitly.

saying these in an interview costs you the question

  • Claiming a single add invalidates the entire collection region for all owners.
  • Claiming Hibernate appends the new id to the cached list incrementally.
  • Believing that changing a child's fields invalidates the parent's cached collection.
  • Assuming a bulk JPQL update keeps collection regions consistent.
  • Recommending collection caching for a high-churn association because 'caching is always faster'.

context