skip to content

Hibernate's second-level cache stores entities in a dehydrated form rather than storing the entity objects themselves. What exactly is stored, why was it designed that way, and what does it cost?

level: seniorimportance: should knowfreq 30%

answer

  1. Disassembled array of column values keyed by id
  2. Associations stored as identifiers, not references
  3. Collections cached separately, separate opt-in
  4. Sharing objects would leak uncommitted state
  5. Cost: hydration per hit, fan-out, serialisable and layout-stable

basics

~20 s

It stores a flat array of column values keyed by identifier, with to-one associations reduced to identifiers and collections held in separate regions. Sharing live objects would leak uncommitted changes between sessions, so each session hydrates its own instance — at the cost of rebuilding it on every hit.

solid answer

~50 s

A cached entry is a disassembled snapshot: the scalar property values in a fixed order, the version if the entity has one, to-one associations stored as foreign-key identifiers rather than object references, and collections cached separately in their own regions keyed by the owner's identifier. The reason is isolation and safety. If sessions shared one instance, an uncommitted edit in one unit of work would be visible in another, entity identity per persistence context would break, and concurrent mutation of a shared mutable object would be a data race. Storing state instead means every session hydrates a private instance, and the entry is a plain serialisable value that can live off-heap or in a distributed grid. The costs are hydration work on every hit, association identifiers that may cause further loads unless those targets are cached too, and the requirement that cached types be serialisable and stable across deployments.

code

java · 11 lines
java
@Entity
@Cacheable
@Cache(usage = CacheConcurrencyStrategy.READ_WRITE)
public class Order {
    @Id private Long id;
    private Long total;

    @OneToMany(mappedBy = "order")
    @Cache(usage = CacheConcurrencyStrategy.READ_WRITE) // separate region, separate opt-in
    private List<OrderLine> lines;
}

go deeper

for a junior

Know that the shared cache holds data, not objects, and that each session builds its own entity instance from it.

for a middle

Describe the contents — scalar values, version, association identifiers, collections separate — and give isolation as the reason objects are not shared.

for a senior

Add the costs you would actually meet: hydration per hit, fan-out through uncached association targets, collections needing their own opt-in, and cache layout breaking on mapping changes.

for a principal

Use it to set policy: which entity clusters are cached together so navigation stays in cache, how regions are versioned across deploys, and what the hydration and memory profile means for capacity.

## What a cached entry actually contains For an entity marked cacheable, Hibernate stores under the key (entity name, identifier) a structure that is essentially the row as the ORM sees it: - the scalar property values, in the fixed order of the entity's persistent attributes; - the version value, when the entity is versioned; - to-one association targets as **identifiers**, not references — so a cached `Order` holds `customerId = 42`, not a `Customer` object; - nothing for collections: an owned collection is cached, if at all, in a separate collection region keyed by the owner identifier, and it too stores element identifiers rather than objects. This is why the format is called dehydrated or disassembled. Hydration is the reverse: given the array, Hibernate instantiates the entity class, sets the scalar values, and installs proxies or lazy handles for the association identifiers. ## Why not cache the objects Three hard reasons. **Isolation.** A managed entity is mutable and belongs to exactly one persistence context. If two sessions shared one instance, a change made in session A before it commits — or that it later rolls back — would be immediately visible in session B. There would be no way to give each unit of work a consistent view, and the cache would silently publish dirty state. **Identity.** JPA guarantees one instance per identifier *per persistence context*. A shared object would either violate that or force every context to wrap it, and dirty checking needs a per-session loaded snapshot to diff against, which a shared object cannot provide. **Portability and lifetime.** A flat array of values is serialisable and can be stored off-heap, on disk, or in a remote grid, and it survives across sessions without keeping object graphs (and therefore whole class-loader-bound structures) alive. An object graph would pin arbitrary reachable state in memory and could not cross a network. ## What it costs - **Hydration per hit.** Every cache hit builds a new instance and populates it. This is real CPU and allocation — far cheaper than a database round trip, but not free, and it means the shared cache is not a magic zero-cost read. - **Association fan-out.** Because associations are identifiers, resolving them may trigger further lookups. If the target entity is itself cached, those are cache hits; if not, they are database reads, and a cached parent with uncached children can turn one logical read into many queries. - **Collections need their own opt-in.** Caching an entity does not cache its collections. A cached `Order` whose `lines` collection is not cached still queries for the lines every time. - **Serialisation and shape stability.** For off-heap or distributed storage, cached types must be serialisable, and the stored layout is tied to the entity's attribute list. Change the mapping and deploy while old entries survive, and you get deserialization failures or misaligned values — which is why regions are usually versioned or cleared as part of a deployment. - **Mutable value types.** Embeddables, arrays, dates and similar mutable values are copied in and out; treating them as shared references is a bug source, which is another reason the cache deals in copied state. ## Practical implications When someone reports "the second-level cache is enabled but performance barely changed", the dehydration model explains several of the causes: hydration cost on hits, uncached collections still hitting the database, and association identifiers dragging in uncached targets. It also explains a pleasant property: two sessions reading the same cached entity cannot interfere with each other, so caching does not weaken isolation *within* a transaction — the staleness risk comes from the entry being old, not from sharing. ## How to answer Name the contents (scalar values, version, association identifiers, collections separate), give isolation as the primary design reason with identity and portability as supporting ones, and finish with the costs: hydration per hit, fan-out through association identifiers, separate opt-in for collections, and serialisation and layout stability across deployments.

  • A cached Order has an uncached Customer association. What does a cache hit on the Order cost?
    The Order itself is served from the cache, but the association is stored as a customer identifier, so touching it triggers a load of the Customer. Since that entity is not cached, the load goes to the database. One logical read becomes a cache hit plus a query, which is why parents and their commonly navigated targets are usually cached together.
  • Why can a deployment that changes an entity's mapping break a populated second-level cache?
    Entries are stored as a positional array matching the entity's persistent attributes, and for off-heap or distributed providers they are serialised. Adding, removing or reordering attributes makes existing entries misaligned or undeserialisable. Standard practice is to clear or version the cache regions as part of the deploy so old entries are never hydrated by new code.

It is the difference between lending someone your assembled model kit and handing them a boxed set of parts: the parts can be copied, stored on a shelf, and posted to another office, and everybody who receives a box builds their own model that nobody else can knock over.

saying these in an interview costs you the question

  • Saying the second-level cache stores entity objects that sessions share
  • Assuming caching an entity automatically caches its collections
  • Forgetting that association targets are stored as identifiers and may still cause queries
  • Treating a cache hit as free, with no hydration cost
  • Ignoring that cached entry layout is tied to the mapping and breaks across mapping changes

context