skip to content

Compare Hibernate's first-level and second-level caches: what is each one scoped to, who can see its contents, when is it active, and what is stored in it?

level: middleimportance: must knowfreq 62%

answer

  1. L1 = persistence context, per session, always on, live objects
  2. L2 = SessionFactory-scoped, opt-in, shared, dehydrated state
  3. L1 for identity and dirty checking; L2 for round-trip avoidance
  4. Lookup: context, then shared cache, then database
  5. Both keyed by id, so predicates still hit the database

basics

~20 s

The first-level cache is the persistence context: per session, always on, holds live entity objects, dies with the session, private to one unit of work. The second-level cache belongs to the SessionFactory: opt-in, shared by all sessions in the JVM, and holds entity state, not objects.

solid answer

~50 s

**First level** = the persistence context of one `EntityManager`/`Session`. Always on, cannot be disabled, invisible to any other session, discarded on close. It stores the managed instances plus the loaded snapshot used for dirty checking, and gives entity identity inside the unit of work. **Second level** = a cache owned by the `SessionFactory`/`EntityManagerFactory`, so it outlives sessions and is shared by every session in that JVM. It is opt-in per entity (`@Cache`/`@Cacheable` plus enabling it in configuration and choosing a provider). It does not store entity instances; it stores a dehydrated form — an array of column values keyed by identifier, with associations kept as identifiers — so each session builds its own instance from it, keeping sessions isolated. Consequences: L1 helps only within one request; L2 helps across requests and users but introduces staleness, invalidation and, in a cluster, cross-node concerns. A `find` checks L1, then L2, then the database.

code

java · 9 lines
java
@Entity
@Cacheable
@org.hibernate.annotations.Cache(
    usage = CacheConcurrencyStrategy.READ_WRITE,
    region = "customers")
public class Customer {
    @Id private Long id;
    private String name;
}

go deeper

for a junior

Get the scopes right: first level per session and automatic, second level per factory, shared and opt-in.

for a middle

Add stored form (instances plus snapshot versus dehydrated state), lookup order, and why each exists (identity versus round-trip avoidance).

for a senior

Discuss isolation as the reason for dehydration, staleness and invalidation as the price of the shared cache, and why enabling it does not speed up filtered queries.

for a principal

Frame the choice: the first-level cache is part of the unit-of-work contract, while the second-level cache is a capacity decision with a consistency cost that must be justified per entity and per topology.

## Two caches, two purposes They are frequently taught as "levels" of one thing, which hides how different they are. The first-level cache exists for **correctness**: identity and dirty checking within a unit of work. The second-level cache exists for **performance**: avoiding round trips across units of work. Nothing about the first is optional; almost everything about the second is a choice. ## Scope and lifetime | | First level | Second level | |---|---|---| | Owner | one `EntityManager`/`Session` | the `EntityManagerFactory`/`SessionFactory` | | Lifetime | until `close()` (typically one request or transaction) | process lifetime, or provider-controlled expiry | | Visibility | private to that session | every session in the JVM, and across cluster nodes if the provider is clustered | | Enabled | always, cannot be switched off | opt-in: configuration plus a provider plus per-entity annotation | | Contents | managed instances plus loaded snapshots | dehydrated state keyed by identifier | | Concurrency | single-threaded by contract | concurrent, needs a concurrency strategy | The visibility line is the one candidates most often get wrong: two simultaneous requests never share first-level state, so the first-level cache does nothing for throughput across users. ## What "dehydrated state" means and why it matters The second-level cache does not hold your `Customer` object. It holds a disassembled representation: the scalar column values in a fixed order, with to-one associations reduced to their identifiers and collections cached separately in their own regions. When a session takes a hit, Hibernate hydrates a new entity instance from that array and attaches it to that session's persistence context. This is deliberate. If instances were shared, two sessions would hold the same mutable object, and one session's uncommitted edit would be visible to another, destroying isolation. Sharing state instead means each session gets its own instance, changes stay local until commit, and the cached data can be serialised for an off-heap or distributed store. The price is hydration work on every hit — real, but far cheaper than a database round trip. ## Lookup order On `em.find(Customer.class, 42L)`: 1. Persistence context. If present, return that exact instance — no SQL, no shared-cache consultation. 2. Second-level cache, if the entity is cacheable. On a hit, hydrate an instance, put it in the persistence context, return it. 3. Database SELECT. Hydrate, put in the persistence context, and (if cacheable) publish to the shared cache. Step 1 taking priority is what preserves identity even when the shared cache holds newer data: within one session you keep the instance you already have. ## Where the two caches part company - **Correctness risk.** The first-level cache cannot be stale relative to your own transaction in a way that surprises you badly; the second-level cache can serve data another transaction has already changed, so it is an eventual-consistency mechanism. - **Invalidation.** The first-level cache needs none; it dies with the session. The second-level cache needs write-through invalidation, region eviction for bulk statements, and messaging across cluster nodes. - **Memory.** The first-level cache grows within one unit of work and is reclaimed at close. The second-level cache is a long-lived heap or off-heap structure sized as capacity. ## Common misstatement to avoid "The second-level cache makes my queries faster." Both entity caches are keyed by identifier. A JPQL query with a predicate is not answered by them; it executes SQL, and only the returned rows are resolved through the caches. Caching the *results* of a query is a separate, separately enabled mechanism with its own validity rules. ## How to answer Give scope, lifetime, visibility, opt-in status and stored form for each, then the lookup order, then one consequence apiece: the first-level cache helps only within a unit of work, and the second-level cache buys cross-session hits at the price of staleness and invalidation.

  • Why does the second-level cache store dehydrated state instead of entity instances?
    To keep sessions isolated and the data portable. If instances were shared, two sessions would mutate the same object and see each other's uncommitted changes. Storing a flat array of column values keyed by identifier lets each session hydrate its own instance, and makes the entry serialisable for off-heap or distributed storage.
  • If an entity is in the second-level cache with newer data than the copy in your session, which do you get from find()?
    The one already in your persistence context. The first-level lookup happens first and wins, because entity identity within a unit of work takes precedence over freshness. To pick up the newer state you must refresh the entity or use a new persistence context.

saying these in an interview costs you the question

  • Saying the second-level cache is enabled by default
  • Claiming the first-level cache is shared between concurrent requests or users
  • Believing the second-level cache stores ready-made entity objects that sessions share
  • Expecting either entity cache to answer a JPQL query with a WHERE clause
  • Describing the second-level cache as transactionally consistent with the database

context