skip to content

L1 vs L2 Cache Scopes

The scope split every interview starts with: session-scoped L1 vs SessionFactory-scoped shared L2, and what L2 actually stores — dehydrated state, not object instances. Knowing that find() consults L2 but JPQL does not is the follow-up.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

questions

4

Inside one JPA persistence context (a Hibernate Session), you call em.find(Customer.class, 1L) twice. How many SELECT statements does Hibernate issue, and are the two returned references the same Java object?

level: juniorimportance: must knowfreq 68%

answer

  1. Persistence context = first-level cache = identity map
  2. Key is (type, id); value is instance plus loaded snapshot
  3. One SELECT, a == b
  4. Always on; only clear/detach/close shrink it
  5. Long session = memory growth; flush + clear per chunk

basics

~20 s

One SELECT, and the same object both times. The persistence context is an identity map: one managed instance per entity identity for as long as the session lives. It is always on and cannot be disabled; clear() or detach() removes entries.

solid answer

~50 s

Hibernate's first-level cache is the persistence context itself, scoped to a single `EntityManager`/`Session`. The first `find` executes the SELECT and stores the managed instance under a key of entity type plus identifier; the second lookup finds that key and returns the identical reference without touching the database, so `a == b` is true. This gives two guarantees. **Identity**: within a session there is exactly one object per database row, so changes made through one reference are visible through every other, and reference equality is a valid identity test for managed entities. **Repeatable read at session level**: the same row read twice yields the same state, even if another transaction changed it in between. It is not optional; the persistence context is how dirty checking and write ordering work. You can only shrink it: `clear()` empties it, `detach(entity)` removes one, `close()` ends it. `refresh(entity)` forces a reload of one instance from the database.

code

java · 6 lines
java
Customer a = em.find(Customer.class, 1L); // SELECT ...
Customer b = em.find(Customer.class, 1L); // no SQL
assert a == b;                            // same managed instance

a.setName("new");
assert b.getName().equals("new");          // one object, one row

go deeper

for a junior

Answer concretely: one SELECT, the same object, because the session keeps one instance per identifier and it is always on.

for a middle

Add the loaded snapshot and dirty checking, that queries still execute SQL but results are deduplicated against managed instances, and the clear/detach/refresh semantics.

for a senior

Bring in operational consequences: memory and flush-time growth in long sessions, chunked flush-and-clear, and when a stale managed instance shadowing fresh query results causes real bugs.

for a principal

Position the persistence context as the unit-of-work boundary: its lifetime defines transaction scope, memory profile and staleness of in-memory state, so session granularity is an architectural decision, not an implementation detail.

## What the first-level cache is The "first-level cache" is not a separate component you configure. It is the persistence context: the map of managed entities that an `EntityManager` (Hibernate `Session`) keeps for its lifetime. Its key is the pair (entity type, identifier); its value is the managed Java instance plus a snapshot of the state as loaded. Because of that map, the second `find` for the same identifier never reaches the database. It is a hash lookup returning the object already in memory. ## The identity guarantee Within one persistence context there is exactly one instance per row. This is the property that makes an ORM usable: - Two code paths that load customer 1 by different routes — a direct `find`, a JPQL query, navigating `order.getCustomer()` — receive the same object. A change made in one place is visible in all of them. - Reference equality (`==`) works as an identity test for managed entities in the same session, which is why `equals`/`hashCode` on entities is subtle only once instances are detached or compared across sessions. - Dirty checking has something to compare against: at flush Hibernate walks the managed entities, diffs each against its loaded snapshot, and generates UPDATEs only for what actually changed. ## Repeatable read at the session level Read a row, and no later read in that session will show you a different version of it, regardless of what other transactions committed, because the later read never leaves memory. That is a convenience, not a database isolation level; the freshness of the value is whatever it was at first load. If you need current state, `em.refresh(entity)` re-reads that row and overwrites the in-memory instance (discarding any unflushed changes to it). ## What it does not cover - **Queries still run.** A JPQL query always goes to the database, because the persistence context is keyed by identifier and cannot answer arbitrary predicates. What it does do is deduplicate the *results*: if a returned row belongs to an entity already managed, Hibernate discards the freshly read row data and hands you the existing instance. So a stale managed instance can "shadow" newer database values. - **Scalar projections** (`select c.name from Customer c`) are not entities and are not stored in the context at all. - **Nothing is shared between sessions.** Two concurrent requests each have their own persistence context and their own instances; the first-level cache never helps another user. ## Lifecycle and the batch pitfall Entries accumulate: every entity loaded or persisted stays until removed. In a long-running batch that reads a million rows in one session, all of them stay reachable, along with their snapshots, so memory grows roughly at twice the entity footprint and flush time grows because dirty checking scans everything. The standard remedy is to work in chunks, `flush()` then `clear()` after each chunk, so the context stays small. `detach(entity)` handles the single-object case; `close()` discards everything. One caution: `clear()` discards pending changes without writing them, so flush first when the writes matter. ## Relationship to the second-level cache Different scope entirely. The first-level cache is per session, always on, stores live objects, and dies with the session. The second-level cache is per `SessionFactory`, opt-in, shared by all sessions, and stores state rather than objects. A `find` consults the persistence context first and only then the shared cache and finally the database. ## How to answer this in an interview Give the concrete answer first — one SELECT, the same reference — then name the mechanism (persistence context as an identity map), then the two guarantees (identity, session-level repeatable read), then the limits (queries still execute, nothing shared between sessions, memory growth in long sessions). That is a complete junior-to-middle answer.

  • Can you disable the first-level cache?
    No. The persistence context is how JPA works: it provides entity identity, dirty checking and write ordering, so there is no switch to turn it off. The only control you have is scope and size — use short sessions, detach individual entities, or clear the context periodically in long-running work.
  • Another transaction updated the row after you loaded it. Does a second find() in the same session see the new value?
    No. The second find returns the instance already in the persistence context without querying, so you keep the state you first loaded. To see the committed change you must call refresh on that entity, or work in a new persistence context. Note that refresh overwrites any unflushed modifications you made to it.
  • What does clear() do to changes you have made but not flushed?
    It throws them away. clear() detaches every managed instance without flushing, so pending updates are never written. In batch code you flush first and then clear; forgetting the flush silently loses work.

Think of the session as a scratch desk for one unit of work: the first time you need a file you fetch it from the archive and put it on the desk; every later request hands you the very same sheet of paper, with all your pencil marks on it, until you clear the desk.

saying these in an interview costs you the question

  • Saying the first-level cache can be turned off in configuration
  • Believing it is shared between users or requests
  • Expecting a JPQL query to be answered from the persistence context without SQL
  • Assuming clear() flushes pending changes before detaching
  • Ignoring that a long-lived session grows without bound in a batch job

context

open as a page

Explain the order in which Hibernate looks for an entity when you call em.find(Customer.class, 42L), and why a JPQL query such as "select c from Customer c where c.city = :city" still executes SQL even when every matching row is already in the second-level cache.

level: middleimportance: must knowfreq 45%

basics

~20 s

find checks the persistence context, then the second-level cache, then the database. A JPQL query with a predicate cannot be answered by either, because both are keyed by identifier and neither can evaluate a WHERE clause — so the SQL runs and only the returned rows are resolved through the caches.

open as a page

Compare Hibernate's first-level and second-level caches: what is each one scoped to, who can see its contents, when is it active, and what is stored in it?

level: middleimportance: must knowfreq 62%

basics

~20 s

The first-level cache is the persistence context: per session, always on, holds live entity objects, dies with the session, private to one unit of work. The second-level cache belongs to the SessionFactory: opt-in, shared by all sessions in the JVM, and holds entity state, not objects.

open as a page

Hibernate's second-level cache stores entities in a dehydrated form rather than storing the entity objects themselves. What exactly is stored, why was it designed that way, and what does it cost?

level: seniorimportance: should knowfreq 30%

basics

~20 s

It stores a flat array of column values keyed by identifier, with to-one associations reduced to identifiers and collections held in separate regions. Sharing live objects would leak uncommitted changes between sessions, so each session hydrates its own instance — at the cost of rebuilding it on every hit.

open as a page