skip to content

Caching

Hibernate's layered caches beyond the session: the shared second-level cache, its concurrency strategies, and the query cache that is off for a reason. Interviewers use L2 questions to see whether you understand consistency costs, not just cache hit ratios.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

explore

questions

28

Inside one JPA persistence context (a Hibernate Session), you call em.find(Customer.class, 1L) twice. How many SELECT statements does Hibernate issue, and are the two returned references the same Java object?

level: juniorimportance: must knowfreq 68%

answer

  1. Persistence context = first-level cache = identity map
  2. Key is (type, id); value is instance plus loaded snapshot
  3. One SELECT, a == b
  4. Always on; only clear/detach/close shrink it
  5. Long session = memory growth; flush + clear per chunk

basics

~20 s

One SELECT, and the same object both times. The persistence context is an identity map: one managed instance per entity identity for as long as the session lives. It is always on and cannot be disabled; clear() or detach() removes entries.

solid answer

~50 s

Hibernate's first-level cache is the persistence context itself, scoped to a single `EntityManager`/`Session`. The first `find` executes the SELECT and stores the managed instance under a key of entity type plus identifier; the second lookup finds that key and returns the identical reference without touching the database, so `a == b` is true. This gives two guarantees. **Identity**: within a session there is exactly one object per database row, so changes made through one reference are visible through every other, and reference equality is a valid identity test for managed entities. **Repeatable read at session level**: the same row read twice yields the same state, even if another transaction changed it in between. It is not optional; the persistence context is how dirty checking and write ordering work. You can only shrink it: `clear()` empties it, `detach(entity)` removes one, `close()` ends it. `refresh(entity)` forces a reload of one instance from the database.

code

java · 6 lines
java
Customer a = em.find(Customer.class, 1L); // SELECT ...
Customer b = em.find(Customer.class, 1L); // no SQL
assert a == b;                            // same managed instance

a.setName("new");
assert b.getName().equals("new");          // one object, one row

go deeper

for a junior

Answer concretely: one SELECT, the same object, because the session keeps one instance per identifier and it is always on.

for a middle

Add the loaded snapshot and dirty checking, that queries still execute SQL but results are deduplicated against managed instances, and the clear/detach/refresh semantics.

for a senior

Bring in operational consequences: memory and flush-time growth in long sessions, chunked flush-and-clear, and when a stale managed instance shadowing fresh query results causes real bugs.

for a principal

Position the persistence context as the unit-of-work boundary: its lifetime defines transaction scope, memory profile and staleness of in-memory state, so session granularity is an architectural decision, not an implementation detail.

## What the first-level cache is The "first-level cache" is not a separate component you configure. It is the persistence context: the map of managed entities that an `EntityManager` (Hibernate `Session`) keeps for its lifetime. Its key is the pair (entity type, identifier); its value is the managed Java instance plus a snapshot of the state as loaded. Because of that map, the second `find` for the same identifier never reaches the database. It is a hash lookup returning the object already in memory. ## The identity guarantee Within one persistence context there is exactly one instance per row. This is the property that makes an ORM usable: - Two code paths that load customer 1 by different routes — a direct `find`, a JPQL query, navigating `order.getCustomer()` — receive the same object. A change made in one place is visible in all of them. - Reference equality (`==`) works as an identity test for managed entities in the same session, which is why `equals`/`hashCode` on entities is subtle only once instances are detached or compared across sessions. - Dirty checking has something to compare against: at flush Hibernate walks the managed entities, diffs each against its loaded snapshot, and generates UPDATEs only for what actually changed. ## Repeatable read at the session level Read a row, and no later read in that session will show you a different version of it, regardless of what other transactions committed, because the later read never leaves memory. That is a convenience, not a database isolation level; the freshness of the value is whatever it was at first load. If you need current state, `em.refresh(entity)` re-reads that row and overwrites the in-memory instance (discarding any unflushed changes to it). ## What it does not cover - **Queries still run.** A JPQL query always goes to the database, because the persistence context is keyed by identifier and cannot answer arbitrary predicates. What it does do is deduplicate the *results*: if a returned row belongs to an entity already managed, Hibernate discards the freshly read row data and hands you the existing instance. So a stale managed instance can "shadow" newer database values. - **Scalar projections** (`select c.name from Customer c`) are not entities and are not stored in the context at all. - **Nothing is shared between sessions.** Two concurrent requests each have their own persistence context and their own instances; the first-level cache never helps another user. ## Lifecycle and the batch pitfall Entries accumulate: every entity loaded or persisted stays until removed. In a long-running batch that reads a million rows in one session, all of them stay reachable, along with their snapshots, so memory grows roughly at twice the entity footprint and flush time grows because dirty checking scans everything. The standard remedy is to work in chunks, `flush()` then `clear()` after each chunk, so the context stays small. `detach(entity)` handles the single-object case; `close()` discards everything. One caution: `clear()` discards pending changes without writing them, so flush first when the writes matter. ## Relationship to the second-level cache Different scope entirely. The first-level cache is per session, always on, stores live objects, and dies with the session. The second-level cache is per `SessionFactory`, opt-in, shared by all sessions, and stores state rather than objects. A `find` consults the persistence context first and only then the shared cache and finally the database. ## How to answer this in an interview Give the concrete answer first — one SELECT, the same reference — then name the mechanism (persistence context as an identity map), then the two guarantees (identity, session-level repeatable read), then the limits (queries still execute, nothing shared between sessions, memory growth in long sessions). That is a complete junior-to-middle answer.

  • Can you disable the first-level cache?
    No. The persistence context is how JPA works: it provides entity identity, dirty checking and write ordering, so there is no switch to turn it off. The only control you have is scope and size — use short sessions, detach individual entities, or clear the context periodically in long-running work.
  • Another transaction updated the row after you loaded it. Does a second find() in the same session see the new value?
    No. The second find returns the instance already in the persistence context without querying, so you keep the state you first loaded. To see the committed change you must call refresh on that entity, or work in a new persistence context. Note that refresh overwrites any unflushed modifications you made to it.
  • What does clear() do to changes you have made but not flushed?
    It throws them away. clear() detaches every managed instance without flushing, so pending updates are never written. In batch code you flush first and then clear; forgetting the flush silently loses work.

Think of the session as a scratch desk for one unit of work: the first time you need a file you fetch it from the archive and put it on the desk; every later request hands you the very same sheet of paper, with all your pencil marks on it, until you clear the desk.

saying these in an interview costs you the question

  • Saying the first-level cache can be turned off in configuration
  • Believing it is shared between users or requests
  • Expecting a JPQL query to be answered from the persistence context without SQL
  • Assuming clear() flushes pending changes before detaching
  • Ignoring that a long-lived session grows without bound in a batch job

context

open as a page

Hibernate's second-level cache does nothing unless you configure it. Walk through everything required to make one entity class actually cached: the configuration properties, the extra library, and the annotations on the class.

level: juniorimportance: must knowfreq 55%

basics

~10 s

Three things. Set hibernate.cache.use_second_level_cache=true. Plug in a RegionFactory backed by a real provider (usually the JCache bridge plus EhCache, Infinispan or Caffeine). Mark the entity with jakarta.persistence.@Cacheable and Hibernate's @Cache.

open as a page

Hibernate has a query cache that is separate from its entity cache. What does it hold, and what two things must be done before a given JPQL query actually uses it?

level: juniorimportance: must knowfreq 48%

basics

~20 s

It caches the results of a query keyed by the query text and its parameters. You must set hibernate.cache.use_query_cache=true globally and mark each query cacheable individually — setCacheable(true) or the org.hibernate.cacheable hint. It also requires the second-level cache.

open as a page

In Hibernate, you put the org.hibernate.annotations.@Cache annotation on an entity's @OneToMany collection field. What does the resulting second-level cache region actually store, and what else has to be cached for that to save database work?

level: middleimportance: must knowfreq 50%

basics

~20 s

Only the element identifiers — an id list keyed by the owner's id, not the children's column data. Hibernate then resolves each id through the element entity's own cache region, so that entity must be cached too or you still hit the database.

open as a page

Hibernate lets you choose a second-level cache concurrency strategy per entity: READ_ONLY, NONSTRICT_READ_WRITE, READ_WRITE, or TRANSACTIONAL. What consistency does each guarantee, and what kind of data is each meant for?

level: middleimportance: must knowfreq 62%

basics

~20 s

READ_ONLY: never-modified data; updates are rejected; fastest. NONSTRICT_READ_WRITE: no locking, entry evicted on change, short stale window possible. READ_WRITE: soft locks give roughly read-committed consistency for mutable data. TRANSACTIONAL: cache enlists in the JTA transaction; strongest, slowest.

open as a page

You run a bulk JPQL statement such as "update Product p set p.price = p.price * 1.1 where p.category = :c" via Query.executeUpdate(). What happens to Hibernate's second-level cache and to entities already loaded in the current persistence context?

level: middleimportance: must knowfreq 48%

basics

~20 s

The statement runs as one SQL UPDATE without loading entities. Hibernate cannot know which rows changed, so it evicts the whole cache region for the affected tables. Entities already loaded in the session are not refreshed and stay stale until you clear or refresh them.

open as a page

When your code updates an entity through Hibernate and that entity is held in Hibernate's second-level cache, how does Hibernate stop the cached copy from going stale, and at what point in the transaction does that happen?

level: middleimportance: must knowfreq 50%

basics

~20 s

Hibernate issues the UPDATE itself, so it knows which cache key changed. At flush it marks the entry as in-flight; only after the database transaction commits does it replace or remove the entry. On rollback the entry is just dropped, never updated.

open as a page

Explain the order in which Hibernate looks for an entity when you call em.find(Customer.class, 42L), and why a JPQL query such as "select c from Customer c where c.city = :city" still executes SQL even when every matching row is already in the second-level cache.

level: middleimportance: must knowfreq 45%

basics

~20 s

find checks the persistence context, then the second-level cache, then the database. A JPQL query with a predicate cannot be answered by either, because both are keyed by identifier and neither can evaluate a WHERE clause — so the SQL runs and only the returned rows are resolved through the caches.

open as a page

Compare Hibernate's first-level and second-level caches: what is each one scoped to, who can see its contents, when is it active, and what is stored in it?

level: middleimportance: must knowfreq 62%

basics

~20 s

The first-level cache is the persistence context: per session, always on, holds live entity objects, dies with the session, private to one unit of work. The second-level cache belongs to the SessionFactory: opt-in, shared by all sessions in the JVM, and holds entity state, not objects.

open as a page

The JPA property jakarta.persistence.sharedCache.mode accepts ENABLE_SELECTIVE, DISABLE_SELECTIVE, ALL and NONE. Explain what each value means for which entities get cached, and which one you would run in production.

level: middleimportance: must knowfreq 42%

basics

~10 s

ENABLE_SELECTIVE (the practical default) caches only entities marked @Cacheable. DISABLE_SELECTIVE caches everything except @Cacheable(false). ALL caches every entity regardless of annotations, NONE caches none. Production should use ENABLE_SELECTIVE — opt in deliberately.

open as a page

A developer turns on Hibernate's query cache and marks a list query cacheable, but observes that on a cache hit the application still issues one SELECT per row. Explain why that happens and how to fix it.

level: middleimportance: must knowfreq 42%

basics

~20 s

For entity queries the query cache stores only identifiers. On a hit Hibernate must rebuild each entity, taking it from the entity second-level cache if present, otherwise loading it by id. Mark the returned entity types cacheable to remove those SELECTs.

open as a page

Walk through what Hibernate's READ_WRITE second-level cache strategy does to a cached entry while a transaction updates that row — what a soft lock is, what concurrent readers see — and how NONSTRICT_READ_WRITE differs.

level: seniorimportance: must knowfreq 50%

basics

~20 s

READ_WRITE replaces the cached entry with a soft lock before the write; while it is there readers miss the cache and hit the database and other writers cannot put. After commit the lock is swapped for the new state, or just released so the next read reloads. NONSTRICT does none of this: it simply evicts, accepting brief stale reads.

open as a page

You run four instances of the same application against one database, each with its own Hibernate second-level cache. What stale-data problems appear, and how do invalidation-based and replicated or distributed cache topologies differ in handling them?

level: seniorimportance: must knowfreq 40%

basics

~20 s

A commit on one node leaves the other three holding the old rows until they are told. Invalidation topologies broadcast "forget this key" so other nodes reload from the database; replicated or distributed topologies ship the new value instead. Asynchronous messaging leaves a stale window either way.

open as a page

Hibernate maintains an internal region commonly called the update-timestamps cache. Explain the role it plays in deciding whether a previously cached query result may be returned, and what its granularity means in practice.

level: seniorimportance: must knowfreq 35%

basics

~20 s

It maps each table (query space) to the timestamp of its last write. A cached query result carries the timestamp it was produced at; if any table it touched was written later, the result is treated as stale and re-executed. Granularity is per table, not per row.

open as a page

An entity is mapped with Hibernate's READ_ONLY second-level cache concurrency strategy, and some code path modifies a managed instance of it and commits. What happens, and why is READ_ONLY the cheapest strategy?

level: juniorimportance: should knowfreq 42%

basics

~20 s

Hibernate refuses the cache update and throws (an unsupported-operation error) when flushing the change, failing the transaction. READ_ONLY is cheapest because immutable data needs no locking, no version reconciliation and no invalidation — entries are stored once and served forever.

open as a page

Which APIs let you remove entries from Hibernate's second-level cache programmatically, and what is the operational risk of the broadest of them?

level: juniorimportance: should knowfreq 36%

basics

~20 s

The portable jakarta.persistence.Cache from the EntityManagerFactory can evict one entity by id, a whole entity type, or everything. Hibernate's own SessionFactory cache adds region-level methods. Evicting everything on a live system sends all traffic straight to the database.

open as a page

An order entity in Hibernate has a second-level-cached @OneToMany collection of 500 line items. One line item is added to that collection. What happens to that collection's cache entry, and what does the answer imply for write-heavy associations?

level: middleimportance: should knowfreq 40%

basics

~20 s

The whole entry for that one owner is invalidated and rewritten with the new id list — there is no per-element delta. Other owners' entries are untouched. Frequently modified collections therefore churn their cache entry on every write and rarely pay off.

open as a page

A Hibernate entity marks a business key such as a book's ISBN with @NaturalId and code looks rows up with session.bySimpleNaturalId(Book.class).load(isbn). What does adding Hibernate's @NaturalIdCache annotation change, and what SQL runs with and without it?

level: middleimportance: should knowfreq 35%

basics

~20 s

@NaturalIdCache stores the natural-key-to-primary-key mapping in its own second-level region, so the lookup skips the select id where isbn = ? resolution query. The entity state still comes from the entity cache or a load by primary key.

open as a page

When Hibernate stores a cacheable query result, what makes up the key it is stored under? Explain what that implies for queries that take bind parameters or use paging.

level: middleimportance: should knowfreq 30%

basics

~20 s

The key covers the query text, every bind parameter value and type, the paging window (first result and max results), the result-set transformation, and the tenant identifier. Each distinct combination is a separate entry, so varied parameters mean no reuse.

open as a page

A Hibernate application has enabled second-level caching on a parent entity's @OneToMany association, yet loading each parent still emits one SELECT per child row. How would you diagnose and fix that?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Almost always the child entity is not cached: the collection region stores ids, so each id resolves with a primary-key select. Confirm with per-region cache statistics, then add caching to the child entity, or add batch fetching, or stop caching the collection.

open as a page

Hibernate's second-level cache stores entities in a dehydrated form rather than storing the entity objects themselves. What exactly is stored, why was it designed that way, and what does it cost?

level: seniorimportance: should knowfreq 30%

basics

~20 s

It stores a flat array of column values keyed by identifier, with to-one associations reduced to identifiers and collections held in separate regions. Sharing live objects would leak uncommitted changes between sessions, so each session hydrates its own instance — at the cost of rebuilding it on every hit.

open as a page

Hibernate implements no cache itself — it delegates to a provider through a RegionFactory. Explain how that plug-in point works, and how you would choose between EhCache, Infinispan and Caffeine behind the JCache (JSR-107) bridge.

level: seniorimportance: should knowfreq 33%

basics

~20 s

RegionFactory is Hibernate's SPI for named cache regions; hibernate-jcache adapts it to any JSR-107 provider. Caffeine is a fast in-JVM cache, EhCache adds sizing/disk tiers, Infinispan adds real clustering — pick by deployment topology, not benchmarks.

open as a page

How do you control the maximum size and expiry of one individual Hibernate second-level cache region, and how do you work out which region name corresponds to a given entity or collection?

level: seniorimportance: should knowfreq 28%

basics

~20 s

Regions default to the fully-qualified entity name; collection regions append the field name. Override with @Cache(region="..."), prefix all with hibernate.cache.region_prefix. Size and TTL are set per region in the provider's config file, not by Hibernate.

open as a page

You are deciding the Hibernate second-level cache concurrency strategy for each entity in a system: some are immutable reference tables, some change a few times a day, some are written constantly. How do you make that call per entity, and when do you decide not to cache an entity at all?

level: principalimportance: should knowfreq 34%

basics

~20 s

Match strategy to mutability and read/write ratio: immutable → READ_ONLY; rarely changed and staleness-tolerant → NONSTRICT_READ_WRITE; mutable, read-mostly, staleness-sensitive → READ_WRITE; atomic cache/DB under JTA → TRANSACTIONAL. Write-heavy or low-reuse entities should not be cached at all.

open as a page

How do you decide whether a given entity belongs in Hibernate's second-level cache, and in what situations does enabling that cache make a system worse rather than faster?

level: principalimportance: should knowfreq 34%

basics

~20 s

Cache entities that are read far more than written, small, reused across sessions, and tolerant of brief staleness. It hurts for write-heavy or huge datasets, for data that must be exactly fresh, and for workloads that read by query rather than by identifier, where the cache adds cost and no hits.

open as a page

You are deciding which parts of a domain model should get a shared second-level cache and how much heap to budget for it. How do you make that call, and what would make you leave the second-level cache switched off entirely?

level: principalimportance: should knowfreq 28%

basics

~20 s

Cache small, hot, rarely-written, id-accessed data whose staleness the domain tolerates. Skip write-heavy, huge, or externally-written tables. Leave L2 off entirely when the database is not the bottleneck, or when a correctness-critical read cannot tolerate any staleness window.

open as a page

Hibernate's query cache is disabled by default. Explain why that default is chosen, and describe the workload shape where enabling it genuinely pays off versus where it quietly makes an application slower.

level: principalimportance: should knowfreq 30%

basics

~20 s

It is off by default because it is easy to make things worse: fine-grained keys, table-coarse invalidation, dependence on the entity cache, and real write cost. It pays off only for repeated queries with low-cardinality parameters over rarely written tables.

open as a page

How do you decide which associations in a Hibernate application deserve a second-level collection cache region and which should be left uncached?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Cache read-mostly, bounded associations that are hot on a navigation path, and only when the element entity is cached too. Skip large or frequently modified collections: each write rewrites the whole id list and, in a cluster, broadcasts invalidation.

open as a page