skip to content

Inside a single Hibernate Session you load the same database row twice by primary key. What do you get back the second time, and what mechanism guarantees that result?

level: juniorimportance: must knowfreq 70%

answer

  1. identity map: (type, id) -> one instance
  2. second find() = no SQL, same reference
  3. exists for dirty checking, not just speed
  4. session-scoped, always on, never expires
  5. key is the id — arbitrary queries always hit the DB

basics

~20 s

You get the very same Java instance — the two references are ==, and the second load usually issues no SQL. The persistence context is an identity map keyed by entity type plus primary key, holding at most one instance per row for the life of the session.

solid answer

~50 s

The persistence context — Hibernate's first-level cache — is an **identity map** from (entity type, primary key) to one instance. The first `find()` runs a SELECT, hydrates an object and registers it under that key; the second `find()` for the same id finds the entry and returns **the same object reference**, with no SQL at all. So `a == b` holds, which is what makes dirty checking coherent: there is exactly one place where changes to that row can live. The cache is on by default, cannot be turned off, and is scoped strictly to one session — a different session gets its own instances, so `==` no longer holds across them. That is also why entities compared or put into sets across sessions need an `equals`/`hashCode` based on a stable business key rather than a generated id.

code

java · 7 lines
java
Book a = em.find(Book.class, 1L);   // SELECT ... FROM book WHERE id = 1
Book b = em.find(Book.class, 1L);   // served from the persistence context
assert a == b;

em.detach(a);
Book c = em.find(Book.class, 1L);   // SELECT again
assert a != c;                       // a is now a stale detached instance

go deeper

for a junior

Say the same instance comes back and no second SELECT is issued, because the session keeps one object per id.

for a middle

Name the identity map explicitly, explain why it is required for dirty checking, and note that only primary-key lookups can be served from it.

for a senior

Add the operational consequences: never expires within a session, grows without bound, and the equals/hashCode implications for entities crossing session boundaries.

for a principal

Position session scope as the design lever — how long a unit of work should keep identity, when to prefer short sessions or stateless access, and where a second-level cache belongs instead.

## What the first-level cache actually is Every `EntityManager`/`Session` owns a persistence context. Internally that is, above all, an **identity map**: a map whose key is the entity type plus the primary key, and whose value is the single Java instance representing that row within this session. It is called the first-level cache because it sits in front of the database and can satisfy loads by id without SQL — but describing it as a cache undersells it. Its primary job is *identity*, not speed. ## The guarantee ``` Book a = em.find(Book.class, 1L); // SELECT ... Book b = em.find(Book.class, 1L); // no SQL assert a == b; // reference equality ``` One row, one object, for the lifetime of the session. Two consequences follow: 1. **Dirty checking is well-defined.** Hibernate keeps a snapshot of the loaded values next to the instance. If two different objects could represent row 1, two conflicting sets of changes could exist and Hibernate would have to guess which to write. The identity map makes that impossible. 2. **Object graphs are consistent.** If an `Order` and an `Invoice` both reference customer 5, they end up holding *the same* `Customer` object, so a change made through one path is visible through the other. ## It is not optional and it is not shared The first-level cache cannot be disabled — it is structural, not a feature. It is also strictly **session-scoped**: it is created with the session and dies with it. Nothing is shared between sessions, between threads or between application nodes; that is what the second-level cache (a different, optional, session-factory-scoped mechanism) is for. Two concurrent sessions each hold their own instance of row 1, and mutating one has no effect on the other until the change reaches the database. ## Which operations consult it - `em.find(Type.class, id)` — checks the map first; SELECT only on a miss. - `em.getReference(Type.class, id)` — returns the managed instance if present, otherwise a proxy; no SELECT until the proxy is used. - Association navigation by id — resolving `order.getCustomer()` goes through the same map. - Loading by **any other criterion** (a JPQL query, a lookup by email) always goes to the database; the map is keyed by primary key and cannot answer arbitrary predicates. ## Where it bites **Identity versus equality across sessions.** `find` in session A and `find` in session B return two objects. If your `equals`/`hashCode` uses the default object identity, those two are unequal — so an entity stored in a `HashSet` in one session cannot be found there after a round trip through another. The standard fix is an `equals`/`hashCode` built on a natural/business key, or on the id with a `hashCode` that stays constant before and after the id is generated. **It does not expire.** The instance sits in the map with whatever values were read. If another transaction commits a change to that row, your session keeps showing the old values until you `refresh()` the entity or drop it from the context. Reads by id in the same session are therefore repeatable regardless of the database's isolation level. **It only grows.** Every entity loaded or persisted in a session is retained, together with its dirty-check snapshot, until `evict`, `clear` or the end of the session. In a long-running or batch session that is a memory-growth problem; the standard remedy is periodic `flush()` + `clear()`. ## Controlling membership - `em.detach(entity)` / `session.evict(entity)` — remove one entry; the next `find()` re-reads from the database and produces a *new* instance. - `em.clear()` — empty the whole map. - `em.refresh(entity)` — keep the same instance but overwrite its fields from a fresh SELECT, discarding local changes. - `em.close()` — end the context entirely. Note the difference between `refresh` and evict-then-reload: `refresh` preserves object identity for anything already holding a reference; evict-and-reload creates a new object and leaves the old references pointing at a stale detached instance. ## Interview framing A strong answer states the guarantee (one instance per id per session), names the mechanism (identity map keyed by type + id), explains why it exists (coherent dirty checking, consistent object graphs) and immediately notes its two boundaries: it only answers primary-key lookups, and it never expires within the session.

  • Does the first-level cache help across HTTP requests or across application nodes?
    No. It lives and dies with a single Session/EntityManager, which typically spans one transaction or one request at most, and it is never shared between threads or JVMs. Caching across those boundaries is the job of the second-level cache, which is scoped to the SessionFactory and can be clustered, and it comes with invalidation concerns the first-level cache simply does not have.
  • Why can entities cached this way break HashSet behaviour when they cross session boundaries?
    Each session hydrates its own instance for a row, so two objects representing the same row are distinct references. With the default identity-based equals/hashCode they are unequal, so an entity added to a HashSet in one session will not be found there using an instance loaded in another. Defining equals/hashCode over a stable business key — or an id-based equals with a constant hashCode — restores correct set and map behaviour.

A coat check with one hook per ticket number. Hand over ticket 42 twice and you get the same coat back, not a copy — the desk simply cannot hold two coats under one number. Walk to a different venue (another session) and there is a different desk with a different coat.

saying these in an interview costs you the question

  • Calling it a performance cache and missing that its purpose is object identity.
  • Believing it is shared between sessions, threads or servers.
  • Thinking it can be disabled or configured with a TTL.
  • Expecting it to answer lookups by columns other than the primary key.
  • Assuming it refreshes itself when another transaction commits a change.

context