skip to content

A Hibernate entity marks a business key such as a book's ISBN with @NaturalId and code looks rows up with session.bySimpleNaturalId(Book.class).load(isbn). What does adding Hibernate's @NaturalIdCache annotation change, and what SQL runs with and without it?

level: middleimportance: should knowfreq 35%

answer

  1. Resolve natural key -> id, then load by id
  2. ##NaturalId region holds ids only
  3. No cache: select id from book where isbn = ?
  4. Immutable by default; mutable = true tracks changes
  5. Both caches warm = zero SQL

basics

~20 s

@NaturalIdCache stores the natural-key-to-primary-key mapping in its own second-level region, so the lookup skips the select id where isbn = ? resolution query. The entity state still comes from the entity cache or a load by primary key.

solid answer

~50 s

A natural-id lookup is always two steps: **resolve** the natural key to the primary key, then **load** the entity by that key. Without `@NaturalIdCache`, resolution is cached only for the lifetime of the session; a cold lookup issues `select id from book where isbn = ?`, then loads the entity — from the entity cache if it is cached there, otherwise `select ... from book where id = ?`. With `@NaturalIdCache`, the mapping `isbn -> id` is kept in a dedicated second-level region, named `com.acme.Book##NaturalId` by default. A hit removes the resolution query entirely. Combined with `@Cache` on the entity, a repeated lookup by ISBN executes **no SQL at all**; with `@NaturalIdCache` alone you still pay one primary-key select. Natural ids are immutable by default — altering one fails at flush unless the field is declared `@NaturalId(mutable = true)`, in which case Hibernate removes the stale mapping and inserts the new one when the value changes.

code

java · 14 lines
java
@Entity
@Cache(usage = CacheConcurrencyStrategy.READ_WRITE)
@NaturalIdCache
public class Book {
    @Id @GeneratedValue Long id;

    @NaturalId
    @Column(unique = true, nullable = false)
    String isbn;

    String title;
}

Book b = session.bySimpleNaturalId(Book.class).load("978-0134685991");

go deeper

for a junior

Know that @NaturalId marks a business key, that bySimpleNaturalId loads by it, and that @NaturalIdCache caches the key-to-id mapping.

for a middle

Walk the resolve-then-load flow and state exactly which SQL disappears with @NaturalIdCache versus with entity caching, plus the default immutability.

for a senior

Discuss the dedicated ##NaturalId region, mutable natural ids and their bookkeeping, eviction after bulk statements, and why precise invalidation beats a cached query.

for a principal

Weigh natural-id caching as a lookup strategy across services — key stability, unique-constraint enforcement, memory for the mapping region, and when a dedicated lookup index or read model is the better answer.

## What a natural id is A *natural id* is the business key of a row: an ISBN for a book, a SKU for a product, a login for a user. The table still has a surrogate primary key (`id`), because Hibernate wants a stable, narrow identifier for associations and caching, but application code frequently knows the business key instead. `@NaturalId` on one or more fields declares that key to Hibernate and unlocks a dedicated lookup API: ```java Book b = session.bySimpleNaturalId(Book.class).load(isbn); // single-attribute Book b = session.byNaturalId(Book.class) .using("isbn", isbn).load(); // general form ``` (`bySimpleNaturalId` only exists when the entity has exactly one `@NaturalId` attribute; a composite natural id must use `byNaturalId(...).using(...)` for each part.) ## Why this beats an equivalent JPQL query `select b from Book b where b.isbn = :isbn` returns the same row, but it always goes to the database unless the query cache is on, and it never consults the entity cache. The natural-id API is deliberately shaped as *resolve then load*, so both halves can be served from caches, and it participates in the persistence context: a natural-id load can be satisfied entirely from the current session, and `getReference`-style deferred loading is available via `getReference()` instead of `load()`. ## The two-step resolution in detail **Step 1 — resolve.** Hibernate needs the primary key that corresponds to the natural key. - Session (L1) scope: every session keeps a natural-id resolution map, so repeated lookups within one session resolve for free, and entities loaded by any means register their natural-id mapping there. - Without `@NaturalIdCache`, a cold session must ask the database: `select id from book where isbn = ?`. Note this returns *only the identifier*. - With `@NaturalIdCache`, the mapping is stored in the second-level cache in a region of its own, whose default name is the entity name suffixed with `##NaturalId` (e.g. `com.acme.Book##NaturalId`). A hit there skips the resolution SQL. **Step 2 — load.** With the primary key in hand, Hibernate performs the ordinary entity load: persistence context → entity region (if the entity is annotated `@Cache`/`@Cacheable`) → `select ... from book where id = ?`. This is why the region layout matters: **natural-id regions hold only identifiers, exactly like collection regions hold only element ids**. Entity state is never duplicated into them. The practical corollary is the same as for collections: `@NaturalIdCache` without entity caching buys you one query instead of two, not zero. ``` @NaturalIdCache only : select ... from book where id = ? @Cache only : select id from book where isbn = ? (then cache hit) both, warm : no SQL neither, cold session : two selects ``` ## Mutability and staleness By default `@NaturalId` fields are treated as immutable. Change one on a managed entity and Hibernate refuses at flush time with an error along the lines of *"An immutable natural identifier of entity com.acme.Book was altered"*. That is a feature: an immutable key means a cached mapping can never silently rot. If the key genuinely changes — a user renaming their login — declare `@NaturalId(mutable = true)`. Hibernate then tracks the previous value and, at flush, removes the outdated `oldValue -> id` entry from the natural-id region and inserts the new one. The cost is extra bookkeeping per update, so only mark keys mutable when they really are. As with any second-level region, statements that bypass the persistence context (bulk JPQL update/delete, native SQL, or another application writing the same table) leave the mapping stale; those need explicit eviction via `sessionFactory.getCache().evictNaturalIdData(Book.class)`. ## Consistency of the cached mapping A natural-id mapping is invalidated precisely, by entity change, rather than by a coarse timestamp. This is the key advantage over caching the equivalent lookup query: a query-result entry is validated against an update-timestamps region, so any write to the `book` table can invalidate every cached query over books, while a natural-id entry survives unrelated writes and is removed only when *that* entity's natural id changes or the entity is deleted. ## Practical guidance - Use `@NaturalId` wherever a stable business key exists — the API is clearer than a hand-written query and it is the only lookup path that can be answered from the session. - Add `@NaturalIdCache` when lookups by that key are hot and the entity is read-mostly, and pair it with `@Cache` on the entity so both halves are covered. - Batch form: `byMultipleNaturalId(Book.class).multiLoad(isbn1, isbn2, ...)` resolves several keys in one round trip, and it consults the same caches. - Enforce the natural id with a unique constraint in the schema; Hibernate does not add one automatically for you, and a duplicate business key breaks the resolution's one-row assumption.

  • What happens if application code changes a field annotated with @NaturalId?
    By default natural ids are immutable, and Hibernate raises an error at flush time saying an immutable natural identifier of the entity was altered. To allow it you declare `@NaturalId(mutable = true)`, after which Hibernate tracks the old value and, on flush, removes the stale mapping from the natural-id cache region and stores the new one.
  • Why prefer a natural-id lookup over a JPQL query on the same column when caching is involved?
    A JPQL query always goes to the database unless the query cache is enabled, and even then its result entry is validated against an update-timestamps region, so any write to that table can invalidate it. The natural-id path resolves through a dedicated id-mapping region that is invalidated only when that entity's natural id or existence changes, and it can be answered from the current session's resolution map.

saying these in an interview costs you the question

  • Saying @NaturalIdCache caches the entity's data — it caches only the natural-key-to-id mapping.
  • Assuming @NaturalIdCache alone means no SQL, forgetting the load by primary key.
  • Thinking a JPQL query on the natural-id column will use the natural-id cache region.
  • Not knowing natural ids are immutable by default and that altering one fails at flush.
  • Believing @NaturalId creates a unique constraint in the database schema by itself.

context