Hibernate implements no cache itself — it delegates to a provider through a RegionFactory. Explain how that plug-in point works, and how you would choose between EhCache, Infinispan and Caffeine behind the JCache (JSR-107) bridge.
answer
- RegionFactory SPI → regions + access strategies
- hibernate-jcache = JSR-107 bridge, provider-agnostic
- Caffeine local/fast, EhCache tiering, Infinispan clustered
- Sizing/TTL live in the provider's config, not Hibernate
- N nodes local cache = N staleness windows
basics
~20 sRegionFactory is Hibernate's SPI for named cache regions; hibernate-jcache adapts it to any JSR-107 provider. Caffeine is a fast in-JVM cache, EhCache adds sizing/disk tiers, Infinispan adds real clustering — pick by deployment topology, not benchmarks.
solid answer
~50 sHibernate defines `org.hibernate.cache.spi.RegionFactory`. It asks the factory for **regions** by name and performs get/put/evict against them, plus the concurrency-strategy contract (read-only, read-write, nonstrict, transactional). Any cache library that can satisfy that SPI can back L2. The usual bridge is the `hibernate-jcache` module (`hibernate.cache.region.factory_class=jcache`), which maps regions onto JSR-107 `Cache` objects obtained from a `CacheManager`. Then the choice is a JCache provider: - **Caffeine** (via `caffeine-jcache`) — local, in-heap, excellent eviction quality and lowest latency. Right when each node may hold its own copy. - **EhCache 3** — local with richer declarative sizing, heap/offheap/disk tiers, mature per-region XML. - **Infinispan** — the one to pick when nodes must coordinate: invalidation or replication across a cluster, and a native Hibernate region factory beyond the JCache API. The deciding question is topology, not raw throughput: with N nodes and a local cache you have N independently stale copies, and only cluster-aware invalidation fixes that.
code
properties · 5 lineshibernate.cache.use_second_level_cache=true
hibernate.cache.region.factory_class=jcache
hibernate.javax.cache.provider=com.github.benmanes.caffeine.jcache.spi.CaffeineCachingProvider
hibernate.javax.cache.uri=classpath:application.conf
hibernate.javax.cache.missing_cache_strategy=failgo deeper
Know that Hibernate delegates caching to a provider through RegionFactory and that JCache is the usual bridge.
Name the mainstream providers, explain that sizing and TTL live in the provider's config, and that missing regions fail startup by default.
Drive the choice from topology and write patterns, and describe how you verify per-region hit ratio and eviction rates in production.
Weigh the operational cost of introducing a distributed grid against simply narrowing what is cached, and treat cross-node staleness as a domain decision rather than a configuration one.
## The SPI Hibernate's caching is deliberately pluggable. The core contract is `org.hibernate.cache.spi.RegionFactory`, which Hibernate uses to build **regions** — named key/value stores — of a few kinds: entity regions (`DomainDataRegion` in modern versions), collection regions, natural-id regions, the query-results region, and the update-timestamps region. On top of raw storage, the factory supplies **access strategies** implementing the semantics of `READ_ONLY`, `READ_WRITE`, `NONSTRICT_READ_WRITE` and `TRANSACTIONAL`, including the soft-lock protocol that read-write uses to keep a concurrent write from publishing stale state. Hibernate never stores objects there. It stores dehydrated state arrays keyed by a cache key built from the entity id, the entity's role, and (where relevant) the tenant identifier. That is why the provider only needs to be a competent map with eviction — all ORM semantics live on Hibernate's side of the SPI. ## The JCache bridge Rather than maintain an adapter per cache product, Hibernate ships `hibernate-jcache`, which implements `RegionFactory` over **JSR-107**. Configuration is small: ``` hibernate.cache.region.factory_class=jcache hibernate.javax.cache.provider=<CachingProvider FQCN> # needed only if several are on the classpath hibernate.javax.cache.uri=classpath:cache-config.xml # the provider's own config file hibernate.javax.cache.missing_cache_strategy=fail ``` Hibernate resolves a `CachingProvider`, obtains a `CacheManager` from the URI, and calls `getCache(regionName)` per region. Region sizing, expiry and tiering are therefore expressed in the **provider's** configuration language, not in Hibernate properties — a frequent surprise for people expecting `hibernate.cache.*.ttl` knobs. The `missing_cache_strategy` default of `fail` matters: if the provider config has no cache for `com.acme.Product`, startup breaks rather than quietly creating an unbounded region. `create-warn` is a reasonable transitional setting; `create` in production invites unbounded growth. ## Choosing a provider **Caffeine** is a modern in-JVM cache with a near-optimal admission/eviction policy (W-TinyLFU) and very low overhead. Its `caffeine-jcache` module makes it a JSR-107 provider. Choose it for single-node deployments, or multi-node deployments where per-node copies of reference data are acceptable. **EhCache 3** is also local but offers declarative, well-documented tiering — heap, offheap (serialized, outside the GC's reach), and disk — with per-region sizing in units of entries or bytes. Choose it when regions are large enough that keeping them on heap hurts GC, or when operators want XML-level control. **Infinispan** is a distributed data grid. Beyond JCache it ships a Hibernate-specific region factory that understands entity and query regions and does **cluster-wide invalidation**: when a node writes, peers drop their copies. Choose it when correctness across nodes matters more than the last microsecond — and accept that you have added a distributed system to operate, with its own failure modes. ## The real decision axis Raw benchmarks rarely decide this. The questions that do: 1. **How many nodes, and can they diverge?** A local cache on N nodes means N independent staleness windows, each bounded only by TTL. If the domain tolerates that (currency lists, feature catalogues), local is fine and far simpler. If not, you need cluster-aware invalidation — or you should not cache that entity. 2. **Who else writes the data?** A batch job, another service, or Hibernate bulk HQL `update`/`delete` can bypass the cache entirely. No provider fixes that; it is a modelling decision. 3. **Heap versus process memory.** Offheap or a grid keeps large datasets out of the GC's working set at the price of serialization cost per access. 4. **Operability.** Every provider adds config files, metrics and a failure mode. A distributed cache adds split-brain and rebalancing questions your on-call rota now owns. ## Verification Whichever provider you choose, wire `hibernate.generate_statistics=true` and export per-region hit/miss/put counts alongside the provider's own metrics. Two numbers matter: hit ratio per region (is this region earning its heap?) and eviction rate (is it too small to ever warm up?).
- Your application runs on four nodes with a local, in-JVM L2 cache. What consistency do you actually get?Each node has its own copy, so a write on node A updates or invalidates only node A's region; the other three keep serving their previous state until their entry expires or is evicted. The staleness window is bounded by TTL, not by the write. That is acceptable for slowly changing reference data and unacceptable for anything a user expects to see change immediately, which is where cluster-aware invalidation or simply not caching the entity is the answer.
- Why does Hibernate no longer let you configure region size and TTL through hibernate.* properties?With the JCache bridge, Hibernate only asks the CacheManager for a named cache; everything about capacity, eviction policy, tiering and expiry belongs to the provider's own configuration model, which differs widely between EhCache, Caffeine and Infinispan. Hibernate deliberately does not try to abstract that, so sizing lives in ehcache.xml, the Caffeine spec, or Infinispan XML, keyed by region name.
Hibernate is a hotel front desk that only knows how to ask for a numbered pigeonhole; whether the pigeonhole wall is a small fast rack, a big filing cabinet with a basement, or a network of racks in other buildings is the provider's business.
saying these in an interview costs you the question
- Claiming Hibernate has a built-in cache implementation
- Thinking a local cache gives cluster-wide consistency
- Expecting hibernate.* properties to control TTL or maximum entries
- Letting Hibernate auto-create missing regions in production
- Choosing a distributed grid for a single-node app because it benchmarks well