What does enabling the hibernate.generate_statistics setting give you, and what kinds of numbers can you read out of Hibernate's Statistics API?
answer
- generate_statistics=true, read SessionFactory.getStatistics()
- prepareStatementCount = the N+1 detector
- QueryStatistics: count, total/avg/max time, rows
- cache hits/misses/puts per region = is caching earning its keep
- cumulative until clear(); factory-wide, not per request
basics
~20 sIt makes Hibernate count its own work into a Statistics object reachable from the SessionFactory: statements prepared, entities loaded/inserted/updated/deleted, queries executed with their times, collection and cache hit/miss/put counts, connections and flushes. Counters are cumulative until you clear them.
solid answer
~40 s`hibernate.generate_statistics=true` turns on internal instrumentation; you then read `SessionFactory.getStatistics()` (via `entityManagerFactory.unwrap(SessionFactory.class)`). Useful families of numbers: - **Work volume** — `getPrepareStatementCount()`, entity load/insert/update/delete counts, collection fetch/load counts, `getFlushCount()`, `getConnectCount()`, `getTransactionCount()`. - **Queries** — `getQueryExecutionCount()`, `getQueries()` plus per-query `QueryStatistics` (execution count, total/avg/max time, rows returned), and `getQueryExecutionMaxTimeQueryString()`. - **Caches** — second-level cache hit/miss/put counts overall and per region, query-cache hit/miss/put, natural-id cache. It is a diagnostic and metrics feed, not free: every operation touches counters, and `getQueries()` accumulates one entry per distinct query string, so dynamically built SQL can grow that map. Counters are cumulative from startup unless you call `clear()`, which is exactly what makes them convenient in tests: clear, run the code, assert.
code
java · 17 linesStatistics stats = emf.unwrap(SessionFactory.class).getStatistics();
stats.clear();
runTheCodePath();
long statements = stats.getPrepareStatementCount();
long entities = stats.getEntityLoadCount();
long queries = stats.getQueryExecutionCount();
QueryStatistics qs = stats.getQueryStatistics(
"select o from Order o where o.status = :s");
System.out.println(qs.getExecutionCount() + " runs, max "
+ qs.getExecutionMaxTime() + " ms");
CacheRegionStatistics region = stats.getCacheRegionStatistics("com.example.Country");
double hitRatio = region.getHitCount()
/ (double) (region.getHitCount() + region.getMissCount());go deeper
Know the setting name, that you read Statistics from the SessionFactory, and that statement and query counts are the headline numbers.
Name several counter families, explain that they are cumulative and factory-wide, and show the clear-then-assert pattern.
Use them to diagnose: compare load count against statement count, read cache hit ratios per region, and watch session open/close for leaks. Know the collection cost.
Decide whether they belong in production at all — counter contention, unbounded query maps with non-parameterised SQL, and the fact that monotonic counters need rating before they mean anything on a dashboard.
## What the setting does By default Hibernate keeps no bookkeeping about its own behaviour — that would cost something on every operation. `hibernate.generate_statistics=true` (or `Statistics.setStatisticsEnabled(true)` at runtime) switches on a `StatisticsImplementor` that Hibernate notifies as it works: each prepared statement, each entity loaded, each flush, each cache lookup. The accumulated numbers live on a single `Statistics` object owned by the `SessionFactory`. In plain JPA you get to it by unwrapping: `Statistics stats = entityManagerFactory.unwrap(SessionFactory.class).getStatistics();` ## The counters worth knowing **Statement and entity volume.** `getPrepareStatementCount()` and `getCloseStatementCount()` tell you how many JDBC statements the session factory has issued — the single most useful number for spotting a code path that fans out into per-row queries. Alongside it sit `getEntityLoadCount()`, `getEntityFetchCount()` (loads that were *not* satisfied from a cache or the persistence context), `getEntityInsertCount()`, `getEntityUpdateCount()`, `getEntityDeleteCount()`, and the collection equivalents (`getCollectionLoadCount()`, `getCollectionFetchCount()`, `getCollectionRecreateCount()`). **Queries.** `getQueryExecutionCount()` is the total; `getQueries()` returns every distinct query string seen, and `getQueryStatistics(String)` gives that query's execution count, total/average/min/max execution time and rows returned. `getQueryExecutionMaxTime()` and `getQueryExecutionMaxTimeQueryString()` name the slowest single execution since the last clear. **Caches.** Second-level cache activity is reported as hit/miss/put counts, both globally (`getSecondLevelCacheHitCount()` and friends) and per region via `getCacheRegionStatistics(regionName)`. The query cache and the natural-id cache have their own triples. The hit ratio — hits / (hits + misses) — is how you decide whether a cache region is earning its keep. Also note `getQueryCacheHitCount()` versus `getQueryCachePutCount()`: a region with many puts and few hits is pure overhead, and one with a high miss count is either mis-sized or holds data whose keys never repeat. **Sessions and connections.** `getSessionOpenCount()`, `getSessionCloseCount()`, `getConnectCount()`, `getFlushCount()`, `getTransactionCount()` and `getSuccessfulTransactionCount()`. A gap between opened and closed sessions is a leak; a gap between total and successful transactions is a rollback rate. ## Session-scoped view `session.getStatistics()` returns a `SessionStatistics` with `getEntityCount()` and `getCollectionCount()` — how many entities and collections this persistence context is currently holding. That is the direct way to catch a session that is accumulating tens of thousands of managed entities during a batch job, which is a memory and dirty-checking problem rather than a query-count problem. ## Semantics you must state correctly - Counters are **cumulative from startup** (or from the last `clear()`), not per request and not windowed. Exporting them to a monitoring system means exporting monotonic counters and letting the monitoring system rate them; the max-time fields are the exception, being stateful maxima that only reset on `clear()` and therefore misleading if you graph them raw. - They are **factory-wide**, aggregated across all threads. Concurrent requests interleave, so you cannot attribute a delta to one request unless the test is single-threaded. - They cover **Hibernate's** view of the world. Statements issued through a raw JDBC connection or a native query executed outside the session are not necessarily reflected the same way, and cache hits mean 'Hibernate did not need the database', which is precisely why the statement count can stay flat while the load count climbs. ## Cost and where it belongs The instrumentation uses concurrent counters, so it is cheap per operation but not free, and under high concurrency the shared counters are a contended write. The bigger risk is the per-query map: it keys on the query string, so an application that builds SQL by string concatenation with embedded literals can accumulate an unbounded number of `QueryStatistics` entries. With named/parameterised queries the set is bounded and the map is harmless. The pragmatic split: on permanently in test and staging, on in production only when you actually export the numbers and your queries are parameterised, and always paired with `clear()` discipline in tests so each assertion starts from zero.
- The entity load count is high but the prepared-statement count barely moved. What does that tell you?Most of those loads were satisfied without a round trip — from the persistence context, the second-level cache, or a single query that returned many rows and materialised many entities. It is generally a healthy shape. The opposite pattern, statement count tracking load count one-for-one, is the signature of per-row fetching.
- Why is getQueryExecutionMaxTime() a poor metric to graph directly in a dashboard?It is a stateful high-water mark that only resets when someone calls clear(), so once a single cold-start query takes two seconds the value stays at two seconds forever and the graph flatlines. Meaningful latency monitoring needs per-interval percentiles, which you build from the counters and times yourself or take from a JDBC-level or database-level source.
saying these in an interview costs you the question
- Thinking the counters are per-request or per-session — they are cumulative and factory-wide
- Assuming statistics are free and can always be left on, ignoring counter contention and the unbounded per-query map when SQL is built with inlined literals
- Confusing entity load count with statement count, and concluding there is an N+1 problem from the former alone
- Reading a low second-level cache hit count as 'the cache is broken' without checking whether the region is even being put to
- Believing generate_statistics also logs the SQL text — it does not; that is a separate logger