skip to content

Performance & Tuning

Making Hibernate fast on purpose: JDBC batching, stateless sessions, read-only hints, streaming, and the statistics that prove what SQL actually ran. Interviewers close Hibernate rounds with tuning questions because they expose who has operated the ORM under real load.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

You find code that loads every row of a table into a List with a JPQL query and then filters and counts it in Java. Why is that a problem, and what should the code do instead?

level: juniorimportance: must knowfreq 55%

answer

  1. query returns everything, Java keeps twelve
  2. entity + snapshot per row = heap multiple
  3. predicate never reaches the index
  4. count(e) not getResultList().size()
  5. stream/page + clear() for real bulk work

basics

~20 s

The database sends every row over the network and Hibernate turns each into a managed entity, so memory and time scale with table size and no index is used. Push the filter, the aggregate and the limit into the query instead.

solid answer

~60 s

Three separate costs stack up: 1. **Transfer** — every row crosses the network, however few are kept. 2. **Materialisation** — Hibernate instantiates an entity per row and keeps a load-time snapshot for dirty checking, so heap use is a multiple of the raw data. 3. **No index** — the predicate never reaches the database, so it cannot use an index; the engine does a full scan and the application does the comparison. And it degrades invisibly: with 500 rows in a test database it looks fine, at 5 million it is an OutOfMemoryError. The fix is to express the intent in the query: - filtering → a `where` clause with bound parameters; - counting → `select count(e) ...`, never `getResultList().size()`; - "the newest one" → `order by` plus `setMaxResults(1)`; - sums and grouping → `sum`/`group by`; - "does one exist" → `select count(e)` or an `exists` subquery. If a genuinely large set must be processed, page it or stream it rather than materialising it, and clear the persistence context as you go.

code

java · 10 lines
java
// anti-pattern: whole table into heap, filtered in Java
List<Customer> all = em.createQuery("select c from Customer c", Customer.class)
                       .getResultList();
long active = all.stream().filter(c -> c.getStatus() == ACTIVE).count();

// fix: the database answers the question it was asked
long active = em.createQuery(
        "select count(c) from Customer c where c.status = :st", Long.class)
    .setParameter("st", ACTIVE)
    .getSingleResult();

go deeper

for a junior

Say plainly that the database should do the filtering and counting, and name the query equivalents — WHERE, COUNT, ORDER BY with a limit.

for a middle

Explain the three costs — transfer, entity materialisation with dirty-check snapshots, and the lost index access — and mention that flush cost also rises with managed-entity count.

for a senior

Add the scaling-bomb framing, the correct pattern for genuine bulk work (chunking, streaming, clear()), and how you would catch it in review or tests.

for a principal

Treat it as a boundary rule: query layers return bounded, projected results, unbounded reads are forbidden without an explicit batching path, and tests enforce row-count budgets.

## The shape of the anti-pattern ```java List<Customer> all = em.createQuery("select c from Customer c", Customer.class).getResultList(); long active = all.stream().filter(c -> c.getStatus() == ACTIVE).count(); ``` The database was asked for everything; the decision about what mattered was made afterwards in Java. Variants: `getResultList().size()` for a count, `stream().max(...)` for the newest row, `stream().mapToLong(...).sum()` for a total, `contains()` for an existence check, `subList` for a page. ## Why it is expensive **Network and serialization.** Every column of every row is encoded by the database, sent over the wire and decoded by the driver. For a million-row table with wide columns this is easily hundreds of megabytes for a result the caller reduces to a single number. **Entity materialisation.** Hibernate does not hand back rows; it builds managed entities. Each one costs the object itself, its collections' placeholders, an entry in the persistence context, and a **load-time snapshot** of its state used later for dirty checking. A rough rule of thumb is several times the raw row size in heap. All of it is garbage a millisecond later. **Lost index access.** A predicate the database never receives cannot be answered by an index. `where status = 'ACTIVE'` against an index touches only matching rows; filtering in Java forces a full scan of the table regardless. **Lost aggregation.** Databases compute `count`, `sum`, `max` and `group by` while streaming, without materialising anything. Doing it in Java means materialising everything first. **Flush cost.** Because the entities are managed, every subsequent flush in that persistence context dirty-checks all of them, so an unrelated write later in the same unit of work becomes slow too. **A silent scaling bomb.** Nothing is wrong at 500 rows. The behaviour is correct at every size; only the cost changes. That is why this survives code review and shows up as a production incident. ## The replacements | Intent | Wrong | Right | |---|---|---| | filter | load all, `stream().filter` | `where` clause with bound parameters | | count | `getResultList().size()` | `select count(c) from Customer c where ...` | | exists | `list.isEmpty()` | `select count(c) ...` or an `exists` subquery | | newest | `stream().max(...)` | `order by ... desc` + `setMaxResults(1)` | | sum / group | `stream().collect(groupingBy)` | `select c.status, sum(c.total) ... group by c.status` | | page | `subList(a, b)` | `setFirstResult` / `setMaxResults` | Always bind parameters (`setParameter`) rather than concatenating values into the query string — concatenation is both a plan-cache killer and, for anything derived from user input, an injection risk. ## When you really do need many rows Batch jobs, exports and migrations legitimately touch large sets. The rule is not "never read a lot", it is "never hold a lot": - **Page through** with an ordered key so each chunk is bounded. - **Stream** the result rather than materialising a list, so rows are processed as they arrive. - **`em.clear()` between chunks**, otherwise the persistence context grows exactly as if you had loaded everything. - **Prefer projections** (or a stateless read path) when the job does not need managed entities at all. ## A useful heuristic Any time a Java collection operation appears immediately after `getResultList()`, ask whether the database could have done it. `filter`, `count`, `max`, `sum`, `sorted`, `distinct`, `limit` and `skip` all have direct query equivalents, and the query version does the work where the data already lives.

  • The code only needs to know whether any matching row exists. What is the cheapest query?
    A count restricted to the predicate, or an exists subquery — 'select count(c) from Customer c where c.status = :st' compared to zero, or a query returning a constant with setMaxResults(1). Both let the database stop early and neither materialises entities. Loading the list and calling isEmpty() forces every matching row to be transferred and turned into a managed entity just to answer a boolean.
  • A nightly job genuinely has to process every row. How do you do that without loading the table into memory?
    Process it in bounded chunks: page with an ordered key (or stream the result set) so only a chunk is materialised at a time, and call clear() on the persistence context after each chunk so processed entities become garbage. Read-only or projection-based access avoids the dirty-checking snapshots entirely. The goal is that heap use is proportional to the chunk size, not to the table size.

It is like phoning a warehouse and having them ship you the entire inventory so you can pick out one box in your driveway, instead of telling them the item number.

saying these in an interview costs you the question

  • Calling getResultList().size() to obtain a count
  • Claiming it is fine because 'the table is small' with no bound enforcing that
  • Believing Hibernate lazily streams a getResultList() so memory is not an issue
  • Replacing the Java filter with a filter over a cached list instead of a WHERE clause
  • Fixing the symptom by raising the heap

context

open as a page

In JPA/Hibernate, how do you make a JPQL/HQL query return instances of your own DTO class instead of managed entity objects, and what must that class provide?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Use a constructor expression: select new com.app.UserDto(u.id, u.name) from User u. Name the class fully qualified and give it a constructor whose parameter types match the selected expressions in order. Results are plain objects, not managed entities.

open as a page

In a plain JPA/Hibernate application, how do you make Hibernate print the SQL statements it executes, and how do you see the actual parameter values instead of the ? placeholders?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Set hibernate.show_sql=true to print to stdout, or better, set the org.hibernate.SQL logger to DEBUG so SQL goes through your logging framework. hibernate.format_sql pretty-prints it. Parameters stay as ? until you enable TRACE on Hibernate's bind-parameter logger.

open as a page

Some teams mark every JPA association fetch = FetchType.EAGER so that 'nothing fails later when the data is needed'. What does that cost at runtime, and what should they do instead?

level: middleimportance: must knowfreq 65%

basics

~20 s

EAGER fetches the association on every load path, even queries that never use it — extra selects per row or cartesian products from joined collections. It cannot be switched off per query. Map everything LAZY and fetch explicitly per use case with fetch joins or entity graphs.

open as a page

What extra work and memory does Hibernate spend on a query that returns managed entity objects, compared with the same query returning a DTO projection?

level: middleimportance: must knowfreq 58%

basics

~20 s

For each managed entity Hibernate builds the instance, keeps it in the persistence context, and stores a second copy of its loaded state for dirty checking. Every flush then compares all of them property by property. A DTO row costs one small object and none of that.

open as a page

How do you get Hibernate to send INSERT, UPDATE and DELETE statements to the database in JDBC batches, and what decides when a batch is actually executed?

level: middleimportance: must knowfreq 58%

basics

~20 s

Set hibernate.jdbc.batch_size to a positive number, for example 30. Hibernate then buffers statements with addBatch and calls executeBatch when the batch fills, when the SQL string changes, or at flush. Statements only batch at flush time, so nothing batches before a flush occurs.

open as a page

Why does an entity whose primary key is generated with @GeneratedValue(strategy = GenerationType.IDENTITY) not get its INSERT statements batched by Hibernate, and what would you map instead if batching matters?

level: middleimportance: must knowfreq 52%

basics

~20 s

With IDENTITY the database assigns the key during the INSERT, and persist() must return a managed entity that already has its identifier, so Hibernate executes the INSERT immediately instead of queuing it for flush. Nothing accumulates, so nothing batches. Use a sequence with an allocation size instead.

open as a page

In plain Hibernate, what does the query hint "org.hibernate.readOnly" (or Session.setDefaultReadOnly(true)) actually change about how loaded entities are managed, and why does it make large reads cheaper?

level: middleimportance: must knowfreq 55%

basics

~20 s

Entities are put in the persistence context without a loaded-state snapshot, so Hibernate cannot dirty-check them and never flushes an UPDATE for them. That removes roughly half the per-entity memory and all flush-time comparison work.

open as a page

What is Hibernate's StatelessSession, and which capabilities of the regular Session do you give up when you use it?

level: middleimportance: must knowfreq 45%

basics

~20 s

A command-oriented Hibernate API opened from the SessionFactory that has no persistence context. You call insert/update/delete explicitly. You lose the first-level cache, dirty checking, write-behind, cascades, collection handling, lifecycle events and interceptors, and second-level cache interaction.

open as a page

A screen backed by Hibernate takes several seconds to render. Which ORM-level mistakes do you look for first, and how do you confirm each one from evidence rather than guesswork?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Count the statements the request emits first. The usual causes: lazy loads in a loop, EAGER mappings pulling graphs, loading whole tables and filtering in Java, no pagination, entity graphs serialized to JSON, and per-row writes. Evidence before fixes.

open as a page

A job streams a multi-million-row HQL query using a cursor and a correctly configured JDBC fetch size, yet it still runs out of heap. What is filling memory, and how do you fix it?

level: seniorimportance: must knowfreq 45%

basics

~20 s

The persistence context. Every entity pulled from the cursor is registered with a state snapshot and stays strongly referenced. Fix: load read-only and call session.clear() every N rows — or select a projection instead of entities.

open as a page

Working through a Hibernate StatelessSession, you load an entity, change one of its fields, and also add an element to its @OneToMany collection that is mapped with CascadeType.ALL. What is written to the database, and why?

level: seniorimportance: must knowfreq 35%

basics

~20 s

Nothing at all. There is no dirty checking, so the field change is invisible until you call update(entity) — which then writes every mapped column. Collections are ignored and cascades do not run, so the new child is never inserted; you must insert it and set its foreign key yourself.

open as a page

When you run a JPQL/HQL query in Hibernate, what is the difference between calling getResultList() and consuming the results through getResultStream() or Hibernate's scroll()?

level: juniorimportance: should knowfreq 40%

basics

~20 s

getResultList() reads every row and builds every entity before returning. getResultStream() and scroll() keep a database cursor open and hand you rows as you pull them, so peak memory depends on how much you keep, not on the result size. Both must be closed.

open as a page

An import loop persists 200,000 entities through one EntityManager without ever flushing or clearing it. It starts fast, gets progressively slower, and eventually runs out of memory. Explain what is happening and how you would restructure it.

level: middleimportance: should knowfreq 45%

basics

~20 s

Every persisted entity stays managed, so the persistence context grows to 200,000 entities plus dirty-check snapshots. Each flush re-checks all of them, so work grows quadratically. Flush and clear in fixed batches, enable JDBC batching, or use a stateless write path.

open as a page

When a JPQL/HQL query selects a few columns, the result can be taken as an Object array, as a jakarta.persistence.Tuple, or as a DTO built by a constructor expression. How do these three shapes differ and when would you pick each?

level: middleimportance: should knowfreq 42%

basics

~20 s

Object[] gives positional untyped values — fine for a throwaway query. Tuple adds access by alias and element type, still generic. A DTO or record gives named, typed fields checked by the compiler and is the right choice for anything crossing a service boundary.

open as a page

What is the JDBC fetch size, how do you set it for a Hibernate query, and why can it be a hard requirement rather than a tuning knob when streaming a very large result set?

level: middleimportance: should knowfreq 40%

basics

~20 s

Fetch size is how many rows the JDBC driver pulls from the server per round trip. Hibernate sets it globally with hibernate.jdbc.fetch_size or per query. Some drivers buffer the entire result unless a fetch size (plus driver-specific conditions) makes them open a server-side cursor.

open as a page

What does enabling the hibernate.generate_statistics setting give you, and what kinds of numbers can you read out of Hibernate's Statistics API?

level: middleimportance: should knowfreq 45%

basics

~20 s

It makes Hibernate count its own work into a Statistics object reachable from the SessionFactory: statements prepared, entities loaded/inserted/updated/deleted, queries executed with their times, collection and cache hit/miss/put counts, connections and flushes. Counters are cumulative until you clear them.

open as a page

A background process keeps a single Hibernate Session open for its whole lifetime, and a web layer serializes still-managed entities to JSON while the persistence context is also still open. What problems does keeping a persistence context alive that long create, and how do you restructure it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The context accumulates every entity it ever loaded: unbounded memory, flush cost proportional to that count, stale data that is never re-read, and accidental writes flushed at commit. Lazy loads fired during serialization also hide fan-out. Use one short session per unit of work.

open as a page

You are asked to speed up a read-heavy list endpoint that currently loads mapped entities and maps them to a response in Java. Walk through converting it to a DTO read path and what you have to watch out for.

level: seniorimportance: should knowfreq 48%

basics

~20 s

Start from the response payload, write a query selecting exactly those columns into a DTO, push filtering, sorting, paging and aggregates into SQL, and drop the entity load. Watch for lost lazy navigation, formatting logic that lived on the entity, and per-row work now missing.

open as a page

You configured a JDBC batch size in Hibernate, but a transaction writing several different entity types still produces tiny batches. What do the hibernate.order_inserts and hibernate.order_updates settings do about that, and what do they cost?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A batch can only hold one SQL string, so interleaved entity types break it after every statement. order_inserts and order_updates sort the flush-time actions by entity type (and updates by primary key) so identical statements become contiguous and fill batches. Cost is sorting at flush plus changed statement order.

open as a page

After enabling JDBC batching in Hibernate, how do you actually prove that statements are being batched, given that hibernate.show_sql prints the same number of lines either way?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Statement logging prints each statement as Hibernate prepares it, so it never shows batching. Prove it with a proxying datasource such as datasource-proxy or p6spy that reports batch size, with TRACE logging on Hibernate's JDBC batch internals, or by counting round trips at the database.

open as a page

Your team keeps re-introducing code paths that fire one SQL statement per row, and it is only noticed in production. How would you turn the number of SQL statements a code path executes into an automatically enforced, failing assertion in the test suite?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Count statements around the code and assert on the count. Either clear Hibernate's Statistics and assert getPrepareStatementCount(), or wrap the DataSource with datasource-proxy/p6spy and assert its captured query count. Run the assertion in a single-threaded test against a real database dialect.

open as a page

A job must load 20 million rows from a file into one table. How do you choose between Hibernate's StatelessSession, a regular Session with periodic flush and clear, and plain JDBC batch inserts?

level: principalimportance: should knowfreq 35%

basics

~20 s

Decide by how much ORM machinery the rows actually need. Flat rows with no cascades or callbacks: StatelessSession, or plain JDBC if mappings add nothing. Graphs, events or auditing: batched Session with flush/clear. Pure single-table dumps: the database's native bulk loader beats all three.

open as a page

Hibernate 6 added an upsert() operation to StatelessSession. What does it do, what does it require of the entity, and when would you use it instead of insert() or update()?

level: middleimportance: nice to knowfreq 20%

basics

~20 s

upsert() writes a row without you knowing whether it already exists — one insert-or-update statement, typically SQL MERGE where the dialect supports it. It needs the identifier already assigned on the entity, and like all stateless writes it writes every mapped column.

open as a page

What does Hibernate's hibernate.jdbc.batch_versioned_data setting control, and why might a JDBC driver's behaviour force you to turn it off?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

It controls whether UPDATE and DELETE statements for entities carrying a version column are sent in JDBC batches. Hibernate decides a row was changed concurrently from the affected-row count; if a driver returns SUCCESS_NO_INFO from executeBatch, that check cannot be made, so batching must be disabled.

open as a page

Hibernate can log a warning for every query slower than a configured number of milliseconds. Which setting turns that on, what exactly does the reported time measure, and how does it differ from the database server's own slow-query log?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

Set hibernate.session.events.log.LOG_QUERIES_SLOWER_THAN_MS to a millisecond threshold. Hibernate then logs each slower statement to the org.hibernate.SQL_SLOW category with its elapsed time, measured client-side around JDBC execution — so it includes network and server queueing, unlike the database's own log.

open as a page

Someone proposes turning on Hibernate's second-level cache for every entity to fix slow reads, with no size or expiry configured per region. How do you evaluate that proposal, and what would you actually cache?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Caching makes existing lookups cheaper; it does not remove queries a bad fetch plan issues. Cache only small, read-mostly, by-id-accessed reference data, with an explicit size cap and expiry per region, and measure the hit ratio. Fix fetch plans first.

open as a page

How would you organise a service so that writes go through mapped JPA entities while reads are served by DTO projections — a lightweight read/write model split — and what does that approach cost you?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Keep one database and one schema. Commands load entities by id and mutate them; queries run projection queries into per-use-case DTOs and never return entities. Costs: two sets of types over one schema, duplicated knowledge of the mapping, and weaker reuse of entity behaviour and caching.

open as a page

How would you choose the JDBC batch size and shape the transactions for a nightly job that writes several million rows through Hibernate?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Measure rather than guess. Start around 30 to 50, raise it while throughput improves, and stop when gains flatten. Chunk the work into bounded transactions, flush and clear per chunk, avoid identity keys, and check whether the driver rewrites batches before tuning further.

open as a page

For a nightly export of tens of millions of rows through Hibernate, how would you decide between holding one cursor open for the entire job versus chunking the read into many short transactions?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Trade consistency against operability. One cursor gives a single consistent read but pins a connection and a transaction for hours and restarts from zero on failure. Chunking by a stable key gives restartability and short transactions, at the cost of a moving snapshot.

open as a page

showing 1–30 of 31