You need to render 100 parent records, each with its child collection, from a Hibernate application. Compare loading the children with a fetch join, with batch fetching, and with subselect fetching — how do you choose?
answer
- Round trips vs rows vs pagination
- join fetch: 1 query, duplicate rows, HHH000104 in memory
- batch: N/size queries, paging-safe, composes
- subselect: 2 queries, ignores the outer limit
- Global batch size as the baseline, join fetch per query
basics
~20 sFetch join: one query, but duplicates parent rows and breaks database-side pagination. Batch fetching: a handful of IN-clause queries, safe with paging and partial traversal. Subselect: exactly one extra query but replays the parent query and ignores its row limit. Paginated screens use batching; unpaginated full traversal favours join or subselect.
solid answer
~60 sAll three remove the per-parent query; they differ in what they cost. **Fetch join** — `join fetch` returns everything in one statement. It multiplies parent rows by collection size, so the driver transfers far more rows than entities; you need `distinct`/deduplication, only one collection can be joined without a cartesian explosion, and `setMaxResults` forces Hibernate to paginate **in memory** after loading the whole result — the HHH000104 warning. **Batch fetching** — `@BatchSize` or `hibernate.default_batch_fetch_size` groups the lazy loads into `where fk in (…)` queries. Round trips are N/batchSize, but there is no row multiplication, pagination is safe, several collections compose fine, and untouched collections cost nothing. **Subselect** — one extra query for all parents by replaying the parent query as a subquery. Excellent for full traversal of a cheap parent query; wrong under pagination, because the subselect ignores the outer limit. My default: lazy mappings plus a global batch fetch size, `join fetch` per query for the one collection a screen always needs, subselect only for unpaginated batch jobs.
code
java · 11 lines// 1. fetch join - one statement, duplicated parent rows, no DB-side paging
em.createQuery("select o from Order o join fetch o.lines where o.status = :s", Order.class);
// 2. batch fetching - mapping stays lazy
@OneToMany(mappedBy = "order") @BatchSize(size = 25)
private Set<OrderLine> lines;
// or: hibernate.default_batch_fetch_size = 25
// 3. subselect - one extra query for all parents of the previous query
@OneToMany(mappedBy = "order") @Fetch(FetchMode.SUBSELECT)
private Set<OrderLine> lines;go deeper
Know that a fetch join loads everything in one query and that batching groups lazy loads, and that both beat one query per parent.
Contrast row multiplication against round trips, and know that a collection fetch join plus maxResults paginates in memory.
Give a decision rule per use case, cover multiple collections, detached use, and how you would measure statement count and rows to decide.
Set the systemic policy — lazy mappings plus a global batch size as the floor, per-query fetch plans as the mechanism, and query-count assertions in tests so regressions are caught rather than discovered in production.
## Frame the comparison properly All three strategies eliminate the one-query-per-parent pattern. Choosing between them is a trade of **round trips** against **rows transferred** against **compatibility with pagination and partial traversal**. Say that first; then the specifics land as evidence rather than trivia. ## Fetch join `select o from Order o join fetch o.lines where o.status = :s` produces a single statement with a join. *Strengths.* One round trip. The collection is initialised eagerly for exactly the parents you asked about, and it works no matter what happens to the session afterwards, which is why it is the standard answer to "I need this data outside the transaction". *Costs.* The result set has one row per (parent, child) pair. A hundred parents with twenty lines each is two thousand rows carrying the parent columns twenty times over — bandwidth and driver-side work far beyond the entity count. Hibernate de-duplicates entities, but you must handle duplicate roots in the returned list (`distinct` in JPQL, or Hibernate 6's automatic de-duplication of entity results). Joining **two** collections multiplies the two sizes together and, for bag-typed collections, is rejected outright with `MultipleBagFetchException`. *The pagination problem.* Combine `join fetch` on a collection with `setMaxResults` and the limit cannot be pushed into SQL — a row limit would truncate collections mid-parent. Hibernate therefore loads the entire result set and paginates in memory, logging `HHH000104: firstResult/maxResults specified with collection fetch; applying in memory`. On a large table this is an outage waiting to happen. ## Batch fetching Mappings stay lazy; `@BatchSize(size = n)` or the global `hibernate.default_batch_fetch_size` groups deferred loads into `where parent_id in (?, …)` queries. *Strengths.* No row multiplication — each child row is transferred once. Round trips are 1 + ceil(N/n), so a hundred parents at size 25 costs five queries total. It composes: three lazy collections mean three batched queries, not a cartesian product. It is compatible with pagination, because the IN list contains only the identifiers on the page. It is *conditional*: collections nobody touches are never loaded, so one mapping serves both the list endpoint that needs children and the one that does not. And as a global property it retroactively fixes N+1 patterns nobody has found yet. *Costs.* Still multiple round trips, so on a high-latency link a fetch join may win. Varying IN sizes create SQL variants, mitigated by parameter padding. And it only helps when many proxies are pending in the same persistence context. ## Subselect fetching `@Fetch(FetchMode.SUBSELECT)` replays the parents' query inside an `IN` subquery, initialising every parent's collection in exactly one extra statement. *Strengths.* Constant two-query profile regardless of parent count, and no row multiplication. For an export job that walks every parent and every child, it is the tightest option. *Costs.* The parent query executes a second time on the database — painful if it was expensive. It is mapping-level, so every query on that entity gets it. It needs the same persistence context. And under pagination it loads children for **all** matching parents, not the page, which is the trap that makes it unsuitable for typical screens. ## A decision rule 1. **Default the mappings**: everything lazy, plus a global batch fetch size around 20-50. This makes accidental N+1 a performance annoyance rather than an incident. 2. **Paginated list screen**: page the parents, let batching load the children. Never fetch-join a collection with a row limit. 3. **One collection always needed, unpaginated or paged by keyset over the parent alone**: `join fetch` (or an entity graph) in that specific query. Keep it to one collection; batch the rest. 4. **Multiple collections**: fetch-join at most one and rely on batching for the others, or split into separate queries against the same persistence context. 5. **Full-traversal batch job with a cheap parent query**: subselect is a reasonable, deliberate choice. 6. **Detached use** (data leaves the transaction and gets serialized): the fetch plan must be explicit — fetch join or entity graph — because nothing can be initialised later. ## Measuring rather than guessing All of this is verifiable: count statements and rows. Enable SQL logging or a statement-counting assertion in tests, and assert an upper bound on query count for the endpoint. The distinction between five queries returning 2,100 rows and one query returning 2,000 duplicated rows is only interesting once you can see both numbers; a candidate who says "I'd measure the statement count and rows fetched, then choose" is more convincing than one who recites a preference.
- What is the HHH000104 warning and why does it matter?Hibernate logs it when a query combines a collection fetch join with firstResult/maxResults: a SQL row limit would cut a parent's collection in half, so Hibernate cannot push the limit down and instead loads the entire result set and paginates in memory. The page returned is correct, which is why it hides, but heap use and latency scale with the whole matching table rather than the page. Page the parents and load the children by batching instead.
- You need two collections of the same parent on one screen. What do you do?Fetch-join at most one of them; joining two collections multiplies their sizes into a cartesian result, and for two bag-typed collections Hibernate refuses with MultipleBagFetchException. The practical approaches are to fetch-join one and let batch fetching initialise the other, or to run two queries in the same persistence context so the second query's results are merged into the already-loaded parents.
- How would you prove which strategy is better for a given endpoint?Measure statement count and rows fetched, not opinions. Turn on SQL statement counting in a test around the endpoint and assert a bound, then compare wall-clock time against a realistic dataset and a realistic network latency. The two failure shapes — too many round trips and too many duplicated rows — trade off against each other, and only measurement on your data and link tells you where the crossover lies.
One lorry that carries everything but cannot stop halfway, versus five vans that can, versus one lorry that reruns the whole route plan to work out what to load.
saying these in an interview costs you the question
- Recommending fetch join for a paginated list without mentioning in-memory pagination
- Believing batch fetching duplicates rows the way a join does
- Choosing subselect for a paginated screen
- Assuming fewer queries is always faster regardless of rows transferred
- Trying to fetch-join several collections in one query