skip to content

What goes wrong when an @EntityGraph eagerly fetches collections combined with pagination or multiple collections, and how do you handle it?

level: seniorimportance: must knowfreq 50%

answer

  1. two List bags -> MultipleBagFetchException
  2. Set avoids throw but still cartesian
  3. HHH000104 pagination applied in memory
  4. fix: page ids, then IN-fetch
  5. @BatchSize / default_batch_fetch_size alternative

basics

~20 s

Fetching two List collections at once throws MultipleBagFetchException. Paginating (Pageable) while join-fetching a collection makes Hibernate load all rows and page in memory (HHH000104 warning). Fix: fetch one collection, use Set for bags, or split into a two-step id-then-fetch query.

solid answer

~50 s

Two classic failures. First, joining **two `List`-typed collections** in one graph throws `org.hibernate.loader.MultipleBagFetchException` ("cannot simultaneously fetch multiple bags") because a bag join produces an ambiguous cartesian product Hibernate can't reconstruct. Second, combining a **collection fetch with pagination** (`Pageable`) can't be done correctly in SQL — the join multiplies rows per parent, so `LIMIT` would cut mid-parent. Hibernate falls back to fetching **all** matching rows and paginating **in memory**, logging `HHH000104: firstResult/maxResults specified with collection fetch; applying in memory`, which is a heap and performance risk on large sets. Remedies: fetch at most one collection per graph; change bag collections to `Set` (or add `@OrderColumn`) to sidestep MultipleBagFetchException; for pagination, do a **two-query** approach — first page the entity IDs without the fetch, then a second `@EntityGraph`/`JOIN FETCH` query `WHERE id IN (:ids)`; or use `@BatchSize`/`hibernate.default_batch_fetch_size` so collections load in a few batched IN queries instead of a join.

code

java · 14 lines
java
public interface OrderRepository extends JpaRepository<Order, Long> {

    // Step 1: paginate IDs only (correct SQL LIMIT, no collection join)
    @Query("select o.id from Order o where o.status = :status")
    Page<Long> findIdsByStatus(@Param("status") Status status, Pageable pageable);

    // Step 2: fetch the collection for just this page's ids
    @EntityGraph(attributePaths = "items")
    @Query("select o from Order o where o.id in :ids")
    List<Order> findWithItemsByIds(@Param("ids") List<Long> ids);
}

// Alternative to avoid MultipleBagFetchException on Order.items:
//   private Set<OrderItem> items;   // Set join is permitted; List (bag) x List throws

go deeper

for a junior

Aware that fetching collections has limits; know the term N+1.

for a middle

Know MultipleBagFetchException and that Set helps; recognize the pagination warning.

for a senior

Explain the cartesian-product root cause, HHH000104 in-memory pagination, and the two-step id-then-fetch fix.

for a principal

Weigh two-step vs @BatchSize vs projections at scale; consider heap risk, count queries, and duplicate-root semantics as an architectural fetch policy.

### Background: bags vs sets In Hibernate a `List` without an `@OrderColumn` is a **bag** — an unordered collection that permits duplicates and has no index column. A `Set` is a set. This distinction drives the failures below. ### Failure 1 — MultipleBagFetchException If a single fetch plan eagerly joins **two** bag collections (e.g. `Order.items` and `Order.shipments`, both `List`), Hibernate throws: ``` org.hibernate.loader.MultipleBagFetchException: cannot simultaneously fetch multiple bags ``` Why: joining two independent one-to-many collections yields a **cartesian product** (rows = items x shipments). For a bag, Hibernate cannot de-duplicate/reconstruct the two collections unambiguously, so it refuses. **Fixes:** - Change one or both collections from `List` to `Set` — a set join is allowed (Hibernate can de-dup by identity). But beware: joining two collections as sets still creates a **cartesian product** in the SQL result, so it can be very wide/slow even when it doesn't throw. - Or fetch only **one** collection per query and load the other via a separate query / `@BatchSize`. - Or add `@OrderColumn` to make the `List` an indexed list (also avoids the bag problem, at the cost of maintaining an order column). ### Failure 2 — pagination with a collection fetch With `Pageable` (or `firstResult`/`maxResults`) plus a `JOIN FETCH`/`@EntityGraph` on a **collection**, the joined result has multiple rows per root entity. A SQL `LIMIT` would truncate rows in the middle of a parent's collection, giving wrong data. Hibernate avoids the wrong answer by **loading every matching row and applying the offset/limit in memory (in the Java heap)**, logging: ``` HHH000104: firstResult/maxResults specified with collection fetch; applying in memory ``` On a large table this can pull the whole result set into memory — an OutOfMemory / latency hazard. (Note: fetching a to-**one** association like `@ManyToOne` with pagination is fine — it doesn't multiply rows.) **Fixes:** - **Two-step (id + fetch)**: page the IDs first (`Page<Long>` or a projection, no collection fetch), then run a second query `@EntityGraph(attributePaths="items") ... where e.id in :ids`. The first query paginates correctly in SQL; the second fetches collections for just that page. - **Batch fetching**: keep the collection LAZY and set `@BatchSize(size = N)` on it (or `hibernate.default_batch_fetch_size`). Pagination works normally; the N+1 becomes ceil(N/size) batched `IN` queries — often the simplest good-enough answer. - **DTO/projection queries** that select only needed scalars avoid the whole collection-join issue. ### Additional gotchas - **DISTINCT and duplicates**: a collection join returns duplicate parent rows (one per child). With plain JPQL `JOIN FETCH` you'd add `DISTINCT` (and historically `hibernate.query.passDistinctThrough=false`) to de-dup roots in the returned list; with a `Set` collection Hibernate de-dups automatically. `@EntityGraph` on a Spring `List<T>` return can therefore surface duplicate roots for `List` collections — be aware when counting. - **Only one 'to-many' at a time** is the safe rule of thumb; you can freely combine several to-**one** fetches with one to-**many**. - **countQuery**: for `Page` results Spring also runs a count query; ensure the graph/joins don't break it (Spring generally derives a separate count query, but custom `@Query` may need `countQuery`). ### When to use what - One collection, no paging -> `@EntityGraph` is perfect. - Paging + collection -> two-step id query or `@BatchSize`; never rely on in-memory pagination for large sets. - Multiple collections -> fetch one, batch the rest; don't stack bag joins.

  • Why is fetching a @ManyToOne with Pageable fine, but a @OneToMany collection is not?
    A to-one join returns exactly one row per root, so SQL LIMIT paginates roots correctly. A collection join returns many rows per root, so LIMIT would slice a parent's children — Hibernate must paginate in memory instead.
  • You changed both List collections to Set to stop MultipleBagFetchException. Any remaining concern?
    Yes — two collection joins still form a cartesian product (items x shipments rows). The query no longer throws but can be enormous. Prefer fetching one collection and batching the other.

saying these in an interview costs you the question

  • Believing you can eagerly join multiple List collections in one @EntityGraph
  • Thinking pagination + collection fetch paginates in SQL (it silently goes in-memory)
  • Assuming @EntityGraph never returns duplicate root rows for List collections

context