skip to content

A JPQL query fetch-joins two different List-mapped collections of the same root entity and Hibernate fails with MultipleBagFetchException ("cannot simultaneously fetch multiple bags"). Why is that rejected, and what are the ways to get the data you wanted?

level: seniorimportance: must knowfreq 48%

answer

  1. two collection joins = M×N cartesian product
  2. bag duplicates are legal → cannot deduplicate
  3. thrown at query build, before any SQL
  4. fix: Set, or one query per collection, same context
  5. DISTINCT does not fix it

basics

~20 s

Joining two bags produces a cartesian product, and because bags allow duplicates and carry no index, Hibernate cannot tell real duplicates from join-generated ones. Fix it by mapping the collections as Sets, or by running one query per collection in the same persistence context.

solid answer

~60 s

Fetch-joining one collection multiplies each root row by the number of children. Fetch-joining a second multiplies again — an M×N cartesian product. For a `Set`, Hibernate can safely collapse the repeats; for a **bag** it cannot, because duplicates are legal collection contents and there is no index to distinguish "this element appears twice in the collection" from "this row appeared twice because of the join". Rather than silently produce a corrupted collection, Hibernate rejects the query up front with `MultipleBagFetchException`. Options: 1. **Change the collections to `Set`** — the usual fix; the query then works, though it still fetches an M×N result set. 2. **Run separate queries in the same persistence context** — fetch roots with collection A, then a second query fetching collection B for the same roots. The already-managed roots get their second collection populated, with no cartesian product. 3. **Add `@OrderColumn`**, turning the bags into indexed lists, which removes the exception — but keeps the cartesian product. Even with Sets, two collection fetch joins in one query is usually the wrong shape at scale; sequential queries are safer.

code

java · 13 lines
java
// throws MultipleBagFetchException when comments and tags are both List (bags)
em.createQuery("select p from Post p " +
               "join fetch p.comments join fetch p.tags", Post.class);

// works with bags unchanged: two queries, one persistence context
List<Post> posts = em.createQuery(
        "select distinct p from Post p join fetch p.comments where p.id in :ids",
        Post.class).setParameter("ids", ids).getResultList();

em.createQuery(
        "select distinct p from Post p join fetch p.tags where p.id in :ids",
        Post.class).setParameter("ids", ids).getResultList();
// posts now have both collections initialised

go deeper

for a junior

Know the cause in one sentence — two bags fetched together create ambiguous duplicates — and that mapping them as Sets removes the error.

for a middle

Explain the cartesian product and why set semantics can deduplicate while bag semantics cannot.

for a senior

Push past legality to cost: prefer sequential queries in one persistence context, and quantify M+N versus M×N rows.

for a principal

Treat it as a fetching-strategy decision across the model — collection loading strategy, result-set size budgets, and the conventions that keep multi-collection fetches out of hot paths.

## What a collection fetch join does to the result set ```sql select p.*, c.* from post p join comment c on c.post_id = p.id ``` A post with 5 comments produces 5 rows. Hibernate reads them and assembles one `Post` whose `comments` collection has 5 elements — it knows to expect repeats of the root. Add a second collection: ```sql select p.*, c.*, t.* from post p join comment c on c.post_id = p.id join tag t on t.post_id = p.id ``` Now a post with 5 comments and 4 tags produces **20** rows: each comment paired with each tag. That is a cartesian product between two independent child sets, and it is both a correctness problem and a size problem. ## Why sets survive it and bags do not Hibernate assembles collections from the flattened rows. For each root it must decide what belongs in each collection. - **Set semantics** — duplicates are impossible by definition, so adding the same child 4 times is harmless: the set holds it once. The result is correct even though the SQL returned it repeatedly. - **Indexed list** — each element carries an index, so a repeated row lands at the same position; the list is rebuilt correctly. - **Bag semantics** — duplicates are *legal contents*. If Hibernate sees the same comment 4 times, it has no basis to decide whether the collection genuinely contains it 4 times or whether the join produced the repeats. Deduplicating would corrupt legitimate bags; not deduplicating would corrupt this one. Because there is no correct behaviour available, Hibernate refuses the query at build time: `org.hibernate.loader.MultipleBagFetchException: cannot simultaneously fetch multiple bags`. This is a fail-fast, not a limitation of SQL — and note that it fires when the query is *built*, before any database round trip, so it surfaces immediately and deterministically. ## Fix 1 — map the collections as Sets ```java @OneToMany(mappedBy = "post") private Set<Comment> comments = new HashSet<>(); @OneToMany(mappedBy = "post") private Set<Tag> tags = new HashSet<>(); ``` The exception disappears, and the collections are assembled correctly. Two caveats: - The **cartesian product is still transferred**. 5 comments × 4 tags is 20 rows; 100 × 100 is 10 000 rows for one post. Hibernate deduplicates in memory *after* the database has produced and shipped every row. On a wide entity that is real network and heap cost. - Entity elements in a `HashSet` need sane `equals`/`hashCode`; a hash derived from a generated id is unstable. ## Fix 2 — sequential queries in one persistence context (usually best) ```java List<Post> posts = em.createQuery( "select distinct p from Post p join fetch p.comments where p.id in :ids", Post.class) .setParameter("ids", ids).getResultList(); em.createQuery( "select distinct p from Post p join fetch p.tags where p.id in :ids", Post.class) .setParameter("ids", ids).getResultList(); ``` The second query returns the *same managed instances* — the persistence context guarantees one instance per identifier per context — and initialises their `tags` collection as a side effect. You can ignore its return value entirely. Row count is M + N instead of M × N. This works with bags unchanged, and it scales: two 100-element collections cost 200 rows rather than 10 000. ## Fix 3 — @OrderColumn Adding `@OrderColumn` makes each collection an indexed list, so the exception no longer applies. But it changes the schema, imposes index maintenance on writes, and leaves the cartesian product in place. Choose it when the order genuinely belongs in the domain, not as a workaround for the exception. ## What does *not* fix it - **`SELECT DISTINCT`** — it deduplicates root references in the query result list, not the contents of the fetched collections, and it does not stop the exception being thrown at build time. - **Making one collection eager** — `FetchType.EAGER` on two bags reproduces exactly the same multi-bag fetch and the same exception, only now on every load of the entity rather than in one query you control. - **Wrapping in a DTO projection with the same double join** — that avoids the exception (no collections are being assembled) but you must reassemble the graph yourself, and the cartesian product remains. ## The takeaway The exception is a symptom, not the disease. The disease is fetching two independent collections in one SQL statement. Sets make the query legal; separate queries make it *right*. Batched or subselect collection loading is the other standard route to the same shape.

  • Why does running two separate fetch queries populate both collections on the same objects?
    Because a persistence context guarantees a single managed instance per entity identifier. The second query returns the roots that are already managed from the first query — Hibernate does not build new objects — and initialising the second fetched collection attaches it to those same instances. The second query's return value can therefore be discarded; the effect is on the objects you already hold.
  • If you change both collections to Set, is the query now efficient?
    It is now legal and correct, but not necessarily efficient. The database still computes and returns the full M×N cartesian product; Hibernate deduplicates only after every row has been fetched. With two collections of a hundred elements that is ten thousand rows carrying the root's columns repeatedly. For anything but small collections, sequential queries — or batched/subselect collection loading — transfer far less.
  • Does adding SELECT DISTINCT avoid MultipleBagFetchException?
    No. DISTINCT affects duplicate root references in the query's result list, and in modern Hibernate the in-memory deduplication of roots happens anyway for fetch joins. MultipleBagFetchException is raised while the query is being built, from the mapping semantics of the two collections, before any SQL is generated or executed.

Joining two child tables is like pairing every guest with every dish on a menu: you get one line per pair. A set can throw the repeats away; a bag has no way to know whether a repeated line means the guest really took two helpings.

saying these in an interview costs you the question

  • Claiming SELECT DISTINCT solves MultipleBagFetchException
  • Switching one collection to EAGER as the fix — it triggers the same failure on every load
  • Believing the exception comes from the database or a SQL limitation
  • Treating a Set change as a performance fix rather than a legality fix — the cartesian product remains
  • Not knowing that two queries in the same persistence context populate the same managed instances

context