skip to content

JOIN FETCH & Entity Graphs

The per-query cures for N+1: fetching associations eagerly in JPQL or declaring fetch plans as entity graphs. Interviewers expect you to know the fetchgraph-vs-loadgraph distinction and when a fetch join is the wrong tool.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

questions

4

In JPQL, what does the JOIN FETCH clause do, and how does it differ from writing a plain JOIN in the same query?

level: juniorimportance: must knowfreq 78%

answer

  1. JOIN = where clause; JOIN FETCH = result graph
  2. LEFT JOIN FETCH keeps childless roots
  3. Collection fetch multiplies root rows
  4. firstResult/maxResults + collection fetch = in-memory paging
  5. Never filter on a fetched collection alias

basics

~20 s

JOIN FETCH tells the ORM to load the joined association's rows into the returned entities in the same SQL statement. A plain JOIN only lets you filter or navigate in the query; the association still loads later on access.

solid answer

~50 s

`JOIN FETCH` is a JPQL instruction that both joins the association in SQL **and** initializes it on the returned entities. A plain `JOIN` (or `LEFT JOIN`) only contributes to the SQL join for filtering, ordering or projection — the association on the returned entity keeps whatever fetch plan the mapping declares, so a lazy one is still a proxy or uninitialized collection and triggers extra queries when touched. So `select o from Order o join o.items i where i.sku = :sku` gives you orders and one query per order later if you read `o.getItems()`. `select o from Order o join fetch o.items` gives you orders with items already populated in one round trip. Two practical notes: a fetched collection cannot be given an alias you then filter on without changing the collection's contents (you'd silently load a partial collection), and fetching a collection duplicates the root row per child, so you get repeated roots unless you de-duplicate.

code

java · 9 lines
java
List<Order> a = em.createQuery(
    "select o from Order o join o.items i where i.sku = :sku", Order.class)
  .setParameter("sku", sku).getResultList();
// a.get(0).getItems() -> extra SELECT per order

List<Order> b = em.createQuery(
    "select o from Order o left join fetch o.items where o.id in :ids", Order.class)
  .setParameter("ids", ids).getResultList();
// b.get(0).getItems() -> already initialized

go deeper

for a junior

Say what it does in one sentence: JOIN FETCH loads the association in the same query, plain JOIN only joins for filtering. Give the orders/items example.

for a middle

Add LEFT vs inner semantics, the duplicate-root effect and DISTINCT, and that the mapping stays lazy while the query overrides the plan.

for a senior

Lead with when you reach for it versus batch fetching, name the pagination trap and the HHH000104 in-memory paging behaviour, and explain why filtering a fetched alias corrupts state.

for a principal

Frame it as a fetch-plan strategy: mappings stay lazy, fetch plans belong to use cases, and choose between fetch joins, entity graphs, batch size and projections based on cardinality and query shape.

## The problem JOIN FETCH solves An ORM maps object graphs onto tables. When you load an `Order`, its `items` collection is usually declared lazy, meaning the ORM hands you a placeholder and only issues SQL for the children when you actually touch them. That is fine for one order and terrible for a hundred: one query for the orders plus one per order for its items — the classic per-row query storm. `JOIN FETCH` is the JPQL clause that lets a *specific query* override the mapping's fetch plan and pull the association eagerly, in the same SQL statement. ## Plain JOIN vs JOIN FETCH Both produce a SQL join. The difference is what the ORM does with the joined columns. - **Plain `JOIN o.items i`** — the ORM adds the join so you can write `where i.sku = ...` or `order by i.price`. It does **not** select the child columns into the persistence context. The `items` collection on each returned `Order` remains uninitialized. - **`JOIN FETCH o.items`** — the ORM adds the join *and* selects the child columns, hydrates `OrderItem` entities, and wires them into each `Order`'s collection. After the query the collection is initialized; touching it issues no further SQL. A useful mental rule: plain JOIN is about the **where clause**; JOIN FETCH is about the **result graph**. `LEFT JOIN FETCH` matters when the association may be empty or null. An inner `JOIN FETCH` on a collection silently drops roots that have no children — a common source of "rows disappeared" bugs. ## Duplicate roots and DISTINCT Joining a to-many association multiplies rows: an order with three items produces three result rows. Without intervention the query returns the same `Order` reference three times (the same object, because the persistence context guarantees one instance per identity within a session, but three entries in the list). The fix is a de-duplication step. Historically people wrote `select distinct o from Order o join fetch o.items`, and Hibernate translated that into two things at once: a SQL `DISTINCT` keyword and an in-memory pass that collapses repeated roots. The SQL `DISTINCT` is useless work for the database on a query whose rows differ only by child columns. Modern Hibernate (6.x) performs the in-memory de-duplication of fetched-collection roots **automatically**, so `distinct` is no longer required for that purpose; in Hibernate 5.2+ the hint `hibernate.query.passDistinctThrough=false` was the way to keep the Java-side de-dup while suppressing the SQL keyword. Knowing this history is a good senior signal. ## Pagination is the sharp edge If you combine a collection `JOIN FETCH` with `setFirstResult`/`setMaxResults`, the row limit applies to the *joined* rows, not to root entities — so a limit of 10 might return three orders with partial-looking item sets. Hibernate detects this, refuses to trust the database limit, and instead fetches **all** matching rows and paginates in memory, logging a warning (historically `HHH000104: firstResult/maxResults specified with collection fetch; applying in memory`). On a large table that is an out-of-memory waiting to happen. The standard workaround is a two-step approach: page the root IDs with a plain query, then fetch the graph with `where o.id in :ids`. Hibernate 6 can also use window functions for some of these cases. ## Filtering a fetched collection Aliasing a fetched association and filtering on it is dangerous: ``` select o from Order o join fetch o.items i where i.status = 'SHIPPED' ``` The returned `Order.items` collections now contain only shipped items, yet those entities live in the persistence context and can be flushed. With `orphanRemoval` or a cascading delete you can lose data; at minimum you have an object that lies about the database. The JPA spec explicitly leaves this undefined. Fetch the whole collection and filter in memory, or query the child entity directly and project what you need. ## Multiple collections You can `JOIN FETCH` several to-one associations freely — they don't multiply rows. Fetching **two** `List` collections in one query throws `MultipleBagFetchException`, because the ORM cannot tell which cartesian row belongs to which list position. (That specific exception is a separate topic; the takeaway here is that two collection fetches in one query is a design smell — split into two queries and let the persistence context join them in memory.) ## When to reach for it Use `JOIN FETCH` when *this particular use case* needs the association: a detail screen, an export, a message payload. Do not fix per-query fetch problems by switching the mapping to `EAGER` — that penalises every other query in the application. `JOIN FETCH` is the per-query tool; the mapping should stay lazy.

  • Why can combining JOIN FETCH on a collection with setMaxResults be dangerous?
    The SQL LIMIT would apply to joined rows, not to root entities, so it would return roots with partially loaded collections. Hibernate refuses to do that: it drops the database limit, fetches every matching row and paginates in memory, warning HHH000104. On a large result set this loads the whole table into heap. The fix is to page root IDs in one query, then fetch the graph with `where id in :ids`.
  • Why is filtering on an aliased fetched collection considered unsafe?
    The returned collection then holds only the rows that passed the filter, but it is a managed collection inside the persistence context. Hibernate believes it represents the full database state, so a flush combined with orphanRemoval or a cascading delete can remove the rows you filtered out. JPA leaves this behaviour undefined; fetch the full collection and filter in memory, or query the child entity directly.

A plain JOIN is looking someone up in the phone book to check their address; JOIN FETCH is bringing them home with you.

saying these in an interview costs you the question

  • Believing a plain JOIN initializes the association ("I joined it, so it's loaded")
  • Fixing an N+1 by switching the mapping to FetchType.EAGER instead of using a per-query fetch
  • Assuming `select distinct` is still required in Hibernate 6, or that it only ever affected SQL
  • Using an inner JOIN FETCH on an optional collection and losing roots that have no children
  • Combining a collection JOIN FETCH with pagination and expecting the database to apply the limit

context

open as a page

What is a JPA entity graph, and how do you define and apply one — both declaratively with @NamedEntityGraph and programmatically with EntityManager.createEntityGraph?

level: middleimportance: should knowfreq 58%

basics

~20 s

An entity graph is a reusable, declarative description of which attributes to load with a query. Define it with @NamedEntityGraph on the entity or build it via em.createEntityGraph(Class), then pass it to a query or find() as a hint.

open as a page

An endpoint that returns a page of 50 root records with their child collections is issuing hundreds of SQL statements. Walk through how you would decide between a fetch join, an entity graph, batch fetching, or a projection to fix it.

level: seniorimportance: should knowfreq 45%

basics

~20 s

Measure the query count first. If you need whole entities and one collection, page root IDs then fetch the graph by ID. If you need several collections, use batch fetching or separate queries. If you only render fields, use a DTO projection and load no entities at all.

open as a page

JPA defines two query hints for entity graphs, jakarta.persistence.fetchgraph and jakarta.persistence.loadgraph. What is the semantic difference, and when does the choice actually change the SQL?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Both make the graph's attributes eager. They differ on everything else: with fetchgraph, attributes not in the graph are lazy regardless of the mapping; with loadgraph, they keep their mapped fetch type, so mapped-EAGER associations still load.

open as a page