skip to content

@BatchSize & Subselect Fetching

The global mitigations that turn N+1 into N/k+1 or 2 queries without touching every query. Interviewers ask for the trade-offs between batch IN-loading, subselect fetching, and join fetching — and where each falls down.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

questions

5

What does Hibernate's @BatchSize annotation — or the hibernate.default_batch_fetch_size setting — change about how lazy associations are loaded?

level: middleimportance: must knowfreq 56%

answer

  1. Still lazy — just grouped
  2. where owner_id in (?, ?, ?)
  3. @BatchSize on collection field or on target entity class
  4. default_batch_fetch_size is off by default
  5. Survives setMaxResults; no cartesian product

basics

~20 s

When one uninitialised proxy or lazy collection is touched, Hibernate also loads up to N other pending ones of the same type in a single query using where owner_id in (?, ?, ...). It turns N follow-up selects into roughly N/batchSize selects.

solid answer

~60 s

Batch fetching is a lazy-loading optimisation, not a change of fetch plan. Associations stay lazy; what changes is that initialising **one** proxy makes Hibernate look through the persistence context for other uninitialised proxies of the same entity or other uninitialised collections of the same role, and load up to `batchSize` of them together with an `IN` predicate. You enable it per association or entity with `@BatchSize(size = 25)`, or globally with `hibernate.default_batch_fetch_size = 25` (disabled by default). A loop over 100 parents that would otherwise emit 100 child selects emits about four. It is the least invasive of the fetch strategies: no cartesian products, no duplicate parent rows, it composes with pagination, and it degrades gracefully when only a few proxies are pending. Its cost is that it is still extra round trips, and it only helps when several proxies of the same type are pending in one session. Placement matters: on a collection field it batches that collection; on the entity class it batches proxies of that entity.

code

java · 13 lines
java
@Entity
@BatchSize(size = 25)          // batches proxies of Customer (to-one direction)
class Customer { }

@Entity
class Order {
    @ManyToOne(fetch = FetchType.LAZY)
    private Customer customer;

    @OneToMany(mappedBy = "order")
    @BatchSize(size = 25)      // batches this collection role
    private Set<OrderLine> lines;
}

go deeper

for a junior

Say that it groups the lazy loads with an IN clause so fewer queries run, and that the association is still lazy.

for a middle

Explain the persistence-context scan for pending proxies, both annotation placements, the global property being off by default, and that pagination still works.

for a senior

Add IN-list parameter padding and statement-cache effects, how to pick a size, and why batching complements rather than replaces fetch joins.

for a principal

Treat the global default as a systemic safety net for accidental N+1, and reason about round-trip cost versus over-fetching, plan cache pressure, and database parameter limits at scale.

## The problem it addresses Lazy loading defers a query until the association is touched. When code walks a list of parents and touches a lazy association on each, the deferral fires once per parent: one query for the parents, then one per row. Batch fetching does not remove the extra queries; it **groups** them. ## The mechanism When Hibernate loads entities lazily it does not immediately go to the database for the one you touched. It first scans the persistence context for other objects of the same batchable kind that are still uninitialised — other proxies of the same entity class, or other uninitialised collections of the same role (that is, the same mapped attribute of the same entity). It then issues one query for up to `batchSize` of them: ```sql select ... from order_line where order_id in (?, ?, ?, ?, ?) ``` All of those associations are initialised in one round trip; the one you touched is returned to your code and the others are simply ready when you reach them. With 100 parents and `size = 25`, you get four queries instead of 100. Crucially the association remains **lazy**. Nothing is loaded unless something touches it, so a request path that never reads the collection pays nothing. ## Where you configure it **Per collection** — annotate the collection field. This batches initialisation of that specific collection role: ```java @OneToMany(mappedBy = "order") @BatchSize(size = 25) private Set<OrderLine> lines; ``` **Per entity class** — annotate the entity. This batches initialisation of *proxies* of that entity, which is what you want for the many-to-one direction, where each parent row holds a proxy to the same target type: ```java @Entity @BatchSize(size = 25) class Customer { ... } ``` A `@BatchSize` on the `@ManyToOne` field itself is not how the to-one case is configured; the annotation belongs on the target entity class (or you rely on the global setting). **Globally** — `hibernate.default_batch_fetch_size` applies a default to every batchable association. It is off by default (no batching). Setting it to a sane value is one of the highest-value single properties in a Hibernate application, because it silently improves every accidental N+1 rather than only the ones someone remembered to annotate. A per-association `@BatchSize` overrides the global value. ## Which side of the query it affects Batch fetching works on the **loading of associations**, not on the root query. It does not affect how many rows your JPQL returns, does not duplicate roots, and does not interfere with `setMaxResults`. That last property is what makes it the natural companion to paginated queries, where a fetch join forces Hibernate to paginate in memory. ## Practical notes and edge cases **The IN list and parameter churn.** A batch of 25 does not always find 25 pending items, so Hibernate must issue queries with varying numbers of parameters. Different parameter counts mean different SQL strings, which means more entries in the statement cache and more parse work on the database. Hibernate mitigates this by padding IN-clause parameter counts to powers of two, controlled by `hibernate.query.in_clause_parameter_padding` (on by default in modern Spring-managed configurations; verify on plain Hibernate). Padding repeats the last id to fill, so results are unaffected. **It only helps when several proxies are pending.** Loading one parent with `find` and touching its collection cannot batch anything — there is one collection. Batching pays off precisely in the list-processing case. **Order of loading is not guaranteed to be your iteration order.** Hibernate picks pending items from the persistence context; the batch may include entities you have not reached yet, which is the point, but it means you cannot reason about "the first N". **It does not deduplicate work across sessions.** Batching is a persistence-context-scoped optimisation. Across requests, caching is the tool, not batching. **Choosing a size.** Values between 10 and 50 cover most cases. Too small and the round-trip count stays high; too large and you risk long IN lists, plan variance, and loading much more than the request needs. Databases also have parameter limits, so very large batch sizes are not free. ## Interview framing The crisp version: "associations stay lazy; batch fetching groups the deferred loads with an IN predicate, configured per collection, per entity class, or globally, and it is the strategy that survives pagination". Adding the padding detail and the "only helps with many pending proxies" caveat shows operational familiarity.

  • Does @BatchSize make an association eager?
    No. The association stays lazy: nothing is loaded until something touches it. Batch fetching only changes what happens at that moment — instead of loading the one association you touched, Hibernate also initialises up to N other pending ones of the same kind in a single query. A code path that never reads the association still issues no query for it.
  • Why can varying IN-clause sizes hurt the database, and what does Hibernate do about it?
    Each distinct parameter count produces a different SQL string, so the database parses and caches many near-identical statements, wasting plan cache and CPU. Hibernate pads the parameter list up to a power of two, repeating an id to fill the gap, so a batch of 13 and a batch of 15 both execute the 16-parameter statement. Repeated ids do not change the result set.
  • Where do you put @BatchSize for a @ManyToOne association?
    On the *target entity class*, because what is being batched is the initialisation of proxies of that entity, not a collection role. Annotating the @ManyToOne field is the common mistake and has no effect for that direction. Alternatively rely on `hibernate.default_batch_fetch_size`, which covers both directions everywhere at once.

A delivery driver who, when asked for one parcel on a street, brings the whole street's parcels in one trip instead of driving back for each.

saying these in an interview costs you the question

  • Saying @BatchSize makes the association eager
  • Believing batching is enabled by default in Hibernate
  • Putting @BatchSize on a @ManyToOne field and expecting proxies to be batched
  • Thinking it changes the number of rows returned by the root query
  • Setting a very large batch size on the assumption that fewer queries is always better

context

open as a page

You need to render 100 parent records, each with its child collection, from a Hibernate application. Compare loading the children with a fetch join, with batch fetching, and with subselect fetching — how do you choose?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Fetch join: one query, but duplicates parent rows and breaks database-side pagination. Batch fetching: a handful of IN-clause queries, safe with paging and partial traversal. Subselect: exactly one extra query but replays the parent query and ignores its row limit. Paginated screens use batching; unpaginated full traversal favours join or subselect.

open as a page

Hibernate's `@Fetch` annotation accepts `FetchMode.SELECT` and `FetchMode.JOIN`. What is the difference, and why does `FetchMode.JOIN` often appear to be ignored?

level: middleimportance: should knowfreq 34%

basics

~20 s

SELECT loads the association with a separate query when it is needed; JOIN loads it in the same statement with an outer join. JOIN applies only when Hibernate builds the SQL itself — find and navigation — and is ignored by JPQL, HQL and Criteria queries, which use their own explicit fetch plan.

open as a page

What does Hibernate's `@Fetch(FetchMode.SUBSELECT)` do on a collection, and what conditions must hold for it to actually take effect?

level: seniorimportance: should knowfreq 38%

basics

~20 s

When any one collection from a previously-run query's results is initialised, Hibernate loads that collection for all those parents in one query, re-using the original query as a subselect in the WHERE clause. It needs the original query still tracked in the same session, and applies to collections only.

open as a page

How would you choose a value for Hibernate's batch fetch size in a production application, and where would you apply it — globally or per association?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Set a global default around 20-50 as a safety net, then override per association where the shape justifies it. Size trades round trips against IN-list length, plan-cache variance and over-fetching; validate by measuring statement count, rows fetched and latency on realistic data.

open as a page