skip to content

Associations & Fetching

Mapping relationships between entities and controlling how Hibernate actually loads them. The richest source of ORM interview questions, because N+1 selects and LazyInitializationExceptions are what break real systems in production.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

An endpoint that returns a page of 50 root records with their child collections is issuing hundreds of SQL statements. Walk through how you would decide between a fetch join, an entity graph, batch fetching, or a projection to fix it.

level: seniorimportance: should knowfreq 45%

basics

~20 s

Measure the query count first. If you need whole entities and one collection, page root IDs then fetch the graph by ID. If you need several collections, use batch fetching or separate queries. If you only render fields, use a DTO projection and load no entities at all.

open as a page

JPA defines two query hints for entity graphs, jakarta.persistence.fetchgraph and jakarta.persistence.loadgraph. What is the semantic difference, and when does the choice actually change the SQL?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Both make the graph's attributes eager. They differ on everything else: with fetchgraph, attributes not in the graph are lazy regardless of the mapping; with loadgraph, they keep their mapped fetch type, so mapped-EAGER associations still load.

open as a page

How does @MapsId let a JPA @OneToOne share a primary key with its owner, and what problem does that solve compared with a plain @OneToOne @JoinColumn?

level: seniorimportance: should knowfreq 42%

basics

~20 s

@MapsId makes the child's primary key be the foreign key to the parent — one column doing both jobs. It removes the separate surrogate key and unique constraint, and it fixes the lazy one-to-one problem on the parent side by making the child's ID known in advance.

open as a page

You map a bidirectional one-to-one between two entities and set fetch = FetchType.LAZY on both sides, but Hibernate still issues an extra SELECT for the side declared with mappedBy. Why does that side ignore the lazy setting, and what can you do about it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The mappedBy side has no foreign key column, so Hibernate cannot tell whether a related row exists without querying. A proxy cannot represent "maybe absent", so it queries eagerly. Fixes: a shared primary key with @MapsId, optional = false plus bytecode enhancement, or drop the inverse side.

open as a page

Hibernate has a setting, hibernate.enable_lazy_load_no_trans, that lets lazy associations load even when no persistence context is open. Why do experienced teams treat turning it on as an anti-fix rather than a solution?

level: seniorimportance: should knowfreq 35%

basics

~20 s

It makes Hibernate open a temporary session and transaction per lazy access outside the original unit of work. The graph then mixes data from many points in time, every touched object costs a round trip and a connection, and the loud error that used to expose a bad fetch plan disappears.

open as a page

Integration tests routinely pass over code that issues one database query per returned row, and the defect only shows up in production. Why do the tests miss it, and how would you write one that fails when a new per-row query pattern is introduced?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Tests assert data, not statement counts; fixtures have two or three rows so the extra queries are invisible; entities are already in the persistence context from setup, so no SQL runs at all. Fix: clear the context, seed enough distinct rows, and assert a statement budget.

open as a page

Across a large JPA domain model, how do you decide which associations should carry cascade settings and orphanRemoval and which should carry none, and what breaks when that boundary is drawn in the wrong place?

level: principalimportance: should knowfreq 30%

basics

~20 s

Cascade only along composition edges inside one aggregate: root to children that have no identity of their own and are referenced by nothing outside. Across aggregate roots, reference by id with no cascade. Wrong boundaries cause shared-data deletion, unbounded graph loads, and hidden write amplification.

open as a page

A team repeatedly ships bugs where persistent objects leave the data-access layer still holding unfetched associations, and the failure only appears when the response is rendered. As a technical lead, how would you decide whether persistent entities may cross the service boundary at all, versus mandating projections?

level: principalimportance: should knowfreq 28%

basics

~20 s

Decide per read path, then enforce it. Query-side paths should return projections: no proxies, no over-fetching, explicit contracts. Command-side paths keep entities inside the transaction and never publish them. Whatever the rule, make it mechanical — architecture tests on layer boundaries, statement counts and graph assertions in tests.

open as a page

Given an entity that has already left a finished unit of work, which operations on its uninitialized lazy association are still safe, and how do you check or force initialization using Hibernate's own helper methods?

level: middleimportance: nice to knowfreq 35%

basics

~20 s

Safe: reading its identifier via Hibernate.getId or a property-access id getter, and Hibernate.isInitialized. Anything that needs the row's state is not. Hibernate.initialize forces the load but only works while a session is still open, so it is a tool for building the graph before the boundary, not for rescuing detached objects.

open as a page

What does the referencedColumnName attribute of JPA's @JoinColumn do, when would you set it, and what does it cost?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

It names which column of the target table the foreign key points at, instead of the default primary key. Use it for legacy schemas keyed on a natural unique column. Cost: the target column needs a unique constraint, and the association no longer resolves through the primary key.

open as a page

Hibernate can use build-time bytecode enhancement instead of runtime proxy subclasses for lazy loading. What does enhancement change about how an entity holds unloaded state, and what does it make possible that proxies cannot?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Enhancement rewrites entity classes at build time so field reads are intercepted inside the class itself, instead of relying on a generated subclass. That enables lazy individual attributes, lazy groups, laziness on the inverse side of a one-to-one, and in-line dirty tracking.

open as a page

When modelling a JPA domain, how do you decide whether a relationship should be mapped in both directions or only one, and what does the extra direction cost you?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Map the @ManyToOne always — it matches the foreign key. Add the inverse collection only when code genuinely navigates parent to children and the collection is bounded. The extra direction costs a consistency invariant, cascade and loading surface, and unbounded collections; the schema is identical either way.

open as a page

How would you choose a value for Hibernate's batch fetch size in a production application, and where would you apply it — globally or per association?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Set a global default around 20-50 as a safety net, then override per association where the shape justifies it. Size trades round trips against IN-list length, plan-cache variance and over-fetching; validate by measuring statement count, rows fetched and latency on realistic data.

open as a page

For a large Hibernate domain model, how would you set a house standard for mapping to-many associations — Set everywhere, List/bag, or indexed lists — and what systemic costs follow from each choice?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Default to Set for owned collections: targeted DML, no delete-all-reinsert, and multiple collections can be fetched in one query. Use indexed lists only where user-defined order is domain data. The price of Set is disciplined equals/hashCode; the price of bags is fetching and write constraints.

open as a page

You inherit a large service where profiling shows dozens of read paths issue one database query per returned row. You cannot fix them all this quarter. How do you decide which ones to address, and what do you put in place so the pattern stops reappearing?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Rank by impact: rows returned x call rate x round-trip latency, weighted by user-facing criticality and connection-pool pressure. Fix the top few properly, then prevent recurrence with lazy-by-default mappings, use-case fetch plans or projections, and statement budgets enforced in CI.

open as a page

showing 31–45 of 45