Associations & Fetching
Mapping relationships between entities and controlling how Hibernate actually loads them. The richest source of ORM interview questions, because N+1 selects and LazyInitializationExceptions are what break real systems in production.
part ofHibernateoverview, primer and where to startread it →on this pageshowhide
explore
- Association Kinds & Owning Side5 questions
- Join Columns vs Join Tables5 questions
- Cascade Types & orphanRemoval6 questions
- LAZY vs EAGER Defaults & Proxies5 questions
- The N+1 Problem5 questions
- JOIN FETCH & Entity Graphs4 questions
- @BatchSize & Subselect Fetching5 questions
- LazyInitializationException5 questions
questions
page 2 of 2An endpoint that returns a page of 50 root records with their child collections is issuing hundreds of SQL statements. Walk through how you would decide between a fetch join, an entity graph, batch fetching, or a projection to fix it.
basics
~20 sMeasure the query count first. If you need whole entities and one collection, page root IDs then fetch the graph by ID. If you need several collections, use batch fetching or separate queries. If you only render fields, use a DTO projection and load no entities at all.
JPA defines two query hints for entity graphs, jakarta.persistence.fetchgraph and jakarta.persistence.loadgraph. What is the semantic difference, and when does the choice actually change the SQL?
basics
~20 sBoth make the graph's attributes eager. They differ on everything else: with fetchgraph, attributes not in the graph are lazy regardless of the mapping; with loadgraph, they keep their mapped fetch type, so mapped-EAGER associations still load.
You map a bidirectional one-to-one between two entities and set fetch = FetchType.LAZY on both sides, but Hibernate still issues an extra SELECT for the side declared with mappedBy. Why does that side ignore the lazy setting, and what can you do about it?
basics
~20 sThe mappedBy side has no foreign key column, so Hibernate cannot tell whether a related row exists without querying. A proxy cannot represent "maybe absent", so it queries eagerly. Fixes: a shared primary key with @MapsId, optional = false plus bytecode enhancement, or drop the inverse side.
Hibernate has a setting, hibernate.enable_lazy_load_no_trans, that lets lazy associations load even when no persistence context is open. Why do experienced teams treat turning it on as an anti-fix rather than a solution?
basics
~20 sIt makes Hibernate open a temporary session and transaction per lazy access outside the original unit of work. The graph then mixes data from many points in time, every touched object costs a round trip and a connection, and the loud error that used to expose a bad fetch plan disappears.
Integration tests routinely pass over code that issues one database query per returned row, and the defect only shows up in production. Why do the tests miss it, and how would you write one that fails when a new per-row query pattern is introduced?
basics
~20 sTests assert data, not statement counts; fixtures have two or three rows so the extra queries are invisible; entities are already in the persistence context from setup, so no SQL runs at all. Fix: clear the context, seed enough distinct rows, and assert a statement budget.
Across a large JPA domain model, how do you decide which associations should carry cascade settings and orphanRemoval and which should carry none, and what breaks when that boundary is drawn in the wrong place?
basics
~20 sCascade only along composition edges inside one aggregate: root to children that have no identity of their own and are referenced by nothing outside. Across aggregate roots, reference by id with no cascade. Wrong boundaries cause shared-data deletion, unbounded graph loads, and hidden write amplification.
A team repeatedly ships bugs where persistent objects leave the data-access layer still holding unfetched associations, and the failure only appears when the response is rendered. As a technical lead, how would you decide whether persistent entities may cross the service boundary at all, versus mandating projections?
basics
~20 sDecide per read path, then enforce it. Query-side paths should return projections: no proxies, no over-fetching, explicit contracts. Command-side paths keep entities inside the transaction and never publish them. Whatever the rule, make it mechanical — architecture tests on layer boundaries, statement counts and graph assertions in tests.
Given an entity that has already left a finished unit of work, which operations on its uninitialized lazy association are still safe, and how do you check or force initialization using Hibernate's own helper methods?
basics
~20 sSafe: reading its identifier via Hibernate.getId or a property-access id getter, and Hibernate.isInitialized. Anything that needs the row's state is not. Hibernate.initialize forces the load but only works while a session is still open, so it is a tool for building the graph before the boundary, not for rescuing detached objects.
What does the referencedColumnName attribute of JPA's @JoinColumn do, when would you set it, and what does it cost?
basics
~20 sIt names which column of the target table the foreign key points at, instead of the default primary key. Use it for legacy schemas keyed on a natural unique column. Cost: the target column needs a unique constraint, and the association no longer resolves through the primary key.
Hibernate can use build-time bytecode enhancement instead of runtime proxy subclasses for lazy loading. What does enhancement change about how an entity holds unloaded state, and what does it make possible that proxies cannot?
basics
~20 sEnhancement rewrites entity classes at build time so field reads are intercepted inside the class itself, instead of relying on a generated subclass. That enables lazy individual attributes, lazy groups, laziness on the inverse side of a one-to-one, and in-line dirty tracking.
When modelling a JPA domain, how do you decide whether a relationship should be mapped in both directions or only one, and what does the extra direction cost you?
basics
~20 sMap the @ManyToOne always — it matches the foreign key. Add the inverse collection only when code genuinely navigates parent to children and the collection is bounded. The extra direction costs a consistency invariant, cascade and loading surface, and unbounded collections; the schema is identical either way.
How would you choose a value for Hibernate's batch fetch size in a production application, and where would you apply it — globally or per association?
basics
~20 sSet a global default around 20-50 as a safety net, then override per association where the shape justifies it. Size trades round trips against IN-list length, plan-cache variance and over-fetching; validate by measuring statement count, rows fetched and latency on realistic data.
For a large Hibernate domain model, how would you set a house standard for mapping to-many associations — Set everywhere, List/bag, or indexed lists — and what systemic costs follow from each choice?
basics
~20 sDefault to Set for owned collections: targeted DML, no delete-all-reinsert, and multiple collections can be fetched in one query. Use indexed lists only where user-defined order is domain data. The price of Set is disciplined equals/hashCode; the price of bags is fetching and write constraints.
You inherit a large service where profiling shows dozens of read paths issue one database query per returned row. You cannot fix them all this quarter. How do you decide which ones to address, and what do you put in place so the pattern stops reappearing?
basics
~20 sRank by impact: rows returned x call rate x round-trip latency, weighted by user-facing criticality and connection-pool pressure. Fix the top few properly, then prevent recurrence with lazy-by-default mappings, use-case fetch plans or projections, and statement budgets enforced in CI.
showing 31–45 of 45