skip to content

Associations & Fetching

Mapping relationships between entities and controlling how Hibernate actually loads them. The richest source of ORM interview questions, because N+1 selects and LazyInitializationExceptions are what break real systems in production.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

JPA offers four association mappings — @ManyToOne, @OneToMany, @OneToOne and @ManyToMany. Explain what each expresses, and how you would choose between them for a model of orders, customers and tags.

level: juniorimportance: must knowfreq 76%

answer

  1. Read the name outward from the annotated field
  2. @ManyToOne = the real FK column
  3. mappedBy side is a view, not a second relationship
  4. @ManyToMany dies when the link gains attributes
  5. Direction is a Java choice; cardinality is a data fact

basics

~20 s

They declare cardinality from the annotated field outward. @ManyToOne: many rows reference one (Order to Customer) and holds the foreign key. @OneToMany: the inverse collection (Customer to Orders). @OneToOne: at most one each way. @ManyToMany: many both ways, via a link table.

solid answer

~50 s

The annotation names the cardinality **from the annotated field outward**. `@ManyToOne` on `Order.customer` means many orders reference one customer; this side maps directly onto the foreign-key column, so it is the cheapest and most fundamental mapping. `@OneToMany` on `Customer.orders` is the collection view of that same relationship; in a bidirectional pair it is declared `mappedBy = "customer"` and generates no SQL of its own. `@OneToOne` means at most one row on each side (`Order` to `Invoice`); one table carries the FK, typically with a unique constraint, or the two share a primary key via `@MapsId`. `@ManyToMany` means both sides may reference many of the other (`Order` to `Tag`) and needs a separate link table. Practical rule: always map the `@ManyToOne` — it is the one the schema actually has. Add the `@OneToMany` inverse only if code really navigates parent to children, and use `@ManyToMany` only while the link itself carries no data.

code

java · 25 lines
java
@Entity
class Order {
    @Id @GeneratedValue Long id;

    @ManyToOne
    @JoinColumn(name = "customer_id")
    Customer customer;            // many orders -> one customer (owns the FK)

    @OneToOne(mappedBy = "order")
    Invoice invoice;              // invoice table holds the FK

    @ManyToMany
    @JoinTable(name = "order_tag",
        joinColumns = @JoinColumn(name = "order_id"),
        inverseJoinColumns = @JoinColumn(name = "tag_id"))
    Set<Tag> tags;
}

@Entity
class Customer {
    @Id @GeneratedValue Long id;

    @OneToMany(mappedBy = "customer")
    Set<Order> orders;            // inverse view of Order.customer
}

go deeper

for a junior

Name the four annotations, say which side has the foreign key, and give one concrete example of each from an everyday domain.

for a middle

Add that a bidirectional pair maps one column, that mappedBy marks the non-owning view, and that direction is independent of cardinality.

for a senior

Argue from the schema outward, flag @ManyToMany as a temporary shape that dies when the link gains attributes, and treat unneeded collections as a cost.

for a principal

Frame associations as aggregate-boundary decisions: which navigations the domain guarantees, which links should be entities in their own right, and where an association should be replaced by an id reference or a query.

## What an association mapping is A relational database expresses relationships with foreign-key columns and, where needed, link tables. An object model expresses them with references and collections. A JPA association annotation is the declaration that tells the ORM how one maps onto the other: which field corresponds to which column, and how many rows may sit on each end. The annotation name is read **from the annotated field outward**: the first word is the multiplicity of the side you are standing on, the second the multiplicity of the side you are pointing at. `@ManyToOne List<...>` is therefore nonsense, and `@OneToMany Customer` equally so — the plural half of the name must line up with the collection. ## The four kinds **@ManyToOne** — many entities of this type reference one of the target. `Order.customer` is the archetype. This is the only association that maps one-for-one onto a plain foreign-key column on the annotated entity's own table, which makes it the most important of the four: if you map nothing else, map this. **@OneToMany** — the reverse view: one entity holds a collection of many. `Customer.orders` is the archetype. Crucially, in a bidirectional pair the `@OneToMany` describes the *same* database column as the matching `@ManyToOne`; it is not an extra relationship. It is declared with `mappedBy` naming the field on the other side that owns the column. **@OneToOne** — at most one row on each side: `Order` to `Invoice`, `User` to `UserProfile`. Relationally this is a foreign key on one of the two tables plus a unique constraint, or a shared primary key where the dependent table's PK is also an FK to the parent (expressed with `@MapsId`). The side holding the column is the owner; the other side uses `mappedBy`. **@ManyToMany** — both ends may reference many of the other: orders and tags, users and roles. There is no column that can hold this, so a third table of `(order_id, tag_id)` pairs is required. One side is the owner and the other, if mapped at all, uses `mappedBy`. ## Direction is separate from cardinality Cardinality is a property of the data. **Direction** is a property of your Java model: whether you mapped one side, or both. `@ManyToOne` alone gives a unidirectional many-to-one. Adding a `mappedBy` `@OneToMany` makes the same relationship bidirectional — the schema does not change at all. Direction is therefore free to choose per use case, while cardinality is dictated by the domain. ## Choosing in practice Start from the schema, not from the prose of the requirement. Ask: "which table would hold the foreign key?" That table's entity gets the `@ManyToOne` (or the owning `@OneToOne`). Then ask whether any code path genuinely needs to navigate the other way; only then add the inverse collection. Unneeded collections are pure liability — they are loaded, cascaded, and reasoned about forever, and a large `@OneToMany` is a well-known source of memory and query problems. `@ManyToMany` deserves particular caution. It is correct only while the association carries no attributes of its own. The moment the business wants "when was this tag added, and by whom", the link table needs columns, and a `@ManyToMany` cannot express them. The standard move is to promote the link table into a first-class entity (`OrderTag`) with two `@ManyToOne` associations, converting one many-to-many into two many-to-ones. Many teams do this pre-emptively because the promotion later is a breaking model change. `@OneToOne` is also worth questioning: if the two tables always exist together and are always loaded together, the relationship may be a sign that the two tables should be one, or that the second is really an `@Embeddable` component. Legitimate uses are optional detail rows, rarely-loaded large payloads, and table-splitting for storage reasons. ## What the annotations do not decide The kind of association does not, by itself, decide the join strategy, the column names, or the write behaviour. Those come from `@JoinColumn`/`@JoinTable`, from `mappedBy` (which end owns the write), and from `cascade`/`orphanRemoval`. Cardinality is only the first of several independent decisions.

  • Does adding the @OneToMany inverse side change the database schema?
    No. A bidirectional pair describes one and the same foreign-key column; the `mappedBy` side is purely a Java navigation path. The schema for `Order.customer` alone and for `Order.customer` plus `Customer.orders` is identical. Direction is a modelling choice with no relational cost — its costs are in loading and lifecycle behaviour instead.
  • When would you replace a @ManyToMany with two @ManyToOne associations?
    As soon as the link itself needs attributes — an added-at timestamp, a quantity, a role, a soft-delete flag — because a `@ManyToMany` link table can only hold the two keys. You introduce an explicit link entity with its own identifier and two `@ManyToOne` fields. It also gives you finer control over writes, since each link row becomes an ordinary managed entity.

Cardinality is how the rooms of a building actually connect; direction is which doors you bothered to cut. The wall is the same either way.

saying these in an interview costs you the question

  • Saying @OneToMany and @ManyToOne are two different relationships that must both be persisted
  • Reading the annotation name backwards, e.g. putting @ManyToOne on the collection
  • Assuming @ManyToMany can carry extra columns such as a quantity or timestamp
  • Claiming a bidirectional mapping requires an extra column or table
  • Modelling every relationship in both directions by default

context

open as a page

In JPA, if you call EntityManager.persist() on a parent entity whose collection holds brand-new child entities, what happens to those children by default, and what does configuring cascade on that association change?

level: juniorimportance: must knowfreq 62%

basics

~20 s

By default persist applies only to the instance you pass. The new children stay transient, and at flush Hibernate typically fails with "object references an unsaved transient instance". Declaring cascade = PERSIST (or ALL) makes the provider run persist on each child too.

open as a page

In JPQL, what does the JOIN FETCH clause do, and how does it differ from writing a plain JOIN in the same query?

level: juniorimportance: must knowfreq 78%

basics

~20 s

JOIN FETCH tells the ORM to load the joined association's rows into the returned entities in the same SQL statement. A plain JOIN only lets you filter or navigate in the query; the association still loads later on access.

open as a page

In JPA, what does the @JoinColumn annotation declare, and which side of an association must carry it?

level: juniorimportance: must knowfreq 75%

basics

~20 s

@JoinColumn names the foreign-key column that stores the association. It goes on the owning side — the side whose table holds the FK, which for @ManyToOne is always the many side. The inverse side uses mappedBy instead.

open as a page

In JPA, what is the default fetch type for each of the four association annotations — @ManyToOne, @OneToOne, @OneToMany and @ManyToMany — and why does the default differ between them?

level: juniorimportance: must knowfreq 82%

basics

~20 s

To-one associations (@ManyToOne, @OneToOne) default to EAGER. To-many associations (@OneToMany, @ManyToMany) default to LAZY. The spec assumes fetching one extra row is cheap and fetching a whole collection is not. Most teams override the to-one defaults to LAZY.

open as a page

A service loads 200 order rows and then reads each order's customer name, and the database log shows 201 statements. Explain what is happening, why it is called the N+1 problem, and why it hurts more than the statement count suggests.

level: juniorimportance: must knowfreq 80%

basics

~20 s

One query fetched the 200 orders; touching each order's lazy customer triggered one more query per order, so 1 + N = 201. The cost is dominated by 200 network round trips, not by the work each tiny query does.

open as a page

In a bidirectional JPA association, what does the `mappedBy` attribute mean, and which side's changes actually cause Hibernate to write the foreign key?

level: middleimportance: must knowfreq 80%

basics

~20 s

mappedBy marks the inverse (non-owning) side and names the owning field on the other entity. Only the owning side — the one holding the foreign-key column, normally the @ManyToOne — produces INSERT/UPDATE of that column. Changing only the inverse collection writes nothing.

open as a page

What does Hibernate's @BatchSize annotation — or the hibernate.default_batch_fetch_size setting — change about how lazy associations are loaded?

level: middleimportance: must knowfreq 56%

basics

~20 s

When one uninitialised proxy or lazy collection is touched, Hibernate also loads up to N other pending ones of the same type in a single query using where owner_id in (?, ?, ...). It turns N follow-up selects into roughly N/batchSize selects.

open as a page

Which cascade operations can be configured on a JPA association, what does each one propagate, and why does the specification make cascading opt-in rather than the default?

level: middleimportance: must knowfreq 55%

basics

~20 s

PERSIST, MERGE, REMOVE, REFRESH, DETACH, and ALL (all five). Each forwards the matching EntityManager call along the association. Cascading is opt-in because propagation is only correct where the parent truly owns the target — REMOVE especially would destroy shared data.

open as a page

What is the difference between mapping a JPA association with CascadeType.REMOVE and mapping it with orphanRemoval = true, and when does each one actually delete a row?

level: middleimportance: must knowfreq 60%

basics

~20 s

CascadeType.REMOVE only fires when you explicitly remove the parent — the children go with it. orphanRemoval = true does that too, and additionally deletes a child the moment it is taken out of the parent's collection or its reference is set to null, even though the parent survives.

open as a page

In a Hibernate entity, a to-many association can be mapped as a java.util.List or a java.util.Set. What does Hibernate call a plain List mapping, and how do the two differ in semantics and in the SQL they produce?

level: middleimportance: must knowfreq 52%

basics

~20 s

A List without an index column is a bag: unordered, duplicates allowed, and Hibernate cannot address a row by position. A Set forbids duplicates using equals/hashCode. For collections that own their rows, a bag change often forces a delete-all-then-reinsert; a Set does not.

open as a page

How do you map a many-to-many association in JPA with @JoinTable, and why do experienced teams often replace @ManyToMany with an explicit link entity?

level: middleimportance: must knowfreq 68%

basics

~20 s

@ManyToMany with @JoinTable maps a link table of two FK columns: joinColumns points at the owner, inverseJoinColumns at the target; the other side uses mappedBy. Teams replace it with a link entity when the relationship needs its own columns or stable identity.

open as a page

When Hibernate hands you a lazily-fetched to-one association, what object do you actually hold, how is that object created, and what constraints does that place on the entity class?

level: middleimportance: must knowfreq 68%

basics

~20 s

You hold a proxy: a runtime-generated subclass of the entity that stores only the identifier. Every inherited method call triggers loading of the real row on first use. That requires a non-final class, non-final methods and an accessible no-arg constructor.

open as a page

Hibernate hands back a proxy object for a lazily mapped association instead of the real entity. What does that proxy hold internally, and why does calling a getter on it after the Hibernate Session that produced it has been closed fail?

level: middleimportance: must knowfreq 70%

basics

~20 s

A lazy proxy is a generated subclass holding an initializer with the target's identifier and a reference to the Session that created it. The first real getter call triggers a deferred SELECT; if that Session is closed the proxy has no connection to run it, so Hibernate throws instead of silently opening a new one.

open as a page

A JPQL query returns 500 entities and the log shows 501 statements, even though the code never dereferences any association. The entity has a @ManyToOne mapped with fetch = FetchType.EAGER. Why does that mapping produce extra statements for a query result when loading the same entity by primary key produces a single joined statement?

level: middleimportance: must knowfreq 62%

basics

~20 s

find() builds its own fetch plan and can add a join. A JPQL query is executed largely as written, so an EAGER association that is not join-fetched must be satisfied afterwards — one secondary select per distinct associated row. No code needs to touch it.

open as a page

You need to render 100 parent records, each with its child collection, from a Hibernate application. Compare loading the children with a fetch join, with batch fetching, and with subselect fetching — how do you choose?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Fetch join: one query, but duplicates parent rows and breaks database-side pagination. Batch fetching: a handful of IN-clause queries, safe with paging and partial traversal. Subselect: exactly one extra query but replays the parent query and ignores its row limit. Paginated screens use batching; unpaginated full traversal favours join or subselect.

open as a page

A JPQL query fetch-joins two different List-mapped collections of the same root entity and Hibernate fails with MultipleBagFetchException ("cannot simultaneously fetch multiple bags"). Why is that rejected, and what are the ways to get the data you wanted?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Joining two bags produces a cartesian product, and because bags allow duplicates and carry no index, Hibernate cannot tell real duplicates from join-generated ones. Fix it by mapping the collections as Sets, or by running one query per collection in the same persistence context.

open as a page

Code fails because a lazily mapped association is touched after the unit of work that loaded it has already finished. Rank the realistic fixes — loading the data inside the transaction, using a fetch join or entity graph, and projecting straight into a DTO — and explain how you choose between them.

level: seniorimportance: must knowfreq 60%

basics

~20 s

Fix it at the query, not at the point of failure. Best: select exactly the fields you need into a DTO. Next: load the entity with a fetch join or entity graph so the needed associations arrive in one query. Acceptable: touch what you need while the context is open. Never: force loading after the fact.

open as a page

How do you confirm and quantify that a running Java application is issuing one database query per returned row — what do you switch on or instrument, and what specifically do you measure?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Measure statements per logical operation, not query duration. Turn on SQL logging (org.hibernate.SQL), or Hibernate Statistics for counts, or wrap the DataSource with datasource-proxy/p6spy to count and assert per request. Then vary the row count and see whether the count scales with it.

open as a page

Why do teams add helper methods such as `addOrder(...)` / `removeOrder(...)` on entities with bidirectional JPA associations instead of letting callers mutate the collection directly?

level: middleimportance: should knowfreq 58%

basics

~20 s

Because a bidirectional pair is one relationship stored in two fields, and both must agree. A helper sets the owning reference and updates the inverse collection in one call, so writes reach the foreign key and the in-memory graph stays truthful for the rest of the session.

open as a page

Hibernate's `@Fetch` annotation accepts `FetchMode.SELECT` and `FetchMode.JOIN`. What is the difference, and why does `FetchMode.JOIN` often appear to be ignored?

level: middleimportance: should knowfreq 34%

basics

~20 s

SELECT loads the association with a separate query when it is needed; JOIN loads it in the same statement with an outer join. JOIN applies only when Hibernate builds the SQL itself — find and navigation — and is ignored by JPQL, HQL and Criteria queries, which use their own explicit fetch plan.

open as a page

In JPA, what is the difference between annotating a collection with @OrderBy and annotating it with @OrderColumn, and what does each one cost?

level: middleimportance: should knowfreq 36%

basics

~20 s

@OrderBy adds an ORDER BY clause when the collection loads — nothing is stored, and the collection stays a bag. @OrderColumn persists each element's position in a dedicated column, making it an indexed list, which costs extra UPDATEs whenever positions shift.

open as a page

What is a JPA entity graph, and how do you define and apply one — both declaratively with @NamedEntityGraph and programmatically with EntityManager.createEntityGraph?

level: middleimportance: should knowfreq 58%

basics

~20 s

An entity graph is a reusable, declarative description of which attributes to load with a query. Define it with @NamedEntityGraph on the entity or build it via em.createEntityGraph(Class), then pass it to a query or find() as a hint.

open as a page

A developer maps a one-to-many collection in JPA with only @OneToMany and no mappedBy or @JoinColumn. What schema does the provider generate, why is that surprising, and what are the alternatives?

level: middleimportance: should knowfreq 47%

basics

~20 s

The provider creates a join table instead of a foreign key in the child table — the default for a unidirectional one-to-many. Alternatives: add @JoinColumn to keep the FK in the child (costs extra UPDATEs), or make it bidirectional with mappedBy, which is cleanest.

open as a page

Given an entity instance returned by Hibernate, how do you determine whether one of its associations has already been loaded without causing it to load, and how do you deliberately force it to load or obtain the underlying non-proxy instance?

level: middleimportance: should knowfreq 42%

basics

~10 s

Use Hibernate.isInitialized(x) (or PersistenceUnitUtil.isLoaded) to test without loading, Hibernate.initialize(x) to force the load while a session is open, and Hibernate.unproxy(x) to get the real instance behind a proxy.

open as a page

In a bidirectional JPA @ManyToMany between, say, orders and tags, which side owns the relationship, and what SQL does Hibernate emit when you remove one element from the owning collection?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The side declaring @JoinTable owns it; the other uses mappedBy and never writes. Removing from a Set-typed owning collection emits one targeted DELETE of that link row. With a List (bag), Hibernate deletes every link row for that owner and re-inserts the survivors.

open as a page

What does Hibernate's `@Fetch(FetchMode.SUBSELECT)` do on a collection, and what conditions must hold for it to actually take effect?

level: seniorimportance: should knowfreq 38%

basics

~20 s

When any one collection from a previously-run query's results is initialised, Hibernate loads that collection for all those parents in one query, re-using the original query as a subselect in the WHERE clause. It needs the original query still tracked in the same session, and applies to collections only.

open as a page

You load an order with its lines, let the persistence context close, edit the detached graph (change one line, add a new one, drop another), then call EntityManager.merge() on the order. What does the cascade configuration control here, and what happens to the line you dropped?

level: seniorimportance: should knowfreq 38%

basics

~20 s

CascadeType.MERGE decides whether the children are merged at all — edited ones update, new ones become inserts. The dropped line is deleted only if the association declares orphanRemoval = true; otherwise its row survives. And merge returns a managed copy: your detached object stays detached.

open as a page

Why is declaring CascadeType.REMOVE (or CascadeType.ALL) on a JPA @ManyToMany association — or on the @ManyToOne side of a parent-child relationship — considered dangerous?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Because the target is shared. Removing one post would delete tags other posts still use, or on a @ManyToOne would delete the parent and, through it, every sibling. Cascade must only follow edges where the source truly owns the target.

open as a page

You remove a single element from a Hibernate-owned collection — an @ElementCollection or a unidirectional @OneToMany — and the SQL log shows a DELETE for every row of that collection followed by re-INSERTs of the survivors. Why does Hibernate do this, and how do you stop it?

level: seniorimportance: should knowfreq 34%

basics

~20 s

The collection is a bag: a List with no index column. Hibernate cannot tell which row an element corresponds to, so it recreates the whole collection. Fix it by mapping the collection as a Set, adding @OrderColumn, or making a one-to-many bidirectional with orphanRemoval.

open as a page

showing 1–30 of 45