In a bidirectional JPA @ManyToMany between, say, orders and tags, which side owns the relationship, and what SQL does Hibernate emit when you remove one element from the owning collection?
answer
- @JoinTable side owns; mappedBy side is silent
- Set = one targeted DELETE
- List bag = delete all, re-insert survivors
- Never cascade REMOVE across a link
- Attributes on the link means promote to an entity
basics
~20 sThe side declaring @JoinTable owns it; the other uses mappedBy and never writes. Removing from a Set-typed owning collection emits one targeted DELETE of that link row. With a List (bag), Hibernate deletes every link row for that owner and re-inserts the survivors.
solid answer
~50 sWith `@ManyToMany`, neither table holds a foreign key — a link table does — so ownership is a free choice. The owner is whichever side declares `@JoinTable`; the other declares `mappedBy` and produces no SQL at all, so mutating only the inverse collection changes nothing in the database. What the owner emits depends on the collection type. A `Set` gives Hibernate identity-based semantics, so removing one element produces a single `delete from order_tag where order_id=? and tag_id=?`. A `List` without `@OrderColumn` is a **bag**: Hibernate cannot address a single row reliably, so it deletes all link rows for that owner and re-inserts the remaining ones. On a tag list of a hundred entries, dropping one becomes one DELETE plus ninety-nine INSERTs, and it churns the table under concurrency. Practically: map `@ManyToMany` as `Set`, choose as owner the side your code actually mutates, and never cascade REMOVE across it — that would delete the other entity, not the link.
code
java · 14 lines@Entity
class Order {
@ManyToMany
@JoinTable(name = "order_tag",
joinColumns = @JoinColumn(name = "order_id"),
inverseJoinColumns = @JoinColumn(name = "tag_id"))
private Set<Tag> tags = new HashSet<>(); // owner
}
@Entity
class Tag {
@ManyToMany(mappedBy = "tags")
private Set<Order> orders = new HashSet<>(); // writes nothing
}go deeper
Know that a link table stores the pairs, that one side declares @JoinTable and owns it, and that the mappedBy side does not write.
Add the Set-versus-List difference in emitted SQL and why cascade REMOVE is inappropriate across a many-to-many.
Reason about write amplification, lock footprint and concurrent edits on the same owner, plus indexing of the link table and when to promote it to an entity.
Treat the link table as a first-class design decision: whether the association is truly attribute-free, how it will evolve, and the cost of the migration from @ManyToMany to an explicit entity once it is not.
## Ownership when no table holds the key For a many-to-one, ownership is dictated by physics: the foreign key sits in one table and that table's entity owns the mapping. A many-to-many has no such column. The relationship lives in a third table of key pairs, and **either** entity may be nominated as its owner. The owner is the side that declares `@JoinTable` (or accepts the defaults); the other declares `@OneToMany`-style `mappedBy` and is purely a read view. The rule that follows is the same as everywhere else: only the owning collection is inspected at flush. `tag.getOrders().add(order)` on the inverse side writes nothing. `order.getTags().add(tag)` on the owning side writes a link row. Because the choice is arbitrary, make it deliberately: pick the side your service code naturally mutates, and — where the two collections differ hugely in size — the side whose collection is small, because writing through the owner means initialising that collection. ## What removal actually emits This is the part interviewers probe, because it separates people who have watched the SQL log from people who have only read annotations. **Set-typed owning collection.** Hibernate can identify each link row by the pair of keys, so a single removal produces: ```sql delete from order_tag where order_id = ? and tag_id = ? ``` One statement, proportional to the change. **List-typed owning collection with no `@OrderColumn`.** A `List` here is a *bag*: an unordered collection that permits duplicates and carries no index column. Hibernate has no stable way to say "delete the third one", so its recovery strategy is wholesale: delete every link row belonging to the owner, then re-insert the ones that remain. ```sql delete from order_tag where order_id = ? insert into order_tag (order_id, tag_id) values (?, ?) -- repeated for survivors ``` For a collection of size N, removing one element costs 1 + (N-1) statements. This is invisible at N = 3 and ruinous at N = 500. It also widens the write footprint: the transaction touches every link row of that owner, so two transactions editing different tags of the same order now block each other, and the table accumulates dead rows. Adding `@OrderColumn` turns the bag into an indexed list, which restores targeted deletes but shifts index values on removal, generating its own UPDATE traffic. The pragmatic default is `Set`. ## Cascade and @ManyToMany `CascadeType.REMOVE` (and `ALL`, which contains it) is almost always wrong here. Cascade operates on **entities**, not on link rows: removing an order would attempt to delete the tags themselves, which other orders still reference. Hibernate removes the link rows for the deleted entity automatically as part of collection cleanup; you never need cascade to achieve that. `PERSIST` and `MERGE` are sometimes reasonable, `REMOVE` essentially never. The same reasoning rules out `orphanRemoval`, which JPA does not permit on a many-to-many at all. ## Concurrency and the link table Because the link table has no entity, you get no `@Version` on it and no per-row optimistic control. Two users adding different tags to the same order concurrently is fine with `Set` semantics (two independent inserts, subject to the unique constraint on the pair), but with bag semantics both transactions delete and rewrite the whole set — the second overwrites the first's addition or deadlocks. That is another concrete argument for `Set`, and one of the reasons teams eventually promote the link table to an explicit entity. ## When to abandon @ManyToMany A `@ManyToMany` is only valid while the link carries nothing but the two keys. Any attribute on the association — added-at, added-by, position, weight, an active flag — forces an explicit link entity with two `@ManyToOne` fields and its own identifier. That conversion also solves the write-amplification story permanently: each link becomes an ordinary entity with ordinary INSERT/DELETE behaviour, individually addressable, versionable, and queryable. Many teams start with the explicit entity for exactly that reason, treating `@ManyToMany` as a convenience for genuinely static pairings such as user-to-role. ## Answering well Say ownership is a free choice expressed by `@JoinTable`; say the inverse side writes nothing; then contrast `Set` (one targeted DELETE) with bag `List` (delete-all plus re-insert), and close with cascade REMOVE being wrong and the link-entity escape hatch. That is the full arc of the question.
- Why is CascadeType.REMOVE on a @ManyToMany usually a bug?Cascade propagates lifecycle operations to the *associated entities*, so removing an order would try to delete its tags, which other orders still reference — either a constraint violation or silent data loss. The link rows for the removed entity are cleaned up by Hibernate anyway as part of collection handling, so the cascade buys nothing. PERSIST or MERGE can be defensible; REMOVE and orphanRemoval are not.
- How do you keep the link table's rows unique and efficiently queryable?Put a composite primary key or unique constraint on the pair of columns, which both prevents duplicate links and gives an index for one navigation direction, then add a second index leading with the other column so lookups from either side are indexed. Hibernate does not add these for you beyond what schema generation emits, so they belong in your migration.
saying these in an interview costs you the question
- Thinking both sides of a @ManyToMany write to the join table
- Using List for a @ManyToMany and being surprised by delete-all-then-reinsert
- Adding cascade = ALL to a @ManyToMany
- Expecting to store extra columns such as a timestamp in a @ManyToMany join table
- Believing orphanRemoval works on a many-to-many