In a Hibernate entity, a to-many association can be mapped as a java.util.List or a java.util.Set. What does Hibernate call a plain List mapping, and how do the two differ in semantics and in the SQL they produce?
answer
- bag = List, no index, duplicates ok
- set = no duplicates via equals/hashCode
- indexed list = @OrderColumn stores position
- bag + owned rows → delete-all + reinsert
- mappedBy inverse bag drives no DML
basics
~20 sA List without an index column is a bag: unordered, duplicates allowed, and Hibernate cannot address a row by position. A Set forbids duplicates using equals/hashCode. For collections that own their rows, a bag change often forces a delete-all-then-reinsert; a Set does not.
solid answer
~60 sHibernate has three collection semantics, not two: **bag** (a `List` with no `@OrderColumn` — unordered, duplicates allowed), **list** (a `List` with `@OrderColumn` — position stored in a column), and **set** (a `Set` — no duplicates, order not persisted). The difference that matters at runtime is *row addressability*. In a bag, Hibernate has no key for an element's row inside the collection, so for collections it owns — a unidirectional `@OneToMany` with a join column or join table, or an `@ElementCollection` — a change means deleting every row for the parent and re-inserting the survivors. A `Set` identifies rows by their values or ids and issues targeted `INSERT`/`DELETE`. Duplicates also behave differently in query results: fetch-joining a bag can put the same associated instance in the collection more than once, while a `Set` deduplicates by `equals`/`hashCode`. That deduplication depends on the entity's `equals`/`hashCode` being sane for objects whose id is assigned only on flush. A bidirectional `@OneToMany(mappedBy = ...)` bag is inverse and drives no DML, so it escapes delete-all-reinsert — but it still cannot be fetch-joined alongside a second bag.
code
java · 13 lines// bag: unordered, duplicates allowed, no index column
@OneToMany(mappedBy = "post")
private List<Comment> comments = new ArrayList<>();
// indexed list: position persisted in comment_order
@OneToMany(cascade = CascadeType.ALL, orphanRemoval = true)
@JoinColumn(name = "post_id")
@OrderColumn(name = "comment_order")
private List<Comment> ordered = new ArrayList<>();
// set: no duplicates, order not persisted
@ElementCollection
private Set<String> tags = new HashSet<>();go deeper
Know that a plain List is a bag (unordered, duplicates allowed) and a Set forbids duplicates, and that ordering needs an explicit annotation.
Explain row addressability and the delete-all-reinsert behaviour, plus which mappings it applies to.
Add the equals/hashCode consequences of Set, the duplicate-results behaviour, and how the choice constrains fetching strategy later.
Discuss collection semantics as a model-wide convention and the migration cost of changing it once schemas and code depend on the current shape.
## Three collection semantics What Java type you declare selects the semantics Hibernate applies: | Java type + mapping | Hibernate semantics | Ordering | Duplicates | |---|---|---|---| | `List` with no `@OrderColumn` | **bag** | none persisted (`@OrderBy` can sort on load) | allowed | | `List` with `@OrderColumn` | **list** (indexed) | position stored in a column | allowed | | `Set` | **set** | none (`SortedSet` + `@SortNatural`/`@SortComparator` sorts in memory) | forbidden | "Bag" is the relational term for a multiset: a collection with no order and no uniqueness. A table has exactly those properties — rows are unordered and can repeat — which is why a bag is the *natural* mapping of a foreign-key relationship, and also why it has the least information available to Hibernate. ## Why bags cost more: row addressability When Hibernate loads a collection it takes a **snapshot** of it. At flush it diffs the current collection against the snapshot to work out what changed. To issue a targeted statement it must be able to say *which row* an element corresponds to. - **Set** — an element maps to a row identified by its own key (the child's id, or the full tuple for an `@ElementCollection`). Remove one element and Hibernate emits one `DELETE` for that key. - **Indexed list** — the row is identified by the index column, so removals and reorderings become `UPDATE`s on the index plus a `DELETE`. - **Bag** — no index, and duplicates are permitted, so "this element" does not identify a row. If the collection owns its rows, Hibernate takes the only sound route: delete every row belonging to the parent and re-insert the remaining elements. That delete-all-and-recreate applies to collections whose DML Hibernate drives: - `@ElementCollection` (a value collection) - unidirectional `@OneToMany` with `@JoinColumn` or a join table - the owning side of a `@ManyToMany` (join-table rows) It does **not** apply to a bidirectional `@OneToMany(mappedBy = "parent")`. That collection is the inverse side: it drives no DML at all, so removing an element produces nothing (or a delete via `orphanRemoval`, and an FK update from the owning side). This is the single most common correction to the blanket claim "Lists are always slow in Hibernate". ## Duplicates in results A bag keeps duplicates, and SQL joins produce them. A query that fetch-joins a collection returns one result row per child row, so a parent with three children appears three times. If the collection being filled is a bag, and the query result is a `List` of roots, you can see the same root repeated. A `Set`-mapped collection collapses the repeats automatically. That difference is the reason `SELECT DISTINCT` and result deduplication crop up constantly with bag mappings and rarely with sets. ## The cost of Set: equals and hashCode A `HashSet` needs a stable hash for its elements. Entities whose `hashCode` derives from a `@GeneratedValue` id break this: an object added while transient hashes to one bucket, then its id is populated on flush and it hashes elsewhere — `contains` returns false for an element that is in the set. The standard fixes are a business/natural key, or a constant `hashCode` with an id-plus-class-based `equals` (correct but degrades the set to a linear scan, which is acceptable for the small collections you should be loading anyway). ## Choosing - Owned collection whose contents change (`@ElementCollection`, unidirectional one-to-many, many-to-many join rows) → **Set**, unless order is part of the domain. - Order is genuinely part of the data (steps in a recipe, slides in a deck) → indexed **list** with `@OrderColumn`. - Just want a deterministic display order → `List` (bag) with `@OrderBy`, or a `SortedSet`. - Bidirectional inverse one-to-many → `List` is fine for DML, but a `Set` costs nothing and keeps multi-collection fetching options open.
- Does a bidirectional @OneToMany mapped as a List always suffer delete-all-and-reinsert?No. With `mappedBy` the collection is the inverse side and drives no DML whatsoever — the child's `@ManyToOne` owns the foreign key, so inserts, FK updates and orphan deletes come from the child side. Delete-all-and-reinsert affects collections Hibernate owns: `@ElementCollection`, unidirectional `@OneToMany` with a join column or join table, and the owning side of a many-to-many.
- What breaks when an entity in a HashSet derives hashCode from its @GeneratedValue id?The hash changes when the id is assigned at flush, so the element sits in the wrong bucket: `contains` and `remove` stop finding it, and duplicates can slip in. Use a stable natural or business key for `equals`/`hashCode`, or return a constant `hashCode` with an id-and-type-based `equals` — correct, at the cost of linear lookup inside the set, which is fine for the small collections you should be loading.
A bag is a shopping bag: you can see what's inside and there may be two identical tins, but you can't say "the third item". A set is a keyring — each key is distinct and you can pull out exactly one.
saying these in an interview costs you the question
- Saying "use Set, List is always slower" without distinguishing owned from inverse collections
- Believing a List mapping preserves insertion order in the database without @OrderColumn
- Thinking @OrderBy persists the order rather than just sorting on load
- Deriving hashCode from a generated id and putting the entity in a HashSet
- Assuming a Set mapping eliminates the duplicate rows the SQL join itself returns