For an association table entity such as order_line, how would you decide between a composite primary key built from the two foreign keys and a single surrogate key with a unique constraint on the pair? What drives the call, and what does each choice cost you later?
answer
- Composite: identity is the pair, no extra column
- Surrogate: one stable handle, keep UNIQUE(pair)
- Anything references the row → surrogate
- Key width propagates down child tables
- Composite keys must never change; no @GeneratedValue
basics
~20 sComposite keys model the truth (a pair is unique) with no extra column, but every child table, cache key and query carries both columns and the key must never change. A surrogate key gives one stable, cheap handle to reference and evolve, at the cost of an extra column plus a unique constraint you must not forget.
solid answer
~60 sBoth are defensible; the decision turns on how the row will be *referenced* and how stable the pair is. Favour a **composite key** (`@EmbeddedId` + `@MapsId`) when: nothing else points at the link row, the pair is genuinely immutable, and you want the database to enforce the pairing as the primary key. You save a column and an index, and the natural clustered/primary index directly supports lookups by the leading key part. Favour a **surrogate key** plus `UNIQUE(order_id, product_id)` when: other tables must reference the row, the row grows its own lifecycle (audit columns, soft delete, versioning), the pair might change, or the key would be wide and propagate into many child tables and cache keys. It also removes the composite-key ergonomics tax — no id class to keep correct, no nested JPQL paths, straightforward `find(id)`. My default for a pure link table with payload is the composite key while it stays pure, and a surrogate key the moment anything references it. Whichever you pick, make it a codebase-wide convention rather than a per-entity coin flip.
code
sql · 12 lines-- composite
CREATE TABLE order_line (
order_id BIGINT NOT NULL, product_id BIGINT NOT NULL, quantity INT NOT NULL,
PRIMARY KEY (order_id, product_id)
);
-- surrogate (unique constraint replaces the PK's enforcement)
CREATE TABLE order_line (
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
order_id BIGINT NOT NULL, product_id BIGINT NOT NULL, quantity INT NOT NULL,
UNIQUE (order_id, product_id)
);go deeper
Know both options exist and that a surrogate key still needs a unique constraint on the business pair.
Name the concrete costs: extra column and index versus id-class boilerplate, nested query paths and no generated identifiers.
Lead with referenceability and mutability, discuss key propagation into child tables and caches, and note the ORM fragility of composite key classes.
Frame it as a durable, hard-to-reverse decision: set a house convention, tie it to how generic infrastructure (auditing, caching, events) is built, and justify the default with the migration asymmetry.
## What the choice actually is A link entity connects two parents and may carry payload (`quantity`, `price_at_time`). Its identity can be expressed two ways: 1. **Composite key** — `PRIMARY KEY (order_id, product_id)`, mapped with `@EmbeddedId` plus `@MapsId` on the two associations. 2. **Surrogate key** — `id BIGINT PRIMARY KEY` (generated) with `UNIQUE (order_id, product_id)` preserving the business rule. Both enforce the same invariant. The difference is everything downstream of identity. ## Arguments for the composite key - **The model tells the truth.** Identity *is* the pair; there is no second notion of "which row" to keep consistent. - **No extra column or sequence.** One fewer index than surrogate + unique constraint, which matters on very wide tables. - **Index alignment.** The primary key index already supports "all lines of order X" lookups on the leading column, so a common access path is free. - **Idempotent writes are natural.** An upsert keyed on the pair is expressible without a lookup round trip. ## Arguments for the surrogate key - **Referenceability.** The moment a third table (an allocation, a shipment line, an audit record) points at the link row, a composite key propagates *both* columns into that table, and into anything referencing *that*. Key width compounds down the graph. - **Mutability.** Composite keys must be immutable; Hibernate does not support changing a primary key. If a business change means "move this line to another product", a composite key forces delete+insert and loses row identity (and any history keyed to it). A surrogate key survives the change as a plain UPDATE. - **ORM ergonomics.** No id class to write, keep `Serializable`, and keep `equals`/`hashCode` correct — a class of silent bugs disappears. Queries use flat `find(id)`, DTO projections and API contracts carry one value, and generated identifiers work again (composite keys cannot use `@GeneratedValue`). - **Caching and messaging.** Second-level cache keys, distributed cache keys, and event payloads all get simpler and smaller with one scalar id. - **Uniform conventions.** If every other entity in the system has `Long id`, a lone composite-key entity forces every generic layer (base repository, auditing, soft-delete, outbox) to special-case it. ## How I decide Questions I ask, in order: 1. **Will anything reference this row?** If yes — surrogate key, almost regardless of the rest. 2. **Is the pair truly immutable in the business?** If reassignment is plausible, surrogate key. 3. **Does the row have its own lifecycle** (created/updated audit, versioning, state)? Rows with lifecycle behave like entities; give them their own identity. 4. **How wide is the key?** Two `BIGINT`s are cheap; a pair of `VARCHAR(64)` natural keys propagating into three child tables is not. 5. **What does the rest of the codebase do?** Consistency has real value for generic infrastructure. If the answers are all "pure link table, immutable, nothing references it", the composite key is the cleaner model and I take it — with `@MapsId`, an immutable record as the `@EmbeddedId`, and a house rule that key values are never mutated. ## Things people get wrong in this discussion - **Dropping the unique constraint** when going surrogate. The surrogate key removes the *primary-key* enforcement of the pair; the `UNIQUE` index must replace it or you get duplicate links, which is strictly worse than either design. - **Claiming composite keys are "slow".** At two BIGINTs the storage/comparison difference is negligible; the real cost is key propagation and ORM ergonomics, not comparison speed. - **Treating it as reversible.** Changing an established primary key means migrating every referencing table and every cached/serialised key. Decide once, early, and write it down. - **Mixing both styles** across a codebase, which multiplies special cases in shared infrastructure. ## Migration reality Going from composite to surrogate later is a schema migration plus a mapping change plus rewrites of everything that referenced the pair — doable but never cheap. Going the other direction is rarer and usually only makes sense when deduplicating a table that grew duplicates. That asymmetry is a decent tiebreaker: when genuinely unsure, the surrogate key is the option that keeps more doors open.
- If you choose the surrogate key, what must you not forget?The unique constraint on the pair. The composite primary key was doing double duty — identity and the business rule that a product appears once per order — and a surrogate key only replaces the first half. Without `UNIQUE(order_id, product_id)` the table quietly accumulates duplicate links, which application-level checks will not reliably prevent under concurrency.
- Does the composite-key choice change how the entity behaves in the persistence context?Yes, in one important way: the identity map, second-level cache and merge resolution key on your `@EmbeddedId` class, so its `equals`/`hashCode` correctness and immutability become load-bearing. A surrogate `Long` id gets that for free. That extra fragility is a real, if often understated, part of the cost side of composite keys.
saying these in an interview costs you the question
- Switching to a surrogate key and dropping the uniqueness guarantee on the pair
- Arguing composite keys are inherently slow rather than costly to propagate
- Ignoring that a composite key can never be updated in place
- Choosing per-entity with no codebase-wide convention
- Assuming the decision can be reversed cheaply later