skip to content

You are designing the identity model for a domain whose objects routinely live outside an open persistence context — cached DTO graphs, HTTP sessions, sets nested inside aggregates. How would you decide between assigning identifiers in the constructor (for example UUIDs), relying on a business/natural key, and using database-generated sequence identifiers as the basis for object equality?

level: principalimportance: nice to knowfreq 26%

answer

  1. Discriminator: when does identity first exist?
  2. Assigned UUID = stable hash everywhere, wider + random inserts
  3. UUIDv7 / @UuidGenerator(style = TIME) restores locality
  4. IDENTITY generation kills insert batching; SEQUENCE with allocationSize does not
  5. Business key only if an external authority owns it

basics

~20 s

Pick by whether an identifier exists at construction time. Assigned UUIDs always do, so equality is simple everywhere, at the cost of index locality — mitigate with time-ordered UUIDs. A true immutable business key is best when one genuinely exists. Generated sequence ids are cheapest in the database but need the constant-hashCode workaround.

solid answer

~60 s

The deciding question is: **when does a stable identity value first exist?** - **Assigned identifier (UUID in the constructor)** — identity exists from `new`, so `equals`/`hashCode` are trivial and correct in sets, caches, detached graphs and across services. Costs: 16 bytes instead of 8, and random UUIDs scatter B-tree inserts, inflating index size and cache misses. Time-ordered variants (UUIDv7, `@UuidGenerator(style = TIME)`) recover most of that. It also decouples id generation from the database, which helps offline creation and event-carried ids. - **Business/natural key** — cleanest when the domain truly has an immutable unique attribute (ISBN, ISO code, SKU). Declare it `@NaturalId`. Risk: keys that look immutable rarely stay so; the day it changes you have identity churn across every cache. - **Generated sequence id** — narrow, index-friendly, cheap, and the default for good reason; but identity is null until flush, so equality needs the constant-hashCode-plus-non-null-id pattern, and objects added to sets before persist behave subtly. For a model whose objects live outside the session, I default to assigned time-ordered UUIDs and treat entities as values with stable identity.

code

java · 12 lines
java
@Entity
class Shipment {
    @Id
    @UuidGenerator(style = UuidGenerator.Style.TIME) // Hibernate 6, time-ordered
    @Column(columnDefinition = "uuid")               // native type, not varchar(36)
    private UUID id;

    @Override public boolean equals(Object o) {
        return o instanceof Shipment other && id.equals(other.getId());
    }
    @Override public int hashCode() { return id.hashCode(); }
}

go deeper

for a junior

Recall the three options and the one fact that decides between them: whether the identifier exists before persist.

for a middle

Compare storage and equality consequences for each option, and name the constant-hashCode pattern that makes generated ids usable.

for a senior

Bring measured trade-offs — index locality, native uuid columns, IDENTITY versus SEQUENCE and batching — and state which entities in a real model justify assigned ids.

for a principal

Make it a system-wide stance with an enforcement mechanism, tie it to service boundaries and event-carried identity, and be explicit about what you are buying and paying on each side.

## Frame the decision correctly The interview trap here is to answer "UUID vs sequence" as a database question. It is really an **object-identity** question with a database bill attached. The single discriminator is *when the identity value comes into existence*, because that determines whether `equals`/`hashCode` can be total and stable — and stability is a hard contract requirement the moment an object enters a `HashSet` or a cache. ## Option A — identifier assigned at construction ```java @Id private final UUID id = UUID.randomUUID(); ``` **Why it wins on identity.** The value is non-null and immutable from the first instant, so `hashCode` never changes, `equals` needs no null special-casing, distinct new objects are distinct in a set, and an entity that has been serialized into an HTTP session or a Redis cache and rehydrated compares equal to the managed instance. Every failure mode of the generated-id approach evaporates. **Secondary benefits.** Identity is available before the database is touched, so you can build an entire object graph, publish an event carrying the id, or let an offline client create records — all without a round trip. It also removes an identity-column round trip that blocks JDBC batching: with `GenerationType.IDENTITY`, Hibernate must execute each INSERT immediately to read the generated key, which disables insert batching entirely. Assigned ids (and sequences with an allocation size) do not have that problem. **Costs, honestly.** A `uuid`/`binary(16)` is twice the width of a `bigint`, and it is duplicated into every foreign key and every index entry — meaningful when a table has several indexes and hundreds of millions of rows. More importantly, random UUIDs insert into random points of a B-tree: page splits, poor cache locality, larger indexes, worse write throughput. This is real and measurable on write-heavy tables. **Time-ordered UUIDs (UUIDv7, ULID, or Hibernate's `@UuidGenerator(style = TIME)`) restore sequential insert locality** while keeping client-side generation, and are the sensible default today. Storing them as native `uuid` (PostgreSQL) or `binary(16)` rather than a 36-char string matters too. **Exposure.** Non-guessable ids in URLs are a mild security benefit over sequential ids (no enumeration, no inventory leakage) — a genuine, if secondary, argument. ## Option B — a real business key If the domain hands you an attribute that is genuinely unique, non-null and immutable — ISBN, ISO 4217 currency code, an externally assigned SKU — hashing on it is the most honest model, and Hibernate's `@NaturalId` both documents it and enforces immutability. It reads well and needs no extra column. The failure mode is misjudging immutability. Email addresses, usernames, national identifiers, and "the customer code, which never changes" all have a habit of changing. When the key changes, every cached copy, every set membership, and every external reference is now wrong. My rule: use a business key for equality only when an external authority owns the value and the domain has no mechanism to change it. Otherwise keep the business key as a `@NaturalId` with a unique constraint for *lookup*, and use a surrogate for *identity*. ## Option C — database-generated sequence The default choice, and defensible: 8 bytes, monotonic, index-friendly, no client-side entropy. With `SEQUENCE` and a sensible `allocationSize` (Hibernate's pooled optimizer), ids are cheap and batching works. The cost is entirely on the object side: identity is `null` until flush. That forces the constant-`hashCode` pattern with `equals` returning false for null ids, and it means a transient object added to a set is only equal to itself — usually the right semantics, but you must implement it deliberately. If your entities never leave the transaction, this cost is zero and Option C is right. ## How I would actually decide 1. **Do entities cross the persistence-context boundary?** If not — short transactional units of work, DTOs at the edge, nothing cached — generated sequences with the constant-hashCode pattern are fine, and I would not add a UUID column for aesthetics. 2. **Do they?** (Long-lived graphs, distributed caches, offline clients, event-carried state, ids minted by one service and referenced by another.) Then assigned identity is worth its cost, and I choose time-ordered UUIDs stored natively. 3. **Is there an externally-owned immutable key?** Then use it for equality and for `@NaturalId` lookups, and still keep a surrogate primary key so foreign keys stay narrow and stable. 4. **Whatever I pick, I make it a house rule**, ideally enforced by a base class or an architecture test, because the failure mode of mixed strategies is silent: one entity in the model with a naive id-based `hashCode` is enough to produce a duplicate row in a `Set` a year later. ## The trade-off to state out loud Assigned ids buy correct object semantics with storage and write-locality; generated ids buy database efficiency with a fragile object-identity contract you must implement carefully. Time-ordered UUIDs shrink the first cost enough that, for a model whose objects live outside the session, I take the assigned-identity side — and I say so with the numbers and the mitigation rather than as a preference.

  • How much does the random-UUID insert problem actually cost, and how do you mitigate it?
    Random keys scatter inserts across the B-tree, causing page splits, poor buffer-cache locality and larger indexes; on write-heavy tables of hundreds of millions of rows this shows up as reduced insert throughput and inflated index size. Mitigations are time-ordered identifiers (UUIDv7 / ULID / Hibernate's TIME style), storing them as a native uuid or binary(16) rather than a 36-character string, and keeping the number of secondary indexes on the table honest.
  • Does the identifier strategy affect insert performance beyond index locality?
    Yes. GenerationType.IDENTITY forces Hibernate to execute each INSERT immediately to read back the generated key, which disables JDBC insert batching altogether. Sequences with a pooled allocationSize, and client-assigned identifiers, let Hibernate defer and batch inserts. That difference often dwarfs the storage argument in bulk-write paths.
  • If you use a business key for equality, what breaks the day the key changes?
    Every cached copy keyed by that value, every HashSet or HashMap containing the object, and every external reference become stale — the object silently stops matching itself. That is why a business key should carry equality only when an external authority owns the value and the domain has no path to change it; otherwise expose it as a @NaturalId for lookup and keep a surrogate identifier for identity.

A passport number issued at birth (assigned UUID) lets you be identified anywhere, immediately. A number issued only when you first enter a country (generated id) is cheaper to print but useless in every conversation before that moment.

saying these in an interview costs you the question

  • Treating it purely as a storage question and never mentioning equals/hashCode stability
  • "UUIDs are always slower" — ignoring time-ordered variants and native column types
  • Assuming email or username is immutable enough to carry equality
  • Not knowing that GenerationType.IDENTITY disables JDBC insert batching
  • Mixing strategies across the model with no house rule, so one bad entity poisons a Set years later

context