You assign entity identifiers in application code — for example a UUID set in the constructor — instead of letting the database generate them. What does that change for Hibernate's handling of the entity, and what is the storage and index cost of random UUIDs?
answer
- id == null no longer means new
- merge() must SELECT first
- @Version or Persistable restores the signal
- Random v4 -> page splits, clustered-index scatter
- UUID v7 / time-ordered + native 16-byte column
basics
~20 sAn id that is never null removes Hibernate's usual new-versus-detached signal, so merge() must SELECT before it knows whether to insert. Random UUIDs are 16 bytes and insert at random points in the primary-key index, causing page splits and poor cache locality; time-ordered UUIDs and a native 16-byte column fix most of that.
solid answer
~60 s**Lifecycle.** Hibernate's default "is this instance new?" test is `id == null`. With an assigned id that test always says "detached", so `merge()` must issue a SELECT to find out whether the row exists, and an unguarded merge-based save path costs one extra query per insert. `persist()` still works — you are simply asserting the entity is new — but you must know which it is. Alternatives are a `@Version` field (a null version means new) or implementing Hibernate's `Persistable` interface to answer the question explicitly. **Storage.** A UUID is 16 bytes against 8 for a bigint, and that width is repeated in every secondary index and every foreign key. **Index behaviour.** Version-4 UUIDs are random, so each insert targets a random leaf page: page splits, a large hot working set, and fragmentation. Prefer a time-ordered UUID (Hibernate's `@UuidGenerator` time-based style, or UUID v7) so inserts append, and store it in a native `uuid`/`binary(16)` column rather than a 36-character string. **Upside:** ids exist before any database contact — no round trip, offline graph construction, ids mintable by clients or other services.
code
java · 16 lines@Entity
public class Document {
@Id
private UUID id = UUID.randomUUID(); // assigned in the constructor
@Version
private Integer version; // null version => Hibernate knows it is new
}
@Entity
public class Event {
@Id
@GeneratedValue
@UuidGenerator(style = UuidGenerator.Style.TIME) // roughly ascending values
private UUID id;
}go deeper
Know that assigning the id yourself means no database round trip, and that a UUID is bigger than a bigint.
Explain that Hibernate's new-versus-detached check relies on a null id, so merge triggers a SELECT, and that random UUIDs hurt index locality.
Discuss the fixes — @Version or Persistable, time-ordered UUIDs, native 16-byte columns — and quantify the clustered-index and secondary-index cost.
Decide who owns the id space across services and clients, weigh idempotent creates and offline graph construction against storage and locality, and consider hybrid surrogate plus external-id designs.
## What changes inside Hibernate Hibernate distinguishes three states: transient (never persisted), managed (in the persistence context), detached (was persisted, context closed). Several operations need to know which state an arbitrary instance is in, and the cheap default heuristic is the identifier's "unsaved value": a `null` id means transient. An application-assigned id is never null, so that heuristic collapses: - `merge(entity)` can no longer shortcut. It must load the row by id — a SELECT — and then either copy state onto the loaded instance (update) or treat it as new (insert). Save-by-merge code therefore performs one extra query per new entity, which is invisible in tests with three rows and very visible in a bulk import. - `persist(entity)` is unaffected but shifts the burden to you: it asserts the entity is new, and calling it on something that already exists in the database surfaces as a constraint violation at flush. Ways to restore the signal: - Add a `@Version` field. A `null` version on a boxed type means "never persisted", and Hibernate uses that as the unsaved-value marker. - Implement Hibernate's `Persistable` interface (`isNew()`), answering explicitly — typically "new until the first flush", using a lifecycle callback such as `@PostPersist`/`@PostLoad` to flip the flag. ## Why the id is worth minting yourself - **No round trip.** Sequence and identity generation each need the database; a UUID does not. You can build an entire object graph — parents, children, join-table links — before opening a transaction. - **Client- and cross-system-generated ids.** A mobile client or an upstream service can mint the id, which makes retries idempotent: replaying a create with the same id hits a primary-key conflict rather than duplicating the row. - **No enumeration leak.** Sequential ids in URLs reveal volume and allow scraping by increment; UUIDs do not. (They are still not a secret — never treat an unguessable id as authorisation.) - **Merge-friendly across databases.** Ids stay unique when data from several instances is combined. ## The index cost, concretely Primary keys live in a B-tree. A monotonically increasing key always inserts into the right-most leaf: that page stays in cache, splits are rare and clean, and the tree stays dense. A random 128-bit value picks a random leaf every time, so: - the working set is effectively the whole index rather than its right edge, and once the index exceeds memory every insert becomes a read plus a write; - pages split in the middle, leaving them half full — the index occupies more space and caches worse; - on engines with clustered primary keys (InnoDB, SQL Server clustered indexes) the *table rows themselves* are stored in key order, so random keys scatter the data too, and every secondary index carries a copy of the 16-byte key. Mitigations, in order of impact: 1. **Use time-ordered UUIDs.** Version 7 (or Hibernate's time-based `@UuidGenerator` style) puts a timestamp in the high bits, so values are roughly ascending and inserts append like a sequence while remaining globally unique. 2. **Store them natively.** PostgreSQL `uuid`, MySQL `binary(16)`, SQL Server `uniqueidentifier` — 16 bytes. A `char(36)` column is more than twice the size, comparisons are slower, and the index bloats accordingly. 3. **Consider a hybrid.** Some systems keep a bigint surrogate as the clustered/primary key for storage locality and a unique, indexed UUID as the external identifier. You pay for one extra unique index and gain a compact key everywhere else. ## Assigned natural keys The same lifecycle rules apply to a business key used as `@Id` — a country code, an ISIN. The extra hazard there is mutability: primary keys propagate into every foreign key, so a "stable" business identifier that turns out to change forces a cascade of updates. That is the standard argument for a surrogate key with the business key as a separate unique constraint. ## How to answer Lead with the lifecycle consequence (`id == null` no longer signals new, so merge must select), then the storage/index cost of randomness with the time-ordered fix, then the upside — ids before the database, idempotent creates, no enumeration. That ordering shows you understand both the ORM mechanics and the physical database behaviour.
- With an assigned id, how can Hibernate still tell a new entity from a detached one?Give the entity a boxed `@Version` field: a null version means it has never been flushed, so Hibernate treats it as transient. Alternatively implement Hibernate's `Persistable` interface and answer `isNew()` yourself, typically with a transient boolean flipped in a `@PostPersist`/`@PostLoad` callback. Either way you avoid the extra SELECT that merge would otherwise need.
- Are random UUID primary keys always a bad idea for write-heavy tables?Not always, but they are the expensive default. The cost is index locality: random values split pages and scatter rows on engines with clustered primary keys. A time-ordered UUID (v7 or Hibernate's time-based style) keeps global uniqueness while restoring append-mostly inserts, and storing it as a native 16-byte type removes the rest of the overhead. If the id must also be compact in many foreign keys, a bigint surrogate with a unique UUID column is a reasonable hybrid.
A sequential key is like filing papers at the back of an always-open drawer; a random UUID makes you open a different drawer in a different cabinet for every sheet.
saying these in an interview costs you the question
- Thinking Hibernate can still distinguish new from detached by a non-null id
- Believing merge() is free for entities with assigned ids
- Storing UUIDs as char(36) and calling the overhead negligible
- Treating an unguessable UUID as an access-control mechanism
- Assuming all UUID versions have the same index behaviour