skip to content

Identifier Generation Strategies

IDENTITY, SEQUENCE, TABLE, and UUID generation and their very different runtime costs. A staple interview question because the wrong pick silently disables JDBC batching or hammers the database for ids.

part ofHibernateoverview, primer and where to startread it →
on this pageshow

questions

5

What identifier generation strategies can JPA's @GeneratedValue use, how do they differ in the SQL they produce, and which one does strategy = AUTO actually pick under Hibernate?

level: juniorimportance: must knowfreq 75%

answer

  1. IDENTITY = id after insert; SEQUENCE = id before
  2. TABLE = counter row + lock, last resort
  3. UUID = no round trip, 16 bytes
  4. Hibernate 6 AUTO -> per-entity <table>_SEQ
  5. No sequences on the dialect -> falls back to table

basics

~20 s

IDENTITY uses an auto-increment column, so the id is known only after the INSERT. SEQUENCE calls a database sequence, so the id is known before. TABLE keeps counters in a table row. UUID generates one in the JVM. Hibernate's AUTO picks SEQUENCE, or UUID for UUID-typed ids.

solid answer

~50 s

- **IDENTITY** — an auto-increment/serial/identity column. The database assigns the value during the INSERT and Hibernate reads it back via generated keys, so the id exists only after the row is written. - **SEQUENCE** — `select nextval(...)` from a database sequence before the INSERT. The id is known up front, and with an allocation size Hibernate caches a block of ids in memory, so N inserts cost far fewer sequence calls. - **TABLE** — a counter row in a dedicated table, fetched with a locking select plus an update. Portable to engines without sequences, but the row is a contention point; a last resort. - **UUID** — generated in the JVM (JPA 3.1 added `GenerationType.UUID`; Hibernate also has `@UuidGenerator`). No database round trip at all. - **Assigned** — no `@GeneratedValue`; the application sets the id. Under Hibernate 6, `AUTO` maps to SEQUENCE for numeric ids (a per-entity sequence) and to UUID generation for UUID-typed ids; on a dialect without sequences it degrades to a table-backed generator.

code

java · 13 lines
java
@Id @GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;                       // auto-increment column

@Id
@GeneratedValue(strategy = GenerationType.SEQUENCE, generator = "order_seq")
@SequenceGenerator(name = "order_seq", sequenceName = "order_seq", allocationSize = 50)
private Long id;                       // database sequence, block of 50

@Id @GeneratedValue(strategy = GenerationType.UUID)
private UUID id;                       // generated in the JVM

@Id
private String code;                   // assigned by the application

go deeper

for a junior

Name the four strategies and the key difference: with IDENTITY the id arrives after the insert, with SEQUENCE before it.

for a middle

Add the SQL each produces, the fact that AUTO is provider-defined and resolves to a per-entity sequence in Hibernate 6, and the dialect fallback.

for a senior

Discuss selection per engine, allocation blocks, gaps and rollback semantics, and the operational consequences of the table generator's hot row.

for a principal

Frame it as an identity-space decision: who mints ids (database, application, other systems), what that costs in round trips and index locality, and whether ids are ever exposed externally.

## The four standard strategies `@GeneratedValue(strategy = ...)` tells the provider where the primary key comes from. The differences matter because they change *when* the id exists, which changes what Hibernate can defer. **IDENTITY.** Backed by an auto-increment column: MySQL `AUTO_INCREMENT`, PostgreSQL `serial`/`GENERATED ... AS IDENTITY`, SQL Server `IDENTITY`. The database allocates the value while executing the INSERT, and the driver returns it through `getGeneratedKeys()`. Because the persistence context is a map keyed by identifier, Hibernate must have an id the moment `persist()` returns a managed instance — so with IDENTITY it executes the INSERT immediately at `persist()` rather than at flush. **SEQUENCE.** A standalone database object producing monotonically increasing numbers, transaction-independent (a rolled-back transaction still consumes values). Hibernate calls `nextval` *before* the INSERT, so the id is available at `persist()` while the INSERT itself is queued until flush. Configured with `@SequenceGenerator(name, sequenceName, allocationSize)`. Supported by PostgreSQL, Oracle, DB2, H2, SQL Server 2012+, MariaDB 10.3+, MySQL 8.0 does not have them. **TABLE.** Emulates a sequence with a table (`hibernate_sequences` by default) holding one row per generator. Fetching a block means `select next_val ... for update` followed by an `update`, usually in a separate connection/transaction so it commits independently. It is fully portable, and universally slow: the counter row serialises every id fetch across the whole cluster and is a classic hotspot and deadlock source. Use only when the engine has no sequences and IDENTITY is unacceptable. **UUID.** Added as `GenerationType.UUID` in JPA 3.1. Hibernate produces an RFC 4122 value in the JVM — random (version 4) by default, with `@UuidGenerator(style = ...)` offering time-based styles. No round trip, ids exist before any database contact, and they can be created offline or in another service. The cost is storage and index behaviour: 16 bytes, and random values scatter B-tree inserts. **Assigned.** Omit `@GeneratedValue` and set the id yourself — a natural key, or a value from another system. ## What AUTO means `AUTO` says "provider's choice", which is why it is a common interview trap. Under Hibernate 6: - For a `UUID` (or UUID-like) id type, it uses UUID generation. - Otherwise it uses the sequence-style generator. Hibernate 6 defaults to a **per-entity** sequence named `<table>_SEQ`; Hibernate 5 used one shared `hibernate_sequence` for the whole application, which is why upgrades often need explicit `@SequenceGenerator` mappings or a migration. - On a dialect with no sequence support, the sequence-style generator falls back to a table-backed structure — i.e. AUTO quietly becomes TABLE on MySQL, which surprises teams who assume it becomes IDENTITY. Because AUTO's meaning shifts with provider, version and dialect, production mappings are usually explicit. ## Choosing On PostgreSQL/Oracle, SEQUENCE with a tuned allocation size is the default choice: ids up front, cheap allocation, and inserts can batch. On MySQL, IDENTITY is idiomatic — sequences are unavailable and the table generator is worse — at the price of no insert batching. UUIDs win when ids must be minted by clients or across services, when you want to build a whole object graph before touching the database, or when you want to avoid leaking row counts in public URLs; prefer a time-ordered variant and a native 16-byte column type. TABLE is a compatibility fallback, not a design choice. One more property distinguishes them: **gaps**. Sequences and allocation blocks are not gapless — rollbacks, restarts and cached blocks lose values. Never expose a generated surrogate id as an invoice or document number that auditors expect to be contiguous; that is a separate, deliberately serialised counter.

  • Your entity uses GenerationType.AUTO and the application runs on MySQL. What does Hibernate actually do?
    MySQL has no sequences, so Hibernate's sequence-style generator degrades to a table-backed generator — a counter row fetched with a locking select and an update. That is usually not what the team wanted; on MySQL you should map IDENTITY explicitly, or accept the table generator with a large allocation size if you need ids before insert.
  • Are generated identifier values guaranteed to be contiguous?
    No. Sequences are transaction-independent, so a rolled-back transaction consumes values permanently; allocation blocks held in memory are lost on restart; and IDENTITY columns skip values after failed inserts. Any requirement for gapless numbering — invoice numbers, legal document sequences — needs a separate serialised counter, not the surrogate primary key.

saying these in an interview costs you the question

  • Believing AUTO always means IDENTITY
  • Thinking generated ids are contiguous with no gaps
  • Assuming a sequence value is rolled back with the transaction
  • Using TABLE generation on a high-throughput table without realising it serialises inserts
  • Claiming MySQL supports database sequences

context

open as a page

What does the allocationSize attribute of JPA's @SequenceGenerator control in Hibernate, and what goes wrong if it disagrees with the INCREMENT BY of the underlying database sequence?

level: middleimportance: must knowfreq 45%

basics

~20 s

allocationSize is how many ids Hibernate reserves per sequence call and hands out from memory — default 50. If the database sequence increments by 1 while Hibernate assumes 50, the reserved block overlaps values a later call returns, producing duplicate primary keys.

open as a page

Why does using an auto-increment (IDENTITY) column for entity identifiers stop Hibernate from batching INSERT statements, and what changes if you switch to a database sequence?

level: middleimportance: must knowfreq 55%

basics

~20 s

With IDENTITY the id only exists once the row is inserted, but Hibernate needs an id to key the entity in the persistence context. So it executes the INSERT immediately at persist(), one statement per entity — nothing is queued, so nothing can batch. A sequence gives the id up front, letting inserts queue and batch at flush.

open as a page

You assign entity identifiers in application code — for example a UUID set in the constructor — instead of letting the database generate them. What does that change for Hibernate's handling of the entity, and what is the storage and index cost of random UUIDs?

level: seniorimportance: should knowfreq 35%

basics

~20 s

An id that is never null removes Hibernate's usual new-versus-detached signal, so merge() must SELECT before it knows whether to insert. Random UUIDs are 16 bytes and insert at random points in the primary-key index, causing page splits and poor cache locality; time-ordered UUIDs and a native 16-byte column fix most of that.

open as a page

Hibernate's sequence-backed id generators can use different optimizers — hi/lo, pooled, and pooled-lo. What does the value returned by the database sequence mean in each, and which are safe when another application inserts into the same table?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

With legacy hi/lo the sequence value is a multiplier, so it bears no relation to stored ids and outside writers collide. With pooled, nextval returns the top of the reserved block; with pooled-lo it returns the bottom. Both keep sequence values inside the id space, so external callers using the same sequence stay safe.

open as a page