skip to content

Why does an entity whose primary key is generated with @GeneratedValue(strategy = GenerationType.IDENTITY) not get its INSERT statements batched by Hibernate, and what would you map instead if batching matters?

level: middleimportance: must knowfreq 52%

answer

  1. persist() must know the id -> immediate INSERT
  2. insert batching disabled for IDENTITY; updates still batch
  3. SEQUENCE + allocationSize (pooled optimizer)
  4. allocationSize must match sequence INCREMENT BY
  5. AUTO resolves per dialect — be explicit

basics

~20 s

With IDENTITY the database assigns the key during the INSERT, and persist() must return a managed entity that already has its identifier, so Hibernate executes the INSERT immediately instead of queuing it for flush. Nothing accumulates, so nothing batches. Use a sequence with an allocation size instead.

solid answer

~50 s

The persistence context is keyed by entity type plus **identifier**, so `persist()` must know the id before it returns. With `GenerationType.IDENTITY` the value only exists once the row is inserted, so Hibernate must run the `INSERT` right there and read the generated key. The insert never reaches the flush-time action queue, and JDBC batching operates on that queue — so insert batching is effectively disabled for these entities, no matter what `hibernate.jdbc.batch_size` says. Subsequent `UPDATE`/`DELETE` for the same entities can still batch. The fix is an identifier strategy that Hibernate can resolve **before** the insert: - `GenerationType.SEQUENCE` with `@SequenceGenerator(allocationSize = 50)`, which uses a pooled optimizer so one sequence call covers 50 identifiers; - `GenerationType.TABLE` where no sequences exist (older MySQL), at the cost of contention on the generator row; - an application-assigned key, for example a UUID or a client-generated identifier. With a sequence the inserts queue up until flush and batch normally.

code

java · 10 lines
java
// no insert batching: the INSERT runs inside persist()
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;

// batches: the id is known before the row is written
@Id
@GeneratedValue(strategy = GenerationType.SEQUENCE, generator = "order_seq")
@SequenceGenerator(name = "order_seq", sequenceName = "order_seq", allocationSize = 50)
private Long id;

go deeper

for a junior

State that IDENTITY forces the insert to run immediately so there is nothing to batch, and that a sequence is the alternative.

for a middle

Explain why persist() needs the identifier, that updates still batch, and how allocationSize reduces sequence calls.

for a senior

Add the allocationSize/INCREMENT BY consistency requirement, table-generator contention, UUID index trade-offs, and how you would diagnose it.

for a principal

Frame identifier strategy as a write-throughput and schema decision spanning key width, index locality, allocation contention and portability.

## Why the identifier decides this Hibernate's persistence context is a map from `EntityKey` — entity name plus identifier — to the managed instance. JPA also specifies that after `persist()` the entity is managed, and for generated identifiers the value must be available at latest by flush; Hibernate's implementation needs it immediately to register the instance. Whatever produces the identifier therefore constrains *when* the row can be written. ## IDENTITY: the database assigns during the insert `GenerationType.IDENTITY` maps to an auto-increment column (`AUTO_INCREMENT`, `SERIAL`/`GENERATED … AS IDENTITY`). The value does not exist until the row is inserted; the only way to learn it is to execute the `INSERT` and read the generated key back (`Statement.getGeneratedKeys()` or a `returning` clause). So `persist()` cannot merely queue an action — it must execute the statement synchronously. That has two visible consequences: 1. **Inserts happen at `persist()`, not at flush.** Ordering with other statements is fixed earlier than you might expect, and the write happens even if you never call `flush()` explicitly. 2. **There is nothing to batch.** JDBC batching accumulates parameter sets against one prepared statement in the flush-time action queue. An immediately executed insert is a batch of one, every time. Hibernate is explicit about this: insert batching is disabled for entities using an identity generator, and older versions logged a warning to that effect. Note what is *not* affected: updates and deletes against those same entities are produced at flush like any other DML and do batch normally. It is only the insert path that is lost. Bidirectional-association fixups after the insert also still batch as updates. ## The alternative: get the id before the insert ### Sequence with an allocation size ```java @Id @GeneratedValue(strategy = GenerationType.SEQUENCE, generator = "order_seq") @SequenceGenerator(name = "order_seq", sequenceName = "order_seq", allocationSize = 50) private Long id; ``` Hibernate calls the sequence, receives a value, and — with `allocationSize` greater than 1 — uses a **pooled optimizer** so that one database call yields a block of 50 identifiers handed out in memory. The identifier is known before any row is written, so the insert is queued and flushed with everything else, and batching works. The sequence call itself is cheap and, importantly, does not have to happen per row. Two warnings about `allocationSize`. It must agree with the sequence's `INCREMENT BY` in the database, or Hibernate's pooled optimizer will hand out values that collide with other clients; Hibernate validates this when schema validation is on. And identifiers become non-contiguous — gaps appear whenever an application instance restarts with unused values in its block. Gaps are harmless unless somebody has (wrongly) attached business meaning to the number. ### TABLE generator Where sequences are unavailable, `GenerationType.TABLE` keeps a counter row in a table and can also allocate in blocks. It restores batching but introduces a hot row that every inserting transaction must update, so it serialises identifier allocation under concurrency. Prefer sequences when the database has them; MySQL gained sequence-like behaviour only via workarounds, which is why the IDENTITY/batching problem is most often met on MySQL. ### Application-assigned identifiers A UUID or client-supplied key removes the round trip entirely and batches fine. The cost is index behaviour: random UUIDs scatter inserts across the primary-key index, which hurts write locality and page cache; time-ordered variants mitigate this. Wider keys also enlarge every secondary index that references them. ## `GenerationType.AUTO` `AUTO` delegates to the dialect. On modern Hibernate 6 with PostgreSQL it resolves to a sequence, which is why the same code can batch on one database and not on another. If batching matters, be explicit rather than relying on `AUTO`. ## Diagnosing it The symptom is that `hibernate.jdbc.batch_size` appears to do nothing for inserts while updates in the same transaction batch fine, and that inserts show up in the log at the moment of `persist()` rather than at flush. Checking the identifier mapping is the first move; a proxying datasource that reports batch size confirms it. If the schema cannot change, at least accept the round trips and reduce them another way — fewer rows, fewer transactions — rather than turning `batch_size` up and expecting an effect.

  • Does IDENTITY also prevent UPDATE and DELETE statements from being batched for those entities?
    No. Only the insert is forced to execute early, because that is where the identifier is produced. Updates and deletes are collected in the action queue and emitted at flush like any other DML, so they batch normally subject to batch size and statement ordering.
  • What can go wrong if allocationSize on the sequence generator does not match the sequence's INCREMENT BY in the database?
    Hibernate's pooled optimizer assumes each sequence call reserves a block of that size and hands out the intermediate values in memory. If the database increments by 1 while Hibernate assumes 50, two application instances — or any other client — will be handed overlapping identifiers and inserts fail on the primary key. Schema validation reports the mismatch, which is one reason to keep it enabled.
  • Why are UUID primary keys not an automatic win even though they batch well?
    A random UUID scatters inserts across the whole primary-key index instead of appending near the right edge, which increases page splits and cache misses on write-heavy tables, and the wider key inflates every secondary index that carries it. Time-ordered UUID variants restore much of the locality, which is why they are usually the version worth choosing.

saying these in an interview costs you the question

  • Claiming batch_size fixes insert throughput for identity-generated entities
  • Not knowing that the INSERT runs inside persist() with IDENTITY
  • Switching to a sequence but leaving allocationSize at 1 and expecting fewer round trips
  • Setting allocationSize without changing the sequence's INCREMENT BY
  • Assuming IDENTITY also blocks update and delete batching

context