You configured a JDBC batch size in Hibernate, but a transaction writing several different entity types still produces tiny batches. What do the hibernate.order_inserts and hibernate.order_updates settings do about that, and what do they cost?
answer
- batch = one SQL string; interleaving kills it
- order_inserts: group by entity type at flush
- order_updates: group by type + sort by PK
- both default false
- PK ordering also reduces deadlocks
basics
~20 sA batch can only hold one SQL string, so interleaved entity types break it after every statement. order_inserts and order_updates sort the flush-time actions by entity type (and updates by primary key) so identical statements become contiguous and fill batches. Cost is sorting at flush plus changed statement order.
solid answer
~50 sA JDBC batch belongs to a single `PreparedStatement`, so a change of SQL string flushes the current batch. If your flush emits `insert into orders`, `insert into order_lines`, `insert into orders`, … you get batches of one however large `hibernate.jdbc.batch_size` is. **`hibernate.order_inserts=true`** sorts pending inserts by entity type before executing them, while preserving dependency order between types, so all `orders` rows go out together, then all `order_lines`. **`hibernate.order_updates=true`** does the same for updates and additionally orders them by primary key. That second effect is valuable independently of batching: several transactions updating the same rows now touch them in a consistent order, which removes a common source of deadlocks. Both default to **false**, so batching in a mixed workload is usually half-configured without them. Costs: a sort of the action queue at flush, and statements no longer following the order in which your code made changes — which matters if you relied on that order, and needs care with self-referencing foreign keys.
code
properties · 3 lineshibernate.jdbc.batch_size=50
hibernate.order_inserts=true
hibernate.order_updates=truego deeper
Know that a batch holds one kind of statement and that Hibernate has settings to group similar statements together.
Name both properties, that they default to false, and why interleaved entity types otherwise yield batches of one.
Add the deadlock-reduction effect of primary-key ordering, the self-referencing-insert risk, and how you would verify batch sizes.
Position ordering within a write-path configuration standard — batch size, ordering, identifier strategy — and the ordering assumptions it invalidates.
## The constraint that creates the problem JDBC batching accumulates parameter sets against one `PreparedStatement`. Hibernate can therefore keep adding to a batch only while consecutive statements share the same SQL string. Any change — different table, different operation, even a different column list for the same table — forces the current batch to be executed and a new one to be started. Realistic write paths interleave. Persisting a hundred orders each with three lines, in natural object order, produces `insert orders`, `insert order_lines`, `insert order_lines`, `insert order_lines`, `insert orders`, … The SQL string alternates constantly, so despite `hibernate.jdbc.batch_size=50` the effective batch size hovers near one and the setting looks broken. ## What the ordering settings do ### `hibernate.order_inserts` When `true`, Hibernate sorts the queued insert actions by entity type before executing the flush. All inserts for one entity type become contiguous, so a single prepared statement can absorb up to `batch_size` of them. The sort is not naive. Hibernate's insert sorter accounts for dependencies between entity types — if `order_lines` carries a foreign key to `orders`, parents must be inserted before children — and orders the type groups accordingly. Where that analysis has limits is **self-referencing** associations, for example a category tree whose rows reference other rows of the same table: within one type group Hibernate cannot always guarantee parent-before-child, and a deferred-constraint-free schema can reject the ordering. When that bites, insert the roots in a separate flush, or make the constraint deferrable if the database supports it. ### `hibernate.order_updates` When `true`, pending updates are sorted by entity type and, within a type, by primary key. The batching benefit is the same as for inserts. The second benefit is subtler and often the real reason to enable it: **deterministic row ordering reduces deadlocks**. Two concurrent transactions that both update rows 7 and 9 will now both take 7 first; without ordering, one may take 7 then 9 while the other takes 9 then 7, and they wait on each other until the database kills one. Ordering does not eliminate deadlocks — different transactions touching different tables in different orders still can — but it removes a common, easily-avoided class. A related setting, `hibernate.order_deletes` (or delete-ordering under `hibernate.order_updates` in some versions), applies the same idea to deletes; deletes are also where foreign-key ordering constraints bite hardest. ## Both default to false This is the most useful fact to carry into an interview. `hibernate.jdbc.batch_size` alone is a half-configuration; for a workload that writes more than one entity type per transaction, the ordering flags are what turn it into real batches. A minimal batching configuration is therefore three properties, not one. ## What it costs - **Sorting work at flush.** Proportional to the number of pending actions; negligible next to the round trips it saves, but not free for enormous flushes. - **Statement order changes.** SQL no longer follows the order of your object manipulations. Code that depended on that ordering — a trigger with side effects, an audit expectation, a deliberately sequenced pair of writes — can behave differently. Hibernate still respects the coarse ordering rules (inserts, then updates, then deletes, in its defined action-type order), so this mainly affects ordering *within* a phase. - **Self-referencing inserts.** As above, the one genuine correctness edge to test. Because the risks are ordering-shaped rather than data-shaped, the safe way to introduce these flags is to enable them and run the write-path integration tests, watching for constraint violations rather than for slowness. ## How to confirm the improvement Statement logging will look identical before and after, since the same statements are prepared either way. Confirm with a proxying datasource or driver-level logging that reports batch size, or by counting round trips at the database. The expected signature after enabling ordering is a small number of large batches per flush instead of many batches of one. ## Interaction with identifier generation Ordering cannot rescue inserts that never reach the flush queue. If the entity uses an identity/auto-increment key, each insert already executed inside `persist()`, so there is nothing left to sort. Fix the identifier strategy first, then ordering and batch size become meaningful.
- Besides larger batches, what other production benefit does ordering updates by primary key give?Fewer deadlocks. When every transaction updates rows of a table in ascending primary-key order, two concurrent transactions touching the same rows acquire them in the same sequence and one simply waits instead of forming a cycle. It does not remove deadlocks caused by different tables being touched in different orders, but it eliminates a frequent and otherwise puzzling class of them.
- Is there a case where enabling order_inserts can cause a failure?Yes, with self-referencing foreign keys. Hibernate orders insert groups by entity type and resolves dependencies between types, but within one type — a category whose parent is another category — the reordering can place a child before its parent and the constraint rejects it. Making the constraint deferrable, or inserting roots in a separate flush, resolves it; the write-path integration tests are where you should catch it.
saying these in an interview costs you the question
- Assuming `hibernate.jdbc.batch_size` alone is sufficient for a mixed workload
- Thinking one batch can span statements against different tables
- Believing the ordering settings default to true
- Missing that ordering updates by primary key also reduces deadlocks
- Concluding batching improved because the SQL log looks the same length