Compare JdbcBatchItemWriter and JpaItemWriter: how each persists a chunk, and when you would choose one over the other.
answer
- Jdbc: one SQL, addBatch/executeBatch, no persistence context
- assertUpdates checks affected rows
- Jpa: merge (or persist) each + single flush
- JPA needs hibernate.jdbc.batch_size to truly batch
- JDBC=raw speed; JPA=cascade/lifecycle
basics
~20 sJdbcBatchItemWriter runs a single parameterized SQL statement as a JDBC batch for the whole chunk. JpaItemWriter persists or merges each entity through an EntityManager and flushes once. Choose JDBC for raw speed on plain SQL, JPA when you already work with managed entities.
solid answer
~40 sJdbcBatchItemWriter takes one SQL string plus an ItemSqlParameterSourceProvider (or a prepared-statement setter) and executes the chunk as a single JdbcTemplate batch (addBatch/executeBatch). It optionally asserts the affected-row count (assertUpdates=true) to detect no-op updates. It bypasses the persistence context, so it's fast and memory-light but you write SQL and it doesn't cascade or manage entity state. JpaItemWriter uses the transaction-bound EntityManager: for each item it calls merge() (or persist() when usePersist=true), then flush() once per chunk. It gives you cascades, dirty checking and entity lifecycle but costs more memory and can trigger N selects on merge. Rule of thumb: bulk inserts/updates of flat rows -> JDBC; domain graphs already mapped with JPA and you need cascade/lifecycle -> JPA.
code
java · 19 lines// JDBC: one statement batched for the whole chunk
@Bean
JdbcBatchItemWriter<Trade> jdbcWriter(DataSource ds) {
return new JdbcBatchItemWriterBuilder<Trade>()
.dataSource(ds)
.sql("INSERT INTO trade (id, price) VALUES (:id, :price)")
.beanMapped() // BeanPropertyItemSqlParameterSourceProvider
.assertUpdates(true)
.build();
}
// JPA: merge/persist each entity, flush once per chunk
@Bean
JpaItemWriter<Trade> jpaWriter(EntityManagerFactory emf) {
return new JpaItemWriterBuilder<Trade>()
.entityManagerFactory(emf)
.usePersist(true) // persist instead of merge for new entities
.build();
}go deeper
Can say one writes via SQL and one via JPA entities.
Should describe batched SQL vs merge+flush and give a basic choose-one rule.
Should cover assertUpdates, persist vs merge, transaction-manager requirements and the hibernate.jdbc.batch_size gotcha.
Should reason about throughput/memory trade-offs, exactly-once/idempotency of the DML, and coordinating transaction managers across mixed writers.
**JdbcBatchItemWriter<T>.** - You configure a single DML statement, e.g. `INSERT INTO trade (id, price) VALUES (:id, :price)`. - Binding: an `ItemSqlParameterSourceProvider` maps each item to named params (commonly `BeanPropertyItemSqlParameterSourceProvider`, which reads bean properties), or an `ItemPreparedStatementSetter` for positional `?` params. - Execution: it uses `NamedParameterJdbcTemplate.batchUpdate(...)`, i.e. one JDBC batch (`addBatch` then `executeBatch`) for the entire chunk — the essence of batched I/O. - `assertUpdates` (default true): after the batch it checks each statement affected at least one row and throws `EmptyResultDataAccessException` otherwise — great for catching updates whose WHERE matched nothing. Turn it off for upserts/inserts where zero rows can be legitimate. - Requires a `DataSource`. It does **not** touch a persistence context — no dirty checking, no cascades, no identity map. That's what makes it fast and low-memory. **JpaItemWriter<T>.** - Requires an `EntityManagerFactory`. It obtains the transaction-bound `EntityManager`. - Per item it calls `entityManager.merge(item)` by default, or `persist(item)` if `usePersist(true)` (persist is cheaper for guaranteed-new entities because merge may issue a SELECT to load current state). - After iterating the chunk it calls `entityManager.flush()` once, so SQL is sent at chunk boundary (JPA/Hibernate batching still needs `hibernate.jdbc.batch_size` set to truly batch at the JDBC level). - You get cascades, `@GeneratedValue`, optimistic locking, lifecycle callbacks — full ORM semantics. - Memory/perf cost: managed entities accumulate in the persistence context; merge can cause extra selects; keep chunk sizes moderate and rely on the transaction clearing context between chunks. **HibernateItemWriter** is the analogous writer for a raw Hibernate `SessionFactory` (save/update/merge + flush/clear). **Transactions.** Both write inside the step's chunk transaction. With JPA the transaction manager must be `JpaTransactionManager` (or JTA) so the EntityManager is joined; with JDBC a `DataSourceTransactionManager` suffices. Mixing writers of different resources needs a shared/coordinated transaction manager. **Choosing.** - Choose **JdbcBatchItemWriter** for: high-volume inserts/updates of tabular data, ETL where you don't need the domain model, best throughput and lowest memory, and when you want explicit SQL/upsert control. - Choose **JpaItemWriter** for: writing already-mapped JPA entities, needing cascade persistence of object graphs, entity lifecycle/validation, or reusing existing repository mappings. Accept the overhead. **Gotchas.** - JDBC `assertUpdates` surprising failures on updates that legitimately match zero rows. - JPA batching won't actually batch at JDBC level unless `hibernate.jdbc.batch_size` is configured; otherwise you get N round-trips despite the single flush. - JpaItemWriter merge returns a new managed instance but the writer discards it — don't rely on the passed item becoming managed. - Very large chunks with JPA bloat the persistence context.
- You configured JpaItemWriter and set a large chunk size but throughput is poor and memory grows. What's likely wrong?Two things: (1) JDBC-level batching isn't enabled — set hibernate.jdbc.batch_size (and order_inserts/order_updates) so the single flush actually batches instead of N round-trips; (2) the persistence context accumulates managed entities within the chunk, so an over-large chunk bloats memory. Reduce chunk size and enable batching.
- When would you turn off assertUpdates on JdbcBatchItemWriter?When zero affected rows is legitimate — e.g. inserts, upserts, or updates whose WHERE may not match. With assertUpdates=true the writer throws EmptyResultDataAccessException if a statement affects no rows, which would falsely fail the step.
- Why prefer persist over merge in JpaItemWriter for new entities?merge may issue a SELECT to load the current DB state before copying, and returns a new managed instance. For entities known to be new, persist avoids that extra select and is cheaper. Set usePersist(true).
saying these in an interview costs you the question
- Claiming JdbcBatchItemWriter goes through the JPA persistence context
- Thinking JpaItemWriter batches at JDBC level automatically without hibernate.jdbc.batch_size
- Saying JPA is always faster because it's 'higher level'
- Not knowing assertUpdates exists or what it guards