How do you choose a chunk (commit-interval) size, and what are the tradeoffs of too-large vs too-small values?
answer
- small = commit overhead, no batching
- large = memory, locks, redo, slow skip-scan
- start 100–1000, then measure
- match writer's JDBC batch size
- variable cost → CompletionPolicy not fixed count
basics
~20 sBigger chunks mean fewer commits and faster throughput but more memory, longer transactions, and more work re-done on failure. Smaller chunks mean more commit overhead but lower memory and cheaper retries. You tune empirically for your data and DB.
solid answer
~40 sChunk size trades commit overhead against transaction cost. Too small (e.g. 1–10): a commit per handful of rows, so commit/round-trip overhead dominates and throughput tanks, and batch writers can't amortize JDBC batching. Too large (e.g. 100k): a huge transaction holding locks longer, high heap use (the whole chunk plus processor outputs sit in memory), bigger rollback and more re-processing on failure, and in a skip-configured step, a rollback triggers slow one-by-one replay of the whole chunk to isolate the bad item. There's no universal number — typical starting points are 100–1000, tuned by measuring throughput, memory, and DB lock contention. Factors: item size, processor cost, writer's batch capability, DB lock behavior, and restart/skip requirements. For highly variable item cost, use a CompletionPolicy instead of a fixed count.
go deeper
Might just say 'pick a reasonable number like 100'.
Knows bigger = fewer commits/faster but more memory.
Articulates both cost curves and ties size to writer batching and memory.
Frames it as an empirical tuning decision balancing throughput, memory, lock contention, restart granularity, and skip-replay cost; knows when to switch to a CompletionPolicy.
**What the number does.** The commit-interval / chunk size sets how many items are read, buffered, and written per transaction. It is the single most impactful throughput knob in a chunk step, and it is a genuine tradeoff with no correct default. **Costs of TOO SMALL (e.g. 1–10):** - **Commit overhead dominates.** Every chunk means a transaction begin/commit round trip plus a `JobRepository` update. At chunk size 1 you effectively pay per-row commit cost — the very thing chunking exists to avoid. - **No write batching.** Bulk writers (`JdbcBatchItemWriter`, JPA batch) can't amortize network/round-trip cost across items, so DB throughput collapses. - **JobRepository churn.** Frequent step-execution updates add DB write load of their own. **Costs of TOO LARGE (e.g. 50k–100k+):** - **Memory pressure.** The framework holds the whole chunk of read items plus the processor's output list in heap simultaneously. Large items × large chunk = OOM risk. - **Long transactions / lock contention.** The chunk's transaction stays open across all reads, processing, and the write, holding DB locks longer and increasing contention and deadlock risk, and stressing the undo/redo log. - **Expensive failure.** A failure late in a giant chunk rolls back the entire chunk; all that read+process work is redone. Bigger chunk = more wasted work per failure. - **Slow skip replay.** In a fault-tolerant step, when a chunk rolls back due to a skippable exception, Spring Batch re-runs the chunk **item-by-item** (scanning) to find and skip the offending item. With a huge chunk that scan is very slow. - **Coarser restart granularity.** Restart resumes at the last committed chunk boundary; larger chunks mean more reprocessing after a crash. **How to actually pick.** 1. **Start in the 100–1000 range** — a common sweet spot for row-oriented DB work. 2. **Measure** throughput (items/sec), peak heap, transaction duration, and DB lock/wait metrics under production-like volume. 3. **Increase** until throughput gains flatten or memory/lock metrics degrade — that plateau is your ceiling. 4. **Account for item size and processor cost.** Large payloads or heavy per-item processing push the size down; tiny uniform rows allow larger. 5. **Match the writer.** If using `JdbcBatchItemWriter`, align chunk size with a healthy JDBC batch size and the driver's `reWriteBatchedInserts`-style options. 6. **Consider restart/skip SLAs.** Tighter recovery requirements favor smaller chunks (less redo, faster skip scan). **Variable-cost data.** When per-item cost/size varies wildly, a fixed count is a poor proxy for transaction cost — switch to a `CompletionPolicy` (e.g. size by bytes, or `CompositeCompletionPolicy` of count + timeout) so chunk boundaries track real cost/latency. **Anti-patterns.** - Copy-pasting `chunk(10)` from a tutorial into a million-row job (commit overhead kills it). - Cranking to `chunk(100000)` 'for speed' and then OOMing or deadlocking in prod. - Ignoring that a fault-tolerant step's rollback replays the whole chunk one-by-one — a hidden cost of large chunks with skips. **Interview framing.** A principal-level answer names the two opposing cost curves (commit overhead vs transaction/memory/replay cost), gives a concrete starting range, insists on measurement, and ties the choice to fault-tolerance and restart requirements — not just raw throughput.
- Why can a large chunk size be especially costly in a fault-tolerant (skip-enabled) step?On a skippable exception the chunk rolls back and Spring Batch re-processes it item-by-item to isolate the bad item. A larger chunk means a much longer one-at-a-time scan, plus more redone read/process work.
- How does chunk size interact with restart after a crash?Restart resumes at the last committed chunk boundary. Larger chunks mean more items were in-flight and uncommitted, so more work is reprocessed on restart; smaller chunks give finer-grained recovery.
saying these in an interview costs you the question
- 'Bigger is always faster' — ignores memory, locks, and replay cost
- Picking chunk size with no measurement
- Ignoring that the whole chunk (reads + processor outputs) sits in heap
- Not aligning chunk size with the writer's batching capability
- Overlooking the item-by-item skip scan cost on rollback