skip to content

How do you choose a chunk (commit-interval) size, and what are the tradeoffs of too-large vs too-small values?

level: principalimportance: should knowfreq 45%

answer

  1. small = commit overhead, no batching
  2. large = memory, locks, redo, slow skip-scan
  3. start 100–1000, then measure
  4. match writer's JDBC batch size
  5. variable cost → CompletionPolicy not fixed count

basics

~20 s

Bigger chunks mean fewer commits and faster throughput but more memory, longer transactions, and more work re-done on failure. Smaller chunks mean more commit overhead but lower memory and cheaper retries. You tune empirically for your data and DB.

solid answer

~40 s

Chunk size trades commit overhead against transaction cost. Too small (e.g. 1–10): a commit per handful of rows, so commit/round-trip overhead dominates and throughput tanks, and batch writers can't amortize JDBC batching. Too large (e.g. 100k): a huge transaction holding locks longer, high heap use (the whole chunk plus processor outputs sit in memory), bigger rollback and more re-processing on failure, and in a skip-configured step, a rollback triggers slow one-by-one replay of the whole chunk to isolate the bad item. There's no universal number — typical starting points are 100–1000, tuned by measuring throughput, memory, and DB lock contention. Factors: item size, processor cost, writer's batch capability, DB lock behavior, and restart/skip requirements. For highly variable item cost, use a CompletionPolicy instead of a fixed count.

go deeper

for a junior

Might just say 'pick a reasonable number like 100'.

for a middle

Knows bigger = fewer commits/faster but more memory.

for a senior

Articulates both cost curves and ties size to writer batching and memory.

for a principal

Frames it as an empirical tuning decision balancing throughput, memory, lock contention, restart granularity, and skip-replay cost; knows when to switch to a CompletionPolicy.

**What the number does.** The commit-interval / chunk size sets how many items are read, buffered, and written per transaction. It is the single most impactful throughput knob in a chunk step, and it is a genuine tradeoff with no correct default. **Costs of TOO SMALL (e.g. 1–10):** - **Commit overhead dominates.** Every chunk means a transaction begin/commit round trip plus a `JobRepository` update. At chunk size 1 you effectively pay per-row commit cost — the very thing chunking exists to avoid. - **No write batching.** Bulk writers (`JdbcBatchItemWriter`, JPA batch) can't amortize network/round-trip cost across items, so DB throughput collapses. - **JobRepository churn.** Frequent step-execution updates add DB write load of their own. **Costs of TOO LARGE (e.g. 50k–100k+):** - **Memory pressure.** The framework holds the whole chunk of read items plus the processor's output list in heap simultaneously. Large items × large chunk = OOM risk. - **Long transactions / lock contention.** The chunk's transaction stays open across all reads, processing, and the write, holding DB locks longer and increasing contention and deadlock risk, and stressing the undo/redo log. - **Expensive failure.** A failure late in a giant chunk rolls back the entire chunk; all that read+process work is redone. Bigger chunk = more wasted work per failure. - **Slow skip replay.** In a fault-tolerant step, when a chunk rolls back due to a skippable exception, Spring Batch re-runs the chunk **item-by-item** (scanning) to find and skip the offending item. With a huge chunk that scan is very slow. - **Coarser restart granularity.** Restart resumes at the last committed chunk boundary; larger chunks mean more reprocessing after a crash. **How to actually pick.** 1. **Start in the 100–1000 range** — a common sweet spot for row-oriented DB work. 2. **Measure** throughput (items/sec), peak heap, transaction duration, and DB lock/wait metrics under production-like volume. 3. **Increase** until throughput gains flatten or memory/lock metrics degrade — that plateau is your ceiling. 4. **Account for item size and processor cost.** Large payloads or heavy per-item processing push the size down; tiny uniform rows allow larger. 5. **Match the writer.** If using `JdbcBatchItemWriter`, align chunk size with a healthy JDBC batch size and the driver's `reWriteBatchedInserts`-style options. 6. **Consider restart/skip SLAs.** Tighter recovery requirements favor smaller chunks (less redo, faster skip scan). **Variable-cost data.** When per-item cost/size varies wildly, a fixed count is a poor proxy for transaction cost — switch to a `CompletionPolicy` (e.g. size by bytes, or `CompositeCompletionPolicy` of count + timeout) so chunk boundaries track real cost/latency. **Anti-patterns.** - Copy-pasting `chunk(10)` from a tutorial into a million-row job (commit overhead kills it). - Cranking to `chunk(100000)` 'for speed' and then OOMing or deadlocking in prod. - Ignoring that a fault-tolerant step's rollback replays the whole chunk one-by-one — a hidden cost of large chunks with skips. **Interview framing.** A principal-level answer names the two opposing cost curves (commit overhead vs transaction/memory/replay cost), gives a concrete starting range, insists on measurement, and ties the choice to fault-tolerance and restart requirements — not just raw throughput.

  • Why can a large chunk size be especially costly in a fault-tolerant (skip-enabled) step?
    On a skippable exception the chunk rolls back and Spring Batch re-processes it item-by-item to isolate the bad item. A larger chunk means a much longer one-at-a-time scan, plus more redone read/process work.
  • How does chunk size interact with restart after a crash?
    Restart resumes at the last committed chunk boundary. Larger chunks mean more items were in-flight and uncommitted, so more work is reprocessed on restart; smaller chunks give finer-grained recovery.

saying these in an interview costs you the question

  • 'Bigger is always faster' — ignores memory, locks, and replay cost
  • Picking chunk size with no measurement
  • Ignoring that the whole chunk (reads + processor outputs) sits in heap
  • Not aligning chunk size with the writer's batching capability
  • Overlooking the item-by-item skip scan cost on rollback

context