As a principal engineer, how would you decide when to use noRollback and how to size chunks, given the rollback-and-rescan cost model of fault-tolerant Spring Batch steps?
answer
- Failure cost ∝ chunk size (N-way rescan)
- Clean data -> big chunks; dirty -> small
- noRollback: processor-phase, no persisted state, skip-expected
- Writer exceptions always roll back — never rely on noRollback there
- Idempotent writer + outbox = precondition; skipLimit bounds blast radius
basics
~20 sUse noRollback only for processor-phase exceptions where nothing transactional was written and re-scanning is wasted work — typically validation skips. Size chunks by balancing happy-path throughput against the cost of rolling back and re-scanning a whole chunk on failure, given your data's error rate and writer idempotency.
solid answer
~50 sThe core trade-off: a chunk transaction gives you bulk-write throughput, but a failure rolls back the whole chunk and re-scans it item-by-item, so failures are expensive and proportional to chunk size. My decision framework: (1) Estimate the expected error/skip rate from the data source. High error rates favor smaller chunks so each rollback wastes less and re-scans fewer items; clean data favors large chunks for throughput. (2) Ensure the writer is idempotent (upsert/dedup) because rescanning replays items — this is a precondition, not optional. (3) Apply noRollback narrowly to processor-phase business exceptions (e.g. ValidationException) that have no persisted state to undo, so a validation skip doesn't trigger a rollback + full rescan; never rely on it for writer exceptions, which always roll back. (4) Push non-idempotent side effects to an outbox/idempotent step. (5) Set skipLimit/retryLimit to bound blast radius and fail loudly on systemic issues rather than silently skipping.
code
java · 23 lines@Bean
public Step ingestStep(JobRepository jobRepository,
PlatformTransactionManager txManager,
ItemReader<Record> reader,
ItemProcessor<Record, Row> processor, // pure transform
ItemWriter<Row> upsertWriter, // idempotent
SkipListener<Record, Row> quarantineListener) {
return new StepBuilder("ingestStep", jobRepository)
.<Record, Row>chunk(200, txManager) // large: data is trusted
.reader(reader)
.processor(processor)
.writer(upsertWriter)
.faultTolerant()
// expected business rejection: skip WITHOUT rolling back the chunk
.skip(ValidationException.class)
.noRollback(ValidationException.class)
// transient faults: retry (these DO roll back for a clean attempt)
.retry(TransientDataAccessException.class)
.retryLimit(3)
.skipLimit(100) // fail loudly if corruption is systemic
.listener(quarantineListener) // dead-letter skipped records
.build();
}go deeper
Not expected; can note chunk size trades throughput vs failure cost.
Knows noRollback is for expected processor rejections and that writers must be idempotent.
Balances chunk size, skip/retry limits, and noRollback coherently with idempotency.
Designs the whole coupled system — throughput, failure cost, idempotency, dead-lettering, and observability — and sets team-wide conventions.
## The cost model to internalize A fault-tolerant chunk step has two regimes: - **Happy path (bulk)**: read N, process N, one bulk write, one commit. Throughput scales with N. - **Failure path (rollback + scan)**: on any chunk exception it rolls back all N, then replays the buffered items **one at a time** (N single-item transactions) to isolate and skip/retry the bad one. So the marginal cost of a failure is roughly *proportional to chunk size*. Large chunks amortize commit overhead beautifully when data is clean, but each poison record triggers an N-way rescan. This is the tension every knob below manages. ## Chunk sizing decision - **Clean, trusted data / rare failures** → larger chunks (hundreds+) for throughput; the occasional rescan is negligible. - **Dirty inbound data / high skip rate** → smaller chunks so each rollback discards and re-scans fewer items, and commit checkpoints are more frequent (better restartability). - **Expensive writes (network round-trips)** → bias larger, but only if idempotent, since rescans re-write. - Always pair with realistic load testing that includes the failure path, not just the happy path. ## When noRollback earns its place Use `.noRollback(Type)` when the exception: 1. is thrown in the **processor/read/listener phase** (not the writer), 2. has **no persisted transactional state** to undo, and 3. represents a routine, expected outcome (business/validation rejection) you'll **skip**. In that case rolling back and re-scanning the whole chunk is pure waste; noRollback lets the good items commit in bulk while the rejected item is skipped. Do **not** register writer exceptions — Batch overrides you and rolls back, because the writer may have partially mutated transactional state. Registering it anyway is a correctness trap that gives false confidence. ## Idempotency is a precondition, not a feature Because rollback triggers replay, the **writer must be idempotent** (upserts, dedup keys, idempotency tokens) and the **processor must be pure**. Non-idempotent external effects belong in an **outbox** consumed by a separate idempotent step, or behind an idempotency key the downstream honors. Skip/retry without idempotency is a latent duplicate-data bug. ## Bounding blast radius - `skipLimit` / `retryLimit` cap how much the step tolerates before failing. Set them to distinguish *expected sparse bad records* from *systemic corruption* — you want the job to fail loudly if 40% of records are bad, not silently skip them into a black hole. - Emit metrics on skip counts and rescan frequency; a rising skip rate is an upstream data-quality signal. - Consider `SkipListener` to route skipped items to a dead-letter/quarantine table for later inspection instead of losing them. ## Interplay with retry Retry re-attempts an item (each attempt can roll back) up to `retryLimit`; a still-failing, skippable exception can then be skipped. Retryable transient faults (deadlocks, timeouts) generally *should* roll back (start clean), so noRollback rarely applies to retryable types. Reserve noRollback for terminal, non-transactional, skippable business rejections. ## The principal-level synthesis Treat chunk size, skip/retry limits, noRollback, and writer idempotency as one coupled design: chunk size sets throughput and failure cost, idempotency makes rescans safe, noRollback trims wasted rollbacks for expected rejections, and limits + listeners bound and observe the blast radius. Get those aligned and a batch job is both fast on the happy path and safe under partial failure.
- Why would you keep chunks small when the input data is known to be dirty?Because a failure rolls back and re-scans the entire chunk item-by-item; with many bad records, large chunks mean frequent expensive N-way rescans and more wasted bulk work. Smaller chunks bound the rescan cost per failure and give more frequent commit checkpoints for restart.
- Would you apply noRollback to a transient deadlock exception you're retrying?No. Retryable transient faults should roll back so the retry starts from a clean transaction. noRollback fits terminal, non-transactional, skippable business rejections in the processor phase, not retryable transient DB errors.
- How do you avoid silently losing skipped records?Register a SkipListener that writes skipped items (and the cause) to a quarantine/dead-letter table, and set skipLimit so a systemic failure aborts the job rather than skipping everything into oblivion.
saying these in an interview costs you the question
- Using noRollback as a blanket 'don't fail my step' switch across all exception types
- Relying on noRollback for writer exceptions
- Choosing huge chunk sizes without accounting for rescan cost or writer idempotency
- Setting an unbounded/huge skipLimit that hides systemic data corruption