skip to content

As a principal engineer, how would you decide when to use noRollback and how to size chunks, given the rollback-and-rescan cost model of fault-tolerant Spring Batch steps?

level: principalimportance: nice to knowfreq 22%

answer

  1. Failure cost ∝ chunk size (N-way rescan)
  2. Clean data -> big chunks; dirty -> small
  3. noRollback: processor-phase, no persisted state, skip-expected
  4. Writer exceptions always roll back — never rely on noRollback there
  5. Idempotent writer + outbox = precondition; skipLimit bounds blast radius

basics

~20 s

Use noRollback only for processor-phase exceptions where nothing transactional was written and re-scanning is wasted work — typically validation skips. Size chunks by balancing happy-path throughput against the cost of rolling back and re-scanning a whole chunk on failure, given your data's error rate and writer idempotency.

solid answer

~50 s

The core trade-off: a chunk transaction gives you bulk-write throughput, but a failure rolls back the whole chunk and re-scans it item-by-item, so failures are expensive and proportional to chunk size. My decision framework: (1) Estimate the expected error/skip rate from the data source. High error rates favor smaller chunks so each rollback wastes less and re-scans fewer items; clean data favors large chunks for throughput. (2) Ensure the writer is idempotent (upsert/dedup) because rescanning replays items — this is a precondition, not optional. (3) Apply noRollback narrowly to processor-phase business exceptions (e.g. ValidationException) that have no persisted state to undo, so a validation skip doesn't trigger a rollback + full rescan; never rely on it for writer exceptions, which always roll back. (4) Push non-idempotent side effects to an outbox/idempotent step. (5) Set skipLimit/retryLimit to bound blast radius and fail loudly on systemic issues rather than silently skipping.

code

java · 23 lines
java
@Bean
public Step ingestStep(JobRepository jobRepository,
                       PlatformTransactionManager txManager,
                       ItemReader<Record> reader,
                       ItemProcessor<Record, Row> processor,   // pure transform
                       ItemWriter<Row> upsertWriter,           // idempotent
                       SkipListener<Record, Row> quarantineListener) {
    return new StepBuilder("ingestStep", jobRepository)
            .<Record, Row>chunk(200, txManager)   // large: data is trusted
            .reader(reader)
            .processor(processor)
            .writer(upsertWriter)
            .faultTolerant()
            // expected business rejection: skip WITHOUT rolling back the chunk
            .skip(ValidationException.class)
            .noRollback(ValidationException.class)
            // transient faults: retry (these DO roll back for a clean attempt)
            .retry(TransientDataAccessException.class)
            .retryLimit(3)
            .skipLimit(100)   // fail loudly if corruption is systemic
            .listener(quarantineListener)  // dead-letter skipped records
            .build();
}

go deeper

for a junior

Not expected; can note chunk size trades throughput vs failure cost.

for a middle

Knows noRollback is for expected processor rejections and that writers must be idempotent.

for a senior

Balances chunk size, skip/retry limits, and noRollback coherently with idempotency.

for a principal

Designs the whole coupled system — throughput, failure cost, idempotency, dead-lettering, and observability — and sets team-wide conventions.

## The cost model to internalize A fault-tolerant chunk step has two regimes: - **Happy path (bulk)**: read N, process N, one bulk write, one commit. Throughput scales with N. - **Failure path (rollback + scan)**: on any chunk exception it rolls back all N, then replays the buffered items **one at a time** (N single-item transactions) to isolate and skip/retry the bad one. So the marginal cost of a failure is roughly *proportional to chunk size*. Large chunks amortize commit overhead beautifully when data is clean, but each poison record triggers an N-way rescan. This is the tension every knob below manages. ## Chunk sizing decision - **Clean, trusted data / rare failures** → larger chunks (hundreds+) for throughput; the occasional rescan is negligible. - **Dirty inbound data / high skip rate** → smaller chunks so each rollback discards and re-scans fewer items, and commit checkpoints are more frequent (better restartability). - **Expensive writes (network round-trips)** → bias larger, but only if idempotent, since rescans re-write. - Always pair with realistic load testing that includes the failure path, not just the happy path. ## When noRollback earns its place Use `.noRollback(Type)` when the exception: 1. is thrown in the **processor/read/listener phase** (not the writer), 2. has **no persisted transactional state** to undo, and 3. represents a routine, expected outcome (business/validation rejection) you'll **skip**. In that case rolling back and re-scanning the whole chunk is pure waste; noRollback lets the good items commit in bulk while the rejected item is skipped. Do **not** register writer exceptions — Batch overrides you and rolls back, because the writer may have partially mutated transactional state. Registering it anyway is a correctness trap that gives false confidence. ## Idempotency is a precondition, not a feature Because rollback triggers replay, the **writer must be idempotent** (upserts, dedup keys, idempotency tokens) and the **processor must be pure**. Non-idempotent external effects belong in an **outbox** consumed by a separate idempotent step, or behind an idempotency key the downstream honors. Skip/retry without idempotency is a latent duplicate-data bug. ## Bounding blast radius - `skipLimit` / `retryLimit` cap how much the step tolerates before failing. Set them to distinguish *expected sparse bad records* from *systemic corruption* — you want the job to fail loudly if 40% of records are bad, not silently skip them into a black hole. - Emit metrics on skip counts and rescan frequency; a rising skip rate is an upstream data-quality signal. - Consider `SkipListener` to route skipped items to a dead-letter/quarantine table for later inspection instead of losing them. ## Interplay with retry Retry re-attempts an item (each attempt can roll back) up to `retryLimit`; a still-failing, skippable exception can then be skipped. Retryable transient faults (deadlocks, timeouts) generally *should* roll back (start clean), so noRollback rarely applies to retryable types. Reserve noRollback for terminal, non-transactional, skippable business rejections. ## The principal-level synthesis Treat chunk size, skip/retry limits, noRollback, and writer idempotency as one coupled design: chunk size sets throughput and failure cost, idempotency makes rescans safe, noRollback trims wasted rollbacks for expected rejections, and limits + listeners bound and observe the blast radius. Get those aligned and a batch job is both fast on the happy path and safe under partial failure.

  • Why would you keep chunks small when the input data is known to be dirty?
    Because a failure rolls back and re-scans the entire chunk item-by-item; with many bad records, large chunks mean frequent expensive N-way rescans and more wasted bulk work. Smaller chunks bound the rescan cost per failure and give more frequent commit checkpoints for restart.
  • Would you apply noRollback to a transient deadlock exception you're retrying?
    No. Retryable transient faults should roll back so the retry starts from a clean transaction. noRollback fits terminal, non-transactional, skippable business rejections in the processor phase, not retryable transient DB errors.
  • How do you avoid silently losing skipped records?
    Register a SkipListener that writes skipped items (and the cause) to a quarantine/dead-letter table, and set skipLimit so a systemic failure aborts the job rather than skipping everything into oblivion.

saying these in an interview costs you the question

  • Using noRollback as a blanket 'don't fail my step' switch across all exception types
  • Relying on noRollback for writer exceptions
  • Choosing huge chunk sizes without accounting for rescan cost or writer idempotency
  • Setting an unbounded/huge skipLimit that hides systemic data corruption

context