skip to content

Design-wise, how do skip, retry, and restart interact, and how do you choose skip limits and exception scope safely in production?

level: principalimportance: should knowfreq 22%

answer

  1. retry=transient, skip=bad-data, restart=resume
  2. retry-exhausted -> becomes skip candidate
  3. restart = same JobParameters resumes after last committed chunk
  4. never skip(Exception.class); whitelist narrow types + noSkip
  5. limit = acceptable defect rate; pair with SkipListener + metrics

basics

~20 s

Retry re-attempts transient failures; if retries are exhausted the item can then be skipped (permanent-bad data). Restart resumes a failed job from the last committed chunk. Choose narrow skippable exception types and a limit tuned to an acceptable defect rate, never 'skip Exception.class' with a huge limit.

solid answer

~50 s

These three are layered defenses. Retry (.retry(TransientEx).retryLimit(n)) handles flaky/transient failures — deadlocks, timeouts — by re-attempting the same item. Skip (.skip(BadDataEx).skipLimit(m)) handles permanently-bad items by discarding them. They compose: an exception can be retried first and, once retries are exhausted, become a skip candidate. Restart is orthogonal: because Spring Batch commits per chunk and stores state in the job repository, a failed job restarted with the same JobParameters resumes after the last committed chunk. Design rules: whitelist narrow, specific exceptions (FlatFileParseException, ValidationException) — never blanket Exception.class; set skipLimit from an acceptable defect ratio, not an arbitrary large number, so systemic corruption still fails loudly; always pair skips with a SkipListener dead-letter store; and keep processors/writers idempotent because both retry and write-skip replay items. Treat skip as 'reject and report', retry as 'try again', restart as 'resume'.

code

java · 21 lines
java
@Bean
public Step resilientStep(JobRepository repo, PlatformTransactionManager tx,
                          ItemReader<Rec> r, ItemProcessor<Rec, Rec> p,
                          ItemWriter<Rec> w, SkipListener<Rec, Rec> deadLetter) {
    return new StepBuilder("resilientStep", repo)
            .<Rec, Rec>chunk(200, tx)
            .reader(r).processor(p).writer(w)
            .faultTolerant()
            // transient failures: re-attempt the same item
            .retry(DeadlockLoserDataAccessException.class)
            .retry(TransientDataAccessException.class)
            .retryLimit(3)
            // permanently-bad data: discard, narrowly scoped
            .skip(FlatFileParseException.class)
            .skip(RecordValidationException.class)
            .noSkip(RecordValidationException.class.getSuperclass() == null
                    ? RecordValidationException.class : RecordValidationException.class) // illustrative
            .skipLimit(50)          // ~ acceptable defect budget, not 'unlimited'
            .listener(deadLetter)   // capture rejects for reconciliation
            .build();
}

go deeper

for a junior

Know skip drops bad items and retry re-attempts; they're different.

for a middle

Explain retry-then-skip composition and that restart resumes from the last committed chunk with the same JobParameters.

for a senior

Justify narrow exception whitelists, idempotency under replay, and pairing skips with a dead-letter listener.

for a principal

Set limits from defect-rate SLAs, decide skip-vs-hard-fail per domain criticality, and standardize retry/skip/restart + observability across the batch platform.

**Three mechanisms, three jobs.** - **Retry** (`.faultTolerant().retry(SomeTransientException.class).retryLimit(3)`) re-executes the *same* item when the failure is likely transient — optimistic-lock/`DeadlockLoserDataAccessException`, network `TransientDataAccessException`, momentary downstream 503s. Retry keeps the item; it just tries again (optionally with a `BackOffPolicy`). - **Skip** discards a permanently-bad item so one poison record doesn't sink a million-row job. - **Restart** is not a per-item mechanism at all: Spring Batch persists `StepExecution`/`JobExecution` state (including the reader's `ExecutionContext` position for restartable readers) in the *job repository* and commits per chunk. Re-launching a **FAILED** job with the **same `JobParameters`** creates a new `JobExecution` that resumes from the last committed chunk rather than from the top. (A COMPLETED job with identical parameters won't rerun — identity is the parameters.) **How they compose.** On a fault-tolerant step you can declare both retry and skip. The order of defense for a given exception is: attempt → if it throws a retryable exception, retry up to the limit → if still failing and the exception is also skippable, skip it (counting against skipLimit) → else fail the step. So an exception can be *both* retryable and skippable: transient blips get retried, genuinely stuck items eventually get skipped. `RetryListener` and `SkipListener` observe the two phases respectively. **Idempotency is the cross-cutting constraint.** Both retry and process/write skip cause items to be re-executed (retry re-runs the item; write-skip rolls back and replays the chunk one item at a time). Any side effect in the processor/writer — external calls, emails, non-transactional counters — can therefore happen more than once. Principal-level design pushes side effects to be idempotent, transactional, or deferred to a final idempotent write. **Choosing the skippable exception scope.** The dangerous anti-pattern is `.skip(Exception.class)` (or `Throwable`) with a large limit — that swallows NPEs, config errors, and OOM-ish issues that indicate *bugs*, not bad data, masking real failures and silently dropping records. Best practice: whitelist the *narrowest* exceptions that genuinely mean "this one record is bad" (`FlatFileParseException`, bean-`ValidationException`, a domain `InvalidRecordException`), and use `.noSkip(...)` to carve out subtypes you must never skip. **Choosing the limit.** Set `skipLimit` from a business-defined *acceptable defect rate*, not a round number. If >0.5% of a feed is malformed, that usually signals an upstream/schema problem the batch should surface by failing — so a limit proportional to expected volume (or a percentage-based custom `SkipPolicy`) fails loudly on systemic corruption while tolerating the odd bad row. An effectively-unlimited skip limit is almost always a mistake. **Observability.** Every fault-tolerant step should: (1) pair skips with a `SkipListener` dead-letter store for reconciliation, (2) surface `getSkipCount()`/read/process/write counts as metrics/alerts, and (3) log retry exhaustion. Silent skips are a data-integrity liability. **When a hard failure beats a skip.** Correctness-critical domains (ledger postings, regulatory filings) should generally *not* skip — a single bad record failing the step, plus retry for transient issues and restart to resume, is safer than dropping records. Skip is for feeds where partial success with a reject report is the accepted contract. **Summary mental model:** retry = "try again (transient)", skip = "reject and report (bad data)", restart = "resume where we left off (after a failure)". They're complementary, not alternatives.

  • Can a single exception type be both retryable and skippable, and what happens?
    Yes. It's retried up to the retryLimit first; if it still fails after retries are exhausted, it becomes a skip candidate and is skipped (counting against skipLimit) rather than failing the step immediately.
  • How does restart know where to resume after a failed fault-tolerant step?
    Spring Batch commits per chunk and persists StepExecution state (including restartable readers' position in the ExecutionContext) to the job repository. Re-launching with the same JobParameters resumes after the last committed chunk.
  • Why is .skip(Exception.class) with a large limit dangerous?
    It skips programming/config errors (NPE, misconfiguration) as if they were bad data, silently dropping records and masking systemic bugs that should fail the job loudly.

saying these in an interview costs you the question

  • Treating skip and retry as interchangeable rather than complementary
  • Setting skipLimit to Integer.MAX_VALUE / effectively unlimited
  • Whitelisting Exception or Throwable as skippable
  • Ignoring idempotency even though retry and write-skip replay items
  • Assuming restart re-runs the whole job from the beginning

context