skip to content

What is skip handling in Spring Batch, and how do you enable it on a chunk-oriented step?

level: juniorimportance: must knowfreq 62%

answer

  1. faultTolerant() unlocks skip/retry
  2. .skip(Ex.class) + .skipLimit(n)
  3. over limit -> SkipLimitExceededException
  4. read skip = drop; process/write skip = rollback + replay single items
  5. counts stored on StepExecution

basics

~10 s

Skip handling lets a step ignore individual bad items instead of failing the whole job. You enable it with faultTolerant() on the step builder, then declare which exceptions to skip and a skipLimit.

solid answer

~40 s

By default a chunk-oriented step fails the moment any read/process/write throws. Skip handling makes the step fault-tolerant: a failing item is bypassed and the step keeps processing the rest. You turn it on with faultTolerant() on the StepBuilder, then chain .skip(SomeException.class) to whitelist an exception type and .skipLimit(n) to cap how many skips are allowed. Exceed the limit and the step fails with SkipLimitExceededException. Use it for dirty input where a few malformed records shouldn't abort a large batch — e.g. a FlatFileParseException on one CSV line. It only applies to the item level within read/process/write; it does not catch framework/config errors.

code

java · 16 lines
java
@Bean
public Step importStep(JobRepository jobRepository,
                       PlatformTransactionManager tx,
                       ItemReader<Person> reader,
                       ItemProcessor<Person, Person> processor,
                       ItemWriter<Person> writer) {
    return new StepBuilder("importStep", jobRepository)
            .<Person, Person>chunk(100, tx)
            .reader(reader)
            .processor(processor)
            .writer(writer)
            .faultTolerant()
            .skip(FlatFileParseException.class)  // whitelist skippable exception
            .skipLimit(10)                        // fail the step after 10 skips
            .build();
}

go deeper

for a junior

Know faultTolerant() + skip(Exception) + skipLimit(n) and that it bypasses bad items instead of failing the step.

for a middle

Explain the SkipLimitExceededException failure, and that read vs process/write skips behave differently.

for a senior

Discuss the rollback-and-replay cost of write skips, idempotency needs, and skip vs retry composition.

for a principal

Frame skip policy as a data-quality/observability decision — reject reporting, limits tuned to acceptable defect rates, and when a hard stop is safer.

**What it is.** In Spring Batch a *chunk-oriented step* reads items one by one (ItemReader), optionally transforms them (ItemProcessor), and writes them in batches (ItemWriter). By default the step is *not* fault tolerant: the first exception thrown by any of these phases fails the chunk, rolls back the transaction, and fails the whole step (and usually the job). **Skip handling** changes that policy so that an individual *item* that causes a recoverable exception is discarded ("skipped") and the step continues with the next item, rather than aborting. **How to enable it.** On the `StepBuilder` you call `.faultTolerant()`, which returns a `FaultTolerantStepBuilder`. On that builder you declare the skip policy fluently: - `.skip(FlatFileParseException.class)` — whitelist an exception type as skippable. You can chain several `.skip(...)` calls. - `.noSkip(SomeOtherException.class)` — explicitly exclude a subtype from being skipped even if a broader parent type is skippable. - `.skipLimit(10)` — the maximum number of items that may be skipped across the whole step. Once the (n+1)th skip would occur, the step throws `SkipLimitExceededException` and fails. Choose this deliberately: a high or unlimited limit can silently swallow massive data-quality problems. **Key semantics.** A skip is *per item*, not per exception occurrence in general — only exceptions thrown while handling a specific item are eligible. Skips are counted and stored in the `StepExecution` (`getReadSkipCount`, `getProcessSkipCount`, `getWriteSkipCount`, and the total `getSkipCount`), so they survive restarts and appear in the job repository. **Read vs process vs write differ importantly.** A skip on *read* simply drops that one item and reads the next — cheap, no rollback of the chunk. A skip on *process* or *write* is more expensive: because the framework can't know which item in the chunk failed, it rolls the chunk back and *re-processes the items one at a time* (single-item chunks) to isolate the offending item, skip just that one, and commit the rest. This is why write skips hurt throughput and why writers/processors should ideally be idempotent. **When to use.** Dirty or semi-structured input (CSV/JSON feeds), integration with flaky downstream systems where a handful of bad records is acceptable, ETL loads where you want a report of rejects rather than a hard stop. **When not to.** Financial correctness where every record matters, or where a skip masks a systemic bug — there a hard failure plus a retry/restart is safer. **Related but distinct.** `.retry(...)`/`.retryLimit(...)` re-attempts a transient failure (e.g. deadlock) before giving up; skip *discards*. They compose: retry first, and if retries are exhausted the item can then be skipped. Don't confuse skip (drop the item) with retry (try again).

  • What happens when the number of skipped items exceeds skipLimit?
    The step fails immediately with SkipLimitExceededException; the item that would have been the (limit+1)th skip is not skipped.
  • Does a skip on read behave the same as a skip on write?
    No. A read skip just drops the item and reads the next with no rollback. A write (or process) skip rolls the chunk back and replays items one-per-chunk to isolate and skip only the bad one, so it's far more expensive.

saying these in an interview costs you the question

  • Thinking faultTolerant() alone skips everything without declaring .skip(...) exceptions
  • Believing skipLimit defaults to unlimited (you must set it; without a policy nothing is skipped)
  • Assuming a write skip is as cheap as a read skip
  • Confusing skip (discard the item) with retry (re-attempt the item)

context