skip to content

Fault Tolerance

Surviving bad data and flaky dependencies: skipping items, retrying transient failures, listeners for visibility, and what actually rolls back. Interviewers ask because a batch that dies on row 900,000 is a real operational cost.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

explore

questions

20

What is a JobExecutionListener in Spring Batch, and when do its callbacks run?

level: juniorimportance: must knowfreq 55%

answer

  1. beforeJob / afterJob
  2. afterJob runs even on FAILED
  3. JobExecution -> BatchStatus/ExitStatus
  4. JobBuilder.listener(...)
  5. @BeforeJob / @AfterJob annotation form

basics

~10 s

A JobExecutionListener lets you run code around a job. beforeJob runs before the job starts; afterJob runs once the job finishes, whether it succeeded or failed. You register it on the job builder.

solid answer

~30 s

JobExecutionListener is a lifecycle hook with two callbacks: beforeJob(JobExecution) runs once before any step starts, and afterJob(JobExecution) runs once after the job completes. Crucially afterJob runs regardless of outcome — COMPLETED, FAILED, or STOPPED — so it's the right place for cleanup, notifications, or inspecting the final BatchStatus/ExitStatus. You register it with JobBuilder.listener(...). You can either implement the interface (its methods are default methods in Spring Batch 5, so override only what you need) or annotate plain methods with @BeforeJob/@AfterJob. Common uses: sending a completion email, logging start/end timestamps, initializing shared resources, or writing summary metrics from the JobExecution's ExecutionContext.

code

java · 21 lines
java
public class NotifyingJobListener implements JobExecutionListener {

    @Override
    public void beforeJob(JobExecution jobExecution) {
        jobExecution.getExecutionContext().put("startedAt", System.currentTimeMillis());
    }

    @Override
    public void afterJob(JobExecution jobExecution) {
        // Runs on success AND failure
        if (jobExecution.getStatus() == BatchStatus.FAILED) {
            // send alert
        }
    }
}

// Registration
Job job = new JobBuilder("reportJob", jobRepository)
        .listener(new NotifyingJobListener())
        .start(step1)
        .build();

go deeper

for a junior

Know the two callbacks and that afterJob runs on success or failure.

for a middle

Know both registration styles and that JobExecution carries BatchStatus/ExitStatus and an ExecutionContext.

for a senior

Discuss idempotent cleanup, seeding vs reading the job ExecutionContext, and behavior when beforeJob throws.

for a principal

Frame it as the outermost of Batch's nested listener scopes and where job-level observability/metrics and resource lifecycles belong.

## What it is Spring Batch models a **Job** as a container of one or more **Steps**. A `JobExecutionListener` is a lifecycle hook that lets your code observe the job as a whole, outside the steps. It has two callbacks: - `void beforeJob(JobExecution jobExecution)` — invoked once, before the first step runs. - `void afterJob(JobExecution jobExecution)` — invoked once, after the last step finishes. Both receive the `JobExecution`, the runtime object describing this run: its `BatchStatus` (COMPLETED / FAILED / STOPPED / etc.), its `ExitStatus`, start/end times, the list of `StepExecution`s, and a job-level `ExecutionContext` (a key/value bag persisted in the job repository). ## Key semantic: afterJob always runs `afterJob` runs **whether the job succeeded or failed** (as long as the job actually started). That makes it the canonical place for cleanup and end-of-job reporting. You typically branch on status: ```java if (jobExecution.getStatus() == BatchStatus.FAILED) { /* alert */ } ``` ## Two ways to write one 1. **Implement the interface.** In Spring Batch 5 the methods are `default`, so you override only the one you care about. 2. **Annotation form.** Put `@BeforeJob` / `@AfterJob` on plain methods of any bean and register that bean; Spring wraps it via `JobListenerFactoryBean`. The annotated method may take a `JobExecution` parameter or no parameter. ## Registration ```java new JobBuilder("myJob", jobRepository) .listener(myJobListener) .start(step1) .build(); ``` ## When to use - Send a start/finish notification or email. - Log or export job-level metrics (durations, counts). - Seed the job `ExecutionContext` before steps, or read aggregated results after. - Acquire/release a resource that spans the whole job. ## Gotchas - Don't put per-record logic here — it fires once per job, not per item; item/step listeners exist for finer scopes. - `beforeJob` is not called if the job fails to launch (e.g., validation of `JobParameters` fails before execution). Once a `JobExecution` exists and starts, both callbacks fire. - Exceptions thrown from `beforeJob` fail the job before any step runs.

  • Does afterJob run if the job fails halfway through a step?
    Yes. As long as the job started, afterJob is always invoked with the final JobExecution — you inspect getStatus()/getExitStatus() to see it FAILED. It only wouldn't run if the job never launched (e.g., pre-launch parameter validation failure).
  • How would you use the annotation form instead of implementing the interface?
    Put @BeforeJob and/or @AfterJob on plain methods of a bean (optionally taking a JobExecution parameter) and pass that bean to JobBuilder.listener(...). Spring wraps it via JobListenerFactoryBean so no interface is needed.

saying these in an interview costs you the question

  • Claiming afterJob only runs on success
  • Thinking beforeJob/afterJob fire per step or per item
  • Confusing JobExecutionListener with StepExecutionListener scope

context

open as a page

In a Spring Batch chunk-oriented step, what does retry(Exception.class).retryLimit(n) do, and when would you use it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

On a faultTolerant() step, retry(SomeException.class) marks that exception as retryable and retryLimit(n) caps the number of attempts. Spring Batch re-runs the failing work instead of failing the job. Use it for transient errors like a brief network or DB blip.

open as a page

In Spring Batch chunk-oriented processing, what happens to the current chunk's transaction when an exception is thrown while processing or writing an item?

level: juniorimportance: must knowfreq 55%

basics

~20 s

The whole chunk's transaction rolls back, so none of the items in that chunk get written. Spring Batch wraps each chunk (read-process-write of N items) in a single transaction, so one bad item undoes the entire chunk.

open as a page

What is skip handling in Spring Batch, and how do you enable it on a chunk-oriented step?

level: juniorimportance: must knowfreq 62%

basics

~10 s

Skip handling lets a step ignore individual bad items instead of failing the whole job. You enable it with faultTolerant() on the step builder, then declare which exceptions to skip and a skipLimit.

open as a page

How does StepExecutionListener work, and what is the @BeforeStep annotation commonly used for?

level: middleimportance: must knowfreq 50%

basics

~20 s

StepExecutionListener has beforeStep and afterStep callbacks that run around a single step. @BeforeStep is often put on a reader or writer method to grab the StepExecution — and through it the JobParameters — so the component can configure itself.

open as a page

When a fault-tolerant Spring Batch step rolls back a chunk, why does it re-read the chunk item-by-item, and how does that isolate the bad item?

level: middleimportance: must knowfreq 50%

basics

~20 s

After a rollback Batch can't tell which item failed, since all N were in one transaction. So it re-processes the chunk one item at a time (chunk size effectively 1). The item that throws again is identified as the bad one and gets skipped or retried; the others commit.

open as a page

How exactly does Spring Batch re-run work when a retryable exception is thrown mid-chunk, and does retryLimit(3) mean three retries or three total attempts?

level: middleimportance: should knowfreq 55%

basics

~20 s

retryLimit(3) means three total attempts (the first try plus two retries), because it maps to SimpleRetryPolicy.maxAttempts, which counts the initial call. On a retryable failure Batch rolls back the chunk transaction and reprocesses the items again, up to that limit.

open as a page

How does Spring Batch handle a skip differently for the read, process, and write phases of a chunk?

level: middleimportance: should knowfreq 45%

basics

~20 s

A read skip just drops the item and reads the next, with no rollback. A process or write skip forces the chunk transaction to roll back and re-run items one at a time to find and skip the single bad item, then commit the rest.

open as a page

What are the ChunkListener callbacks, and how do they relate to the chunk's transaction boundaries?

level: seniorimportance: should knowfreq 40%

basics

~10 s

ChunkListener has beforeChunk, afterChunk, and afterChunkError. beforeChunk runs at the start of the chunk inside its transaction; afterChunk runs after that transaction commits successfully; afterChunkError runs if the chunk failed and rolled back.

open as a page

Explain the item-level listeners (ItemReadListener, ItemProcessListener, ItemWriteListener) and their error callbacks.

level: seniorimportance: should knowfreq 42%

basics

~20 s

These three listeners hook the read, process, and write phases of each chunk. Each has a before, an after, and an onError callback — for example ItemReadListener has beforeRead, afterRead(item), and onReadError(exception) — so you can log, count, or inspect failures at each stage.

open as a page

How do you add exponential backoff between retries in Spring Batch, and what should you know about ExponentialBackOffPolicy?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Spring Batch's chunk retry (retryLimit) has no delay between retries and the FaultTolerantStepBuilder does not expose a BackOffPolicy. To get exponential backoff you use Spring Retry directly — a RetryTemplate with ExponentialBackOffPolicy, or @Retryable(backoff=@Backoff) on the processor/writer method.

open as a page

When would you supply a custom RetryPolicy instead of retryLimit(), and what is a RetryListener good for?

level: seniorimportance: should knowfreq 45%

basics

~20 s

retry()/retryLimit() only build a simple attempt-count policy. A custom RetryPolicy lets you vary the rule — different limits per exception, time-boxed retries, or always/never. A RetryListener hooks the retry lifecycle (open/onError/close) for logging, metrics, or aborting.

open as a page

What does .noRollback(Exception.class) do on a fault-tolerant Spring Batch step, and when would you use it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

noRollback tells the step that a given exception type should NOT trigger a transaction rollback. Use it when a processor throws a business/validation exception you want to handle (e.g. skip that item) without discarding the whole chunk's transaction. It applies to processor exceptions, not writer exceptions.

open as a page

Given Spring Batch re-processes a rolled-back chunk item-by-item, what does that imply for an ItemProcessor that has side effects, and how do you handle it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Because a rolled-back chunk is replayed one item at a time, the same item can pass through the ItemProcessor more than once. Any side effect there (charging a card, sending email, calling an external API) would happen repeatedly. Keep the processor idempotent or pure and push non-idempotent effects to an idempotent writer.

open as a page

When would you implement a custom SkipPolicy instead of using skip()/skipLimit(), and how does it work?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Use a custom SkipPolicy when the skip decision needs logic beyond "is this exception type in the whitelist and under the count limit" — e.g. different limits per exception, inspecting the exception message, or unlimited skips for one type. You implement shouldSkip(Throwable, long) and register it with .skipPolicy(...).

open as a page

How do you record or route items that get skipped, and what are the gotchas with SkipListener?

level: seniorimportance: should knowfreq 30%

basics

~10 s

Register a SkipListener with .listener(...) on the fault-tolerant step. It has onSkipInRead(Throwable), onSkipInProcess(item, Throwable), and onSkipInWrite(item, Throwable). Use them to log or write rejected items to a dead-letter table so nothing is lost silently.

open as a page

When an exception is registered for both retry and skip, how do retry and skip interact, and how do their counters and the writer scan interplay?

level: principalimportance: should knowfreq 35%

basics

~20 s

Retry runs first: Batch retries the item up to retryLimit. Only when retries are exhausted and the same exception is skippable does it skip the item, counting it toward skipLimit. If skipLimit is exceeded the step fails. So the order is retry, then skip, then fail.

open as a page

Design-wise, how do skip, retry, and restart interact, and how do you choose skip limits and exception scope safely in production?

level: principalimportance: should knowfreq 22%

basics

~20 s

Retry re-attempts transient failures; if retries are exhausted the item can then be skipped (permanent-bad data). Restart resumes a failed job from the last committed chunk. Choose narrow skippable exception types and a limit tuned to an acceptable defect rate, never 'skip Exception.class' with a huge limit.

open as a page

How are Batch listeners registered, and what are the design trade-offs between the interface form, the annotation form, and listener ordering/nesting?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

You register listeners on the builders — JobBuilder.listener for job listeners, StepBuilder.listener for step/chunk/item listeners. You can implement the interface or annotate methods (@BeforeStep, @AfterChunk, etc.); Spring's factory beans detect the annotations. Listeners nest job -> step -> chunk -> item and fire in registration order.

open as a page

As a principal engineer, how would you decide when to use noRollback and how to size chunks, given the rollback-and-rescan cost model of fault-tolerant Spring Batch steps?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Use noRollback only for processor-phase exceptions where nothing transactional was written and re-scanning is wasted work — typically validation skips. Size chunks by balancing happy-path throughput against the cost of rolling back and re-scanning a whole chunk on failure, given your data's error rate and writer idempotency.

open as a page