skip to content

How do you add exponential backoff between retries in Spring Batch, and what should you know about ExponentialBackOffPolicy?

level: seniorimportance: should knowfreq 40%

answer

  1. Chunk retry = immediate, NO backoff on the builder
  2. Backoff via Spring Retry: @Backoff or RetryTemplate
  3. Exponential: initialInterval x multiplier^n capped at maxInterval
  4. defaults 100ms / x2 / 30s
  5. add jitter (ExponentialRandomBackOffPolicy) vs thundering herd

basics

~20 s

Spring Batch's chunk retry (retryLimit) has no delay between retries and the FaultTolerantStepBuilder does not expose a BackOffPolicy. To get exponential backoff you use Spring Retry directly — a RetryTemplate with ExponentialBackOffPolicy, or @Retryable(backoff=@Backoff) on the processor/writer method.

solid answer

~40 s

This is a common trap: the fault-tolerant chunk step retries immediately — there is no built-in inter-retry delay and FaultTolerantStepBuilder has no backOffPolicy() method. If you need spacing (to let a downstream recover or avoid hammering it), you drop down to Spring Retry inside your processor/writer: annotate the method with @Retryable(retryFor=..., backoff=@Backoff(delay=500, multiplier=2, maxDelay=10000)) and enable @EnableRetry, or build a RetryTemplate with an ExponentialBackOffPolicy (initialInterval, multiplier, maxInterval) and wrap the risky call in template.execute(...). ExponentialBackOffPolicy grows the wait geometrically: delay = initialInterval * multiplier^attempt, capped at maxInterval; defaults are 100ms initial, x2, 30s max. Because plain exponential backoff can synchronize many workers (thundering herd), prefer ExponentialRandomBackOffPolicy to add jitter. Note backoff blocks the processing thread, so it competes with chunk throughput.

code

java · 22 lines
java
@Configuration
@EnableRetry
class RetryConfig {}

@Component
class RemoteEnrichmentProcessor implements ItemProcessor<Order, Invoice> {

    // Backoff lives here, because the chunk step's retryLimit cannot delay retries.
    @Retryable(retryFor = RemoteAccessException.class,
               maxAttempts = 4,
               backoff = @Backoff(delay = 500, multiplier = 2.0, maxDelay = 10_000))
    @Override
    public Invoice process(Order order) {
        return remoteClient.enrich(order); // 503/timeout -> retried with growing delay
    }

    @Recover
    Invoice recover(RemoteAccessException e, Order order) {
        // final fallback after backoff attempts are exhausted
        return Invoice.deferred(order);
    }
}

go deeper

for a junior

Know that plain chunk retries happen immediately with no delay.

for a middle

Know backoff comes from Spring Retry (@Backoff or RetryTemplate), not the Batch builder.

for a senior

Configure ExponentialBackOffPolicy correctly and place the retry/backoff at the risky call boundary.

for a principal

Design distributed backoff with jitter and budgets, balancing downstream protection against thread-blocking throughput loss.

### The gotcha: chunk retry has no backoff Spring Batch's `faultTolerant().retry(...).retryLimit(n)` retries **immediately** — there is **no delay between attempts**, and `FaultTolerantStepBuilder` does **not** expose a `backOffPolicy(...)` setter. Candidates often assume you can chain a backoff onto the step; you cannot. For transient failures caused by a struggling downstream, immediate retries can make things worse. ### Getting backoff via Spring Retry Spring Batch is built on **Spring Retry**, which owns the `BackOffPolicy` abstraction. You add backoff at the **operation level**, inside your processor or writer, one of two ways: **1. Declarative — @Retryable / @Backoff** ```java @Retryable(retryFor = RemoteAccessException.class, maxAttempts = 4, backoff = @Backoff(delay = 500, multiplier = 2.0, maxDelay = 10_000)) public Invoice call(Order o) { ... } ``` with `@EnableRetry` on a config class. The proxy retries the method with the given backoff. **2. Programmatic — RetryTemplate + ExponentialBackOffPolicy** ```java ExponentialBackOffPolicy backOff = new ExponentialBackOffPolicy(); backOff.setInitialInterval(500); // first wait, ms backOff.setMultiplier(2.0); // each wait x2 backOff.setMaxInterval(10_000); // cap RetryTemplate t = new RetryTemplate(); t.setBackOffPolicy(backOff); t.setRetryPolicy(new SimpleRetryPolicy(4)); ``` Then call `t.execute(ctx -> remote.call(order))` inside the processor. ### ExponentialBackOffPolicy mechanics - Wait before attempt k ≈ `initialInterval * multiplier^(k-1)`, capped at `maxInterval`. - Defaults: `initialInterval=100ms`, `multiplier=2.0`, `maxInterval=30000ms`. - **ExponentialRandomBackOffPolicy** multiplies by a random factor to add **jitter**, avoiding the *thundering herd* where many partitions/instances retry in lockstep. - **FixedBackOffPolicy** is a constant delay; **NoBackOffPolicy** (default in a bare RetryTemplate) is immediate. ### Trade-offs / when to use - Backoff blocks the **current thread**, so long waits reduce chunk throughput — keep maxInterval sane and consider multi-threaded/partitioned steps. - Backoff belongs with **transient, downstream-pressure** failures (429/503, timeouts), not deterministic errors. - If you must retry at the chunk level *and* want backoff, the pragmatic answer is: move the flaky call into a Spring-Retry-wrapped component so backoff is applied there, and let the chunk-level retry/skip handle the residual failure classification. - Always add **jitter** in distributed batch (partitioned/remote-chunking) to prevent synchronized retries.

  • Why add jitter to exponential backoff in a partitioned/remote-chunking job?
    Without jitter, many workers that failed at the same instant retry at the same computed delays, re-synchronizing load onto the recovering downstream (thundering herd). ExponentialRandomBackOffPolicy randomizes each interval so retries spread out.
  • What is the downside of large backoff intervals in a chunk step?
    Backoff sleeps the processing thread, so the step stalls during the wait and overall throughput drops. In single-threaded steps a long maxInterval can dominate wall-clock time; mitigate with bounded maxInterval and multi-threaded or partitioned steps.

saying these in an interview costs you the question

  • Claiming FaultTolerantStepBuilder has a backOffPolicy() method
  • Assuming chunk retries already wait between attempts
  • Forgetting @EnableRetry for @Retryable/@Backoff
  • Using plain exponential backoff across many workers without jitter

context