skip to content

When would you supply a custom RetryPolicy instead of retryLimit(), and what is a RetryListener good for?

level: seniorimportance: should knowfreq 45%

answer

  1. retryLimit -> SimpleRetryPolicy
  2. retryPolicy() overrides, don't mix
  3. TimeoutRetryPolicy / ExceptionClassifierRetryPolicy / Composite
  4. listener: open(veto)/onError/close
  5. getRetryCount() in onError

basics

~20 s

retry()/retryLimit() only build a simple attempt-count policy. A custom RetryPolicy lets you vary the rule — different limits per exception, time-boxed retries, or always/never. A RetryListener hooks the retry lifecycle (open/onError/close) for logging, metrics, or aborting.

solid answer

~40 s

retry(type)+retryLimit(n) internally builds a SimpleRetryPolicy: a flat maxAttempts plus a retryable-exceptions map. When you need richer rules you pass retryPolicy(RetryPolicy) instead: TimeoutRetryPolicy (retry until a wall-clock budget expires), ExceptionClassifierRetryPolicy (different policy per exception type), CompositeRetryPolicy (AND/OR combine several), or AlwaysRetryPolicy/NeverRetryPolicy. You use one or the other, not both, since retryPolicy overrides the generated one. A RetryListener is a callback around each retry: open() is called once before the first attempt and can veto retrying by returning false, onError() fires after every failed attempt (ideal for logging the attempt count and warning-level telemetry), and close() runs once at the end with the final throwable. Register it with listener(new MyRetryListener()) on the fault-tolerant builder. Listeners are for observability and metrics — they do not change the retry decision except via open()'s boolean.

code

java · 21 lines
java
// Different retry aggression per exception type
Map<Class<? extends Throwable>, RetryPolicy> policies = new HashMap<>();
policies.put(LockTimeoutException.class, new SimpleRetryPolicy(5));
policies.put(HttpServerErrorException.class, new SimpleRetryPolicy(3));
ExceptionClassifierRetryPolicy classifier = new ExceptionClassifierRetryPolicy();
classifier.setPolicyMap(policies);

RetryListener logging = new RetryListener() {
    @Override public <T, E extends Throwable> void onError(
            RetryContext ctx, RetryCallback<T, E> cb, Throwable t) {
        log.warn("retry attempt #{} failed: {}", ctx.getRetryCount(), t.toString());
    }
};

return new StepBuilder("step", jobRepository)
        .<In, Out>chunk(20, tx)
        .reader(reader).processor(processor).writer(writer)
        .faultTolerant()
        .retryPolicy(classifier)     // instead of retry()/retryLimit()
        .listener(logging)
        .build();

go deeper

for a junior

Know a RetryListener exists for logging around retries.

for a middle

Name SimpleRetryPolicy vs alternatives and the open/onError/close listener methods.

for a senior

Choose the right RetryPolicy for the failure shape (count vs time vs per-type) and wire listeners for metrics.

for a principal

Standardize retry policy + listener conventions across jobs (telemetry, budgets) and avoid control logic leaking into listeners.

### Why retryLimit is not always enough `retry(...).retryLimit(n)` produces a `SimpleRetryPolicy` — one flat attempt count applied to every registered exception. Real systems often need more nuance, which is where **`retryPolicy(RetryPolicy)`** comes in. The `RetryPolicy` interface (Spring Retry) decides, given a `RetryContext`, whether another attempt is allowed. Built-in implementations: - **SimpleRetryPolicy** — fixed maxAttempts + retryable exception map (what retryLimit builds). - **TimeoutRetryPolicy** — keep retrying until a wall-clock timeout elapses, regardless of count. - **AlwaysRetryPolicy / NeverRetryPolicy** — unconditional yes/no. - **ExceptionClassifierRetryPolicy** — route each exception type to its own delegate policy (e.g. 5 attempts for a lock timeout, 1 for everything else). - **CompositeRetryPolicy** — combine policies with optimistic (any allows) or pessimistic (all must allow) semantics. You provide **either** retry()/retryLimit() **or** retryPolicy(...) — mixing them is a configuration error because the explicit policy replaces the generated one. ### RetryListener lifecycle `RetryListener` (Spring Retry) is the observability hook. Its methods: - `open(RetryContext, RetryCallback)` — called **once** before the first attempt. Returning **false aborts** the whole retry (rarely used, but it is the one method that affects control flow). - `onError(RetryContext, RetryCallback, Throwable)` — called **after each failed attempt**. `context.getRetryCount()` tells you how many attempts have failed. This is the natural place to log a WARN, bump a Micrometer counter, or emit a trace event. - `close(RetryContext, RetryCallback, Throwable)` — called **once** when retrying ends, with the last throwable (null on success). Good for final success/failure metrics. - (Spring Retry 2 also adds a default `onSuccess` callback.) Register via `.listener(retryListener)` on `FaultTolerantStepBuilder` — the overloaded `listener(...)` accepts `RetryListener` (as well as skip/chunk/step listeners; Batch dispatches by type). ### Practical guidance - Use **ExceptionClassifierRetryPolicy** when different transient failures deserve different aggression. - Use **TimeoutRetryPolicy** when what matters is a latency budget, not a count (e.g. retry a remote call for up to 2s). - Keep listeners **side-effect-light**; they run on the hot path of every failure. - Do not put business logic in `open()`'s veto — prefer clear exception classification for the retry decision.

  • Can a RetryListener stop retries from happening?
    Only via open() returning false, which aborts retry before the first attempt. onError() and close() are observational and cannot change the decision; the RetryPolicy owns the keep-retrying? decision.
  • What happens if you set both retryLimit() and retryPolicy()?
    They conflict — retryPolicy replaces the SimpleRetryPolicy that retry()/retryLimit() would build, so the limit is effectively ignored or throws a config error. Pick one approach.

saying these in an interview costs you the question

  • Thinking onError() can cancel the remaining retries
  • Combining retryLimit() and a custom retryPolicy() expecting both to apply
  • Believing RetryListener changes the retry count
  • Not knowing TimeoutRetryPolicy or ExceptionClassifierRetryPolicy exist

context