Compare FixedBackOff and ExponentialBackOff for retry spacing. How do you cap attempts with ExponentialBackOff, and why add jitter?
answer
- Fixed = constant interval, maxAttempts
- Exponential = ×multiplier up to maxInterval
- Plain ExponentialBackOff ~ unbounded by count
- Use ExponentialBackOffWithMaxRetries to cap
- Jitter avoids thundering-herd retry storms
basics
~20 sFixedBackOff retries at a constant interval for a set number of attempts. ExponentialBackOff grows the delay each time (×multiplier up to a max). To bound retries with exponential delays, use ExponentialBackOffWithMaxRetries. Jitter (randomness) spreads retries out to avoid all consumers hammering a recovering dependency at once.
solid answer
~40 sFixedBackOff(interval, maxAttempts) spaces every retry equally — simple and predictable, good when the failure clears on a roughly constant timescale. ExponentialBackOff starts at initialInterval and multiplies by a factor each attempt (capped at maxInterval), which backs off pressure on a struggling dependency. A gotcha: plain ExponentialBackOff defaults to effectively unlimited retries (Long.MAX_VALUE elapsed time), so to bound it in a DefaultErrorHandler you use ExponentialBackOffWithMaxRetries(maxRetries) and set initialInterval/multiplier/maxInterval. Jitter — setting a randomness/multiplier on the backoff — perturbs each delay so that many consumers that failed simultaneously (e.g. a downstream outage) don't all retry in lockstep and create a synchronized thundering-herd spike when the dependency recovers. Choose fixed for transient blips with constant recovery time; exponential+jitter for overloaded or rate-limited downstreams.
go deeper
Know fixed = same delay each time, exponential = growing delay, and that you can cap the number of retries.
Use ExponentialBackOffWithMaxRetries correctly, set initial/multiplier/max, and explain jitter's purpose.
Tie backoff choice to downstream behavior, poll-interval limits, and fleet-wide retry-storm avoidance.
Define backoff/jitter standards per dependency class and model aggregate retry load across many consumers.
## The two BackOff types (from Spring Core's `org.springframework.util.backoff`) ### `FixedBackOff(interval, maxAttempts)` Every retry waits the same `interval` milliseconds; `maxAttempts` bounds the number of retries. `new FixedBackOff(2000L, 3)` = up to 3 retries, 2s apart. It's predictable and easy to reason about, ideal when the underlying failure is expected to clear after a fairly constant time (e.g. a brief network blip). ### `ExponentialBackOff` Delays **grow geometrically**: start at `initialInterval`, multiply by `multiplier` each attempt, capped at `maxInterval`. E.g. initial=1s, multiplier=2, maxInterval=30s → 1s, 2s, 4s, 8s, 16s, 30s, 30s…. This **relieves pressure** on a struggling dependency: the harder it's failing, the less often you poke it. ## The 'unbounded' gotcha Plain `ExponentialBackOff` is bounded by **elapsed time** (`maxElapsedTime`, default `Long.MAX_VALUE`), **not by a retry count**. Used naively in a `DefaultErrorHandler`, that means it could retry essentially forever. To cap by attempts, Spring provides **`ExponentialBackOffWithMaxRetries(int maxRetries)`**: ``` var backOff = new ExponentialBackOffWithMaxRetries(5); backOff.setInitialInterval(1000L); backOff.setMultiplier(2.0); backOff.setMaxInterval(10_000L); var handler = new DefaultErrorHandler(recoverer, backOff); ``` ## Why jitter When a **shared downstream** (DB, API) goes down, *every* consumer instance and *every* in-flight record fails at nearly the same moment. With deterministic backoff they all wake to retry **simultaneously** — a **thundering herd** / synchronized retry storm that can re-knock-over the dependency the instant it recovers. **Jitter** adds randomness to each delay so retries are spread across a window, smoothing the load. In Spring you can introduce randomness via the backoff's `multiplier`/random support (and at the `@Backoff(random = true)` level for `@RetryableTopic`). ## Interaction with `max.poll.interval.ms` For **blocking** retries (DefaultErrorHandler), the cumulative backoff still counts against `max.poll.interval.ms` (default 5 min). A long exponential tail can blow that and trigger a rebalance — another reason to cap `maxInterval` and total retries, or move long backoffs to **non-blocking** retry topics where the delay lives in the topic, not the poll loop. ## Choosing - **FixedBackOff**: transient, constant-timescale failures; predictable SLAs. - **ExponentialBackOff(WithMaxRetries) + jitter**: overloaded/rate-limited downstreams, shared dependencies, large fleets — to avoid synchronized retry storms.
- What's the trap with using ExponentialBackOff directly in DefaultErrorHandler?Plain ExponentialBackOff is bounded by maxElapsedTime (default Long.MAX_VALUE), not by a retry count, so it can retry almost forever. Use ExponentialBackOffWithMaxRetries to bound by number of attempts.
- How does jitter help when a shared database outage causes all consumers to fail at once?Without jitter, all consumers back off by identical deterministic delays and retry in lockstep, creating a synchronized spike that can re-overwhelm the database the moment it recovers. Jitter randomizes each delay, spreading retries over a window and smoothing the load.
saying these in an interview costs you the question
- Saying plain ExponentialBackOff is bounded by retry count — it's bounded by elapsed time by default.
- Ignoring max.poll.interval.ms when using long blocking exponential backoffs.
- Claiming jitter slows recovery — it prevents synchronized retry storms that can delay recovery further.
- Using exponential backoff for a constant-timescale transient failure where fixed is simpler and sufficient.