Implement exponential backoff with jitter for a flaky network Flow using retryWhen, retrying only transient errors. Walk through the design.
answer
- base * 2^attempt, then coerceAtMost(maxDelay)
- Add jitter to avoid thundering herd
- Only retry transient errors (IO, 5xx, 429)
- Fail fast on 4xx / non-retryable
- Upstream must be idempotent; test with runTest virtual clock
basics
~10 sUse retryWhen: check the error is retryable and you haven't hit the cap, wait a growing amount of time (doubling each try) plus a small random delay, then retry; otherwise give up.
solid answer
~40 sretryWhen gives you cause and a zero-based attempt counter and is a suspend lambda, so you compute the delay and call delay() before returning true. Exponential backoff = base * 2^attempt, capped at a max to avoid runaway waits; add random jitter (e.g. +/- a fraction) so many clients don't retry in lockstep (thundering herd). Gate retries on retryability: only retry transient causes (IOException, HTTP 5xx/429) and a maximum attempt count; return false for non-retryable causes (4xx, serialization errors) so they fail fast. Never retry CancellationException — retryWhen already excludes it, but don't write predicates that catch it. Keep the upstream idempotent since it re-runs. Use coerceAtMost for the cap and Random for jitter.
code
kotlin · 9 linesfun <T> Flow<T>.retryTransient(maxRetries: Long = 5): Flow<T> =
retryWhen { cause, attempt ->
if (attempt >= maxRetries || !isTransient(cause)) false
else {
val backoff = (100L * (1L shl attempt.toInt())).coerceAtMost(5_000)
delay(backoff + Random.nextLong(0, backoff / 2 + 1))
true
}
}go deeper
Can call retryWhen with a delay but may retry all errors and forget caps.
Implements capped exponential backoff and a retryable-error gate correctly.
Adds jitter, fail-fast classification, idempotency reasoning, and virtual-clock testing.
Weighs system-wide effects: thundering herd, circuit breaking, budget/SLA on retries, and where retry belongs in the resilience stack.
## The shape of the solution ```kotlin private val RETRYABLE = setOf(500, 502, 503, 504, 429) fun isTransient(cause: Throwable): Boolean = when (cause) { is IOException -> true is HttpException -> cause.code in RETRYABLE else -> false } fun <T> Flow<T>.retryTransient( maxRetries: Long = 5, base: Long = 100, // ms maxDelay: Long = 5_000 // ms cap ): Flow<T> = retryWhen { cause, attempt -> if (attempt >= maxRetries || !isTransient(cause)) { false // give up: too many tries OR non-retryable } else { val exp = base * (1L shl attempt.toInt()) // base * 2^attempt val capped = exp.coerceAtMost(maxDelay) val jitter = Random.nextLong(0, capped / 2 + 1) // full-ish jitter delay(capped - capped / 2 + jitter) true } } ``` ## Why each piece exists - **`retryWhen`, not `retry`** — we need conditional logic (retryable check), a growing `delay`, and access to `attempt`. `retry(n)` cannot delay. - **`attempt` is zero-based** — `1L shl attempt` doubles the base each retry: attempt 0 -> base, 1 -> 2x, 2 -> 4x. This is **exponential backoff**: waits grow geometrically so a struggling server gets breathing room. - **`coerceAtMost(maxDelay)`** — caps the wait so attempt 10 doesn't sleep for minutes. - **Jitter** — adding randomness prevents the **thundering herd**: if 1000 clients fail at the same instant and all back off by exactly the same schedule, they retry in sync and re-overload the server. Randomizing spreads them out. - **Retryability gate** — retrying a 400 Bad Request or a deserialization error is pointless; those will fail identically forever. Return `false` to **fail fast** on permanent errors and reserve retries for **transient** ones (network blips, 503, rate-limit 429). - **CancellationException** — `retryWhen` never retries it; don't let `isTransient` accidentally classify it as retryable (it won't here, but a naive `else -> true` would be wrong — and even then the operator guards it). ## Idempotency Because the upstream `flow { }` re-runs on every retry, the producer must be **idempotent** (safe to repeat). Retrying a non-idempotent POST that already partially succeeded can double-charge a customer. For unsafe operations, add an idempotency key or retry at a layer where repetition is safe. ## Placement Put `retryTransient()` **directly above** the operator that can fail and **below** mapping/UI logic, so only the network step is retried and downstream transforms run once per successful emission. ## Testing Use `runTest` with the virtual clock; `delay` is skipped/advanced automatically, so backoff tests run instantly while still verifying attempt counts and which exceptions are retried.
- Why add jitter instead of pure exponential backoff?To break synchronization between many clients that failed at once, avoiding a thundering-herd retry spike that re-overloads the server.
- How do you test backoff without waiting real seconds?Run in runTest; the virtual time scheduler fast-forwards delay() so backoff completes instantly while attempt counts stay accurate.
- What must be true about the upstream for safe retries?It must be idempotent, because the producer block re-executes on every retry.
Like polite knocking that backs off — knock, wait a bit, wait longer, give a random pause so you and everyone else aren't all banging at once — and stop if the door is permanently locked.
saying these in an interview costs you the question
- Retrying every exception, including non-retryable 4xx or parse errors
- No cap on backoff delay (unbounded waits)
- Omitting jitter and ignoring thundering-herd risk
- Retrying non-idempotent operations without an idempotency key
- Trying to swallow CancellationException in the predicate