skip to content

Implement exponential backoff with jitter for a flaky network Flow using retryWhen, retrying only transient errors. Walk through the design.

level: seniorimportance: should knowfreq 45%

answer

  1. base * 2^attempt, then coerceAtMost(maxDelay)
  2. Add jitter to avoid thundering herd
  3. Only retry transient errors (IO, 5xx, 429)
  4. Fail fast on 4xx / non-retryable
  5. Upstream must be idempotent; test with runTest virtual clock

basics

~10 s

Use retryWhen: check the error is retryable and you haven't hit the cap, wait a growing amount of time (doubling each try) plus a small random delay, then retry; otherwise give up.

solid answer

~40 s

retryWhen gives you cause and a zero-based attempt counter and is a suspend lambda, so you compute the delay and call delay() before returning true. Exponential backoff = base * 2^attempt, capped at a max to avoid runaway waits; add random jitter (e.g. +/- a fraction) so many clients don't retry in lockstep (thundering herd). Gate retries on retryability: only retry transient causes (IOException, HTTP 5xx/429) and a maximum attempt count; return false for non-retryable causes (4xx, serialization errors) so they fail fast. Never retry CancellationException — retryWhen already excludes it, but don't write predicates that catch it. Keep the upstream idempotent since it re-runs. Use coerceAtMost for the cap and Random for jitter.

code

kotlin · 9 lines
kotlin
fun <T> Flow<T>.retryTransient(maxRetries: Long = 5): Flow<T> =
    retryWhen { cause, attempt ->
        if (attempt >= maxRetries || !isTransient(cause)) false
        else {
            val backoff = (100L * (1L shl attempt.toInt())).coerceAtMost(5_000)
            delay(backoff + Random.nextLong(0, backoff / 2 + 1))
            true
        }
    }

go deeper

for a junior

Can call retryWhen with a delay but may retry all errors and forget caps.

for a middle

Implements capped exponential backoff and a retryable-error gate correctly.

for a senior

Adds jitter, fail-fast classification, idempotency reasoning, and virtual-clock testing.

for a principal

Weighs system-wide effects: thundering herd, circuit breaking, budget/SLA on retries, and where retry belongs in the resilience stack.

## The shape of the solution ```kotlin private val RETRYABLE = setOf(500, 502, 503, 504, 429) fun isTransient(cause: Throwable): Boolean = when (cause) { is IOException -> true is HttpException -> cause.code in RETRYABLE else -> false } fun <T> Flow<T>.retryTransient( maxRetries: Long = 5, base: Long = 100, // ms maxDelay: Long = 5_000 // ms cap ): Flow<T> = retryWhen { cause, attempt -> if (attempt >= maxRetries || !isTransient(cause)) { false // give up: too many tries OR non-retryable } else { val exp = base * (1L shl attempt.toInt()) // base * 2^attempt val capped = exp.coerceAtMost(maxDelay) val jitter = Random.nextLong(0, capped / 2 + 1) // full-ish jitter delay(capped - capped / 2 + jitter) true } } ``` ## Why each piece exists - **`retryWhen`, not `retry`** — we need conditional logic (retryable check), a growing `delay`, and access to `attempt`. `retry(n)` cannot delay. - **`attempt` is zero-based** — `1L shl attempt` doubles the base each retry: attempt 0 -> base, 1 -> 2x, 2 -> 4x. This is **exponential backoff**: waits grow geometrically so a struggling server gets breathing room. - **`coerceAtMost(maxDelay)`** — caps the wait so attempt 10 doesn't sleep for minutes. - **Jitter** — adding randomness prevents the **thundering herd**: if 1000 clients fail at the same instant and all back off by exactly the same schedule, they retry in sync and re-overload the server. Randomizing spreads them out. - **Retryability gate** — retrying a 400 Bad Request or a deserialization error is pointless; those will fail identically forever. Return `false` to **fail fast** on permanent errors and reserve retries for **transient** ones (network blips, 503, rate-limit 429). - **CancellationException** — `retryWhen` never retries it; don't let `isTransient` accidentally classify it as retryable (it won't here, but a naive `else -> true` would be wrong — and even then the operator guards it). ## Idempotency Because the upstream `flow { }` re-runs on every retry, the producer must be **idempotent** (safe to repeat). Retrying a non-idempotent POST that already partially succeeded can double-charge a customer. For unsafe operations, add an idempotency key or retry at a layer where repetition is safe. ## Placement Put `retryTransient()` **directly above** the operator that can fail and **below** mapping/UI logic, so only the network step is retried and downstream transforms run once per successful emission. ## Testing Use `runTest` with the virtual clock; `delay` is skipped/advanced automatically, so backoff tests run instantly while still verifying attempt counts and which exceptions are retried.

  • Why add jitter instead of pure exponential backoff?
    To break synchronization between many clients that failed at once, avoiding a thundering-herd retry spike that re-overloads the server.
  • How do you test backoff without waiting real seconds?
    Run in runTest; the virtual time scheduler fast-forwards delay() so backoff completes instantly while attempt counts stay accurate.
  • What must be true about the upstream for safe retries?
    It must be idempotent, because the producer block re-executes on every retry.

Like polite knocking that backs off — knock, wait a bit, wait longer, give a random pause so you and everyone else aren't all banging at once — and stop if the door is permanently locked.

saying these in an interview costs you the question

  • Retrying every exception, including non-retryable 4xx or parse errors
  • No cap on backoff delay (unbounded waits)
  • Omitting jitter and ignoring thundering-herd risk
  • Retrying non-idempotent operations without an idempotency key
  • Trying to swallow CancellationException in the predicate

context