skip to content

Multiple independent clients all start retrying a failing call to the same dependency using the exact same exponential backoff schedule. Why can this still overload the dependency, and how does adding 'jitter' — specifically full jitter versus equal jitter — help?

level: middleimportance: must knowfreq 70%

answer

  1. full jitter: random(0, cap)
  2. equal jitter: cap/2 + random(0, cap/2)
  3. AWS 'Exponential Backoff and Jitter' post
  4. thundering herd across many clients, not just one
  5. periodic spikes = missing jitter signature

basics

~20 s

If every client waits the same amount of time before retrying, they all retry at once, creating a new spike. Jitter adds randomness to each client's wait time so retries spread out instead of arriving together.

solid answer

~50 s

Exponential backoff controls how the delay grows per client, but if every client independently follows the same schedule and failed around the same time, their retries land at the same moments — 1s, 2s, 4s later, in lockstep — recreating a thundering-herd spike at each retry round instead of one continuous spike. Jitter randomizes the actual delay used each attempt. Full jitter picks a delay uniformly at random between 0 and the computed exponential cap (e.g., between 0 and 8s for that round), maximizing spread but sometimes retrying very soon. Equal jitter keeps half the exponential delay fixed and randomizes only the other half (e.g., 4s + random(0,4s)), guaranteeing a minimum wait while still spreading retries, at the cost of less spread than full jitter. Full jitter generally produces the lowest aggregate load and total completion time in AWS's published analysis, but equal jitter avoids near-zero-delay retries.

go deeper

for a junior

Should grasp that many clients retrying on the same schedule can retry 'at the same time' and that jitter adds randomness to spread that out.

for a middle

Should be able to describe at least one jitter formula (full or equal) concretely and explain why it helps beyond plain backoff.

for a senior

Should compare full jitter and equal jitter's trade-offs (spread vs. predictability) and reference how to diagnose the difference from production traffic patterns.

for a principal

Should reason about which jitter strategy fits which class of system (high-fanout backend calls vs. latency-sensitive or single-queue-consumer paths) and design the jitter window as a deliberate SLO trade-off.

## Backoff alone leaves clients correlated Exponential backoff solves the 'how fast should I retry' problem for a single client, but it does not by itself solve a second, distinct problem: correlated retries across a large population of independent clients. Picture a dependency that has a brief availability blip lasting one second, during which every one of a thousand concurrent callers gets a failure at roughly the same moment. If all thousand clients use the identical deterministic exponential schedule — say 1s, 2s, 4s, 8s — then every single one will retry at t+1s, and if that retry also fails (because the dependency, now facing a synchronized spike of a thousand simultaneous requests, buckles under it), all thousand will retry again at t+3s (1+2), then t+7s, and so on. The backoff schedule successfully spaces out each individual client's own retries, but it does nothing to desynchronize different clients from each other — they all 'wake up' together, recreating exactly the **thundering-herd** problem backoff was meant to solve, just delayed and repeated at each retry round instead of happening once. This failure mode is easy to miss in testing (where a single client is exercised in isolation) and only shows up under real production load with many concurrent callers. ## The two most commonly discussed strategies Jitter fixes this by injecting randomness into the delay so clients that failed at the same instant do not retry at the same instant. There are several standard jitter strategies, and the two most commonly discussed are full jitter and equal jitter, both described in AWS's widely-cited 'Exponential Backoff and Jitter' architecture blog post. - **Full jitter** computes the exponential cap for the current attempt (`cap = min(maxDelay, base * factor^attempt)`) and then picks the actual delay as a uniform random value between 0 and that cap: `delay = random(0, cap)`. This maximizes the spread of retry times across the population, since some clients will retry almost immediately and others will wait nearly the full cap, which empirically produces the lowest total number of retries and the shortest aggregate completion time across a population of failing clients, per AWS's own simulation results. Its downside is that any individual retry might land very close to zero delay, so a single client can occasionally retry almost immediately after a failure. - **Equal jitter** addresses that specific concern by keeping half of the computed exponential delay as a fixed floor and randomizing only the remaining half: `delay = cap/2 + random(0, cap/2)`. This guarantees every client waits at least some minimum amount before retrying, trading away some of full jitter's spread (and therefore some throughput advantage) for a more predictable lower bound on retry timing. ## Spread versus predictability The trade-off across jitter strategies is spread versus predictability. | Randomization | What it gives you | |---|---| | More randomness (full jitter) | minimizes the odds of synchronized retry waves and generally produces the best system-wide outcome, but makes any single client's retry timing highly variable, complicating reasoning about worst-case latency for a specific request | | Less randomness (equal jitter, or no jitter) | is easier to reason about per-request but reintroduces correlation risk across the client population | The choice usually favors more jitter for backend-to-backend service calls at scale — where system-wide load pattern matters more than any single request's exact timing — and can favor equal jitter for latency-sensitive paths with fewer concurrent callers, where the goal is a predictable floor on retry delay. ## The signature on a dashboard Failure modes without jitter show up as periodic load spikes on a dependency's dashboard at intervals matching the backoff schedule (a spike every ~1s, ~3s, ~7s after an initial blip, tapering as clients give up or succeed) rather than a smooth decay — a recognizable signature distinguishing 'backoff without jitter' incidents from plain retry-storm-without-backoff incidents, which instead show sustained elevated load with no periodicity. ## Where it shows up A concrete real-world example is AWS's own guidance for its SDKs: default retry behavior in recent AWS SDK major versions uses exponential backoff with jitter specifically because AWS observed synchronized retry spikes from large customer fleets during transient service degradations, and jitter was the fix that spread that load out smoothly instead of in discrete waves.

  • How would you distinguish, from a dependency's traffic graph, a retry storm caused by missing backoff versus one caused by missing jitter?
    Missing backoff typically shows sustained, non-decreasing elevated load because retries keep firing at a constant or near-constant rate without the delay growing. Missing jitter (but present backoff) instead shows a periodic pattern of spikes at intervals matching the backoff schedule, since the population of clients retries in synchronized waves that space out over time as the schedule grows, rather than a smooth continuous elevation.
  • Why might full jitter still cause a brief spike in retry load right after an initial failure event, despite randomizing the delay?
    Because full jitter draws uniformly between 0 and the cap, a nontrivial fraction of clients will draw values close to zero purely by chance, so a burst of near-immediate retries can still occur right after the initial failure. This is generally much smaller and less synchronized than a no-jitter thundering herd, but it's why some systems combine jitter with a small minimum delay floor for the first retry.
  • Would you apply the same jitter strategy to a client-to-client backend call as to a retry on a message queue consumer processing a poison message?
    Not necessarily — a backend-to-backend call under load benefits most from full jitter's population-level spread, since many instances are likely retrying concurrently. A queue consumer retrying a single message is more concerned with eventually distinguishing a transient failure from a permanently failing message, so it typically pairs backoff/jitter with a max-attempt count and a dead-letter queue rather than optimizing for population-level spread.

Like a fire drill where every office on every floor is told to reassemble exactly 5 minutes after the alarm — everyone floods the stairwell at the same instant. Jitter is like staggering each floor's assembly time by a random minute or two so the stairwell never gets a synchronized crush.

saying these in an interview costs you the question

  • Thinks exponential backoff alone prevents thundering herd across many clients
  • Can't explain why identical deterministic schedules across clients are a problem
  • Confuses jitter with backoff itself, treating them as the same mechanism
  • Doesn't know full jitter can occasionally produce near-zero delays
  • Assumes more randomness is always strictly better with no trade-off

context