skip to content

AWS's own SDKs default to a formula sometimes called 'decorrelated jitter': sleep = min(cap, random_between(base, previous_sleep * 3)). How does this differ mechanically from full jitter (random(0, min(cap, base * 2^attempt))), and why might it be preferred for long retry sequences?

level: principalimportance: nice to knowfreq 30%

answer

  1. sleep = random_between(base, previous_sleep*3), capped
  2. depends on previous actual delay, not just attempt index
  3. smoother, less choppy curve than full jitter
  4. AWS SDK default in several implementations
  5. secondary tuning choice vs. having jitter at all

basics

~20 s

Full jitter recalculates a random delay from scratch each time based on the attempt number. Decorrelated jitter instead bases each new random delay on the previous one, which tends to spread retries out more smoothly over a long sequence of attempts.

solid answer

~50 s

Full jitter computes each attempt's delay independently as a uniform random draw between 0 and an exponentially growing cap tied purely to the attempt index, so two consecutive draws are statistically uncorrelated — it's possible to draw a small delay, then a small delay again, back to back. Decorrelated jitter instead makes each delay a function of the previous delay (random_between(base, previous_sleep * 3), capped at a max), letting the delay wander upward with randomness baked into the growth itself, rather than the growth being deterministic and only the final draw being randomized. In AWS's own published simulations, decorrelated jitter produced a similar or slightly better spread of retries across a client population compared to full jitter, with less clustering of back-to-back short delays, making it AWS's SDK default in several official implementations. The practical difference from full jitter is usually secondary to just having jitter at all — the choice between the two is a minor tuning decision, not a correctness-critical one.

go deeper

for a junior

Not expected to know decorrelated jitter by name; full jitter/equal jitter is sufficient at this level.

for a middle

May recognize the name but isn't expected to reproduce the formula from memory.

for a senior

Should be able to describe how decorrelated jitter's dependence on the previous delay differs from full jitter's dependence on the attempt index.

for a principal

Should be able to reason about why this reduces cross-round re-clustering risk relative to full jitter, discuss the minor implementation trade-off (carrying state), and correctly frame it as a secondary tuning decision relative to jitter's primary role.

## How full jitter draws each delay Full jitter, as commonly used elsewhere, computes each retry's delay independently: for attempt number n, it derives a deterministic exponential cap (`cap_n = min(maxDelay, base * factor^n)`) purely from the attempt index, and then draws the actual delay as `random(0, cap_n)`. Because this draw is independent each time — the formula has no memory of what the previous delay actually was, only of the attempt count — it's statistically possible, though not typical, for a client to draw a small delay on attempt 3, and then draw another small delay on attempt 4, purely by chance, even though the theoretical cap has grown between the two. Two consecutive short draws don't violate anything about full jitter's design, but they mean a specific client's actual observed delay sequence can look choppier or less smoothly increasing than the underlying exponential cap curve would suggest. ## What decorrelated jitter changes Decorrelated jitter, the formula AWS documents in its retry SDK guidance, changes the input to the randomization: rather than deriving each delay purely from the attempt index, it derives it from the previous actual delay the client used. The formula is typically written as `sleep = min(cap, random_between(base, previous_sleep * 3))` — the next delay is drawn uniformly between the base delay and three times whatever delay was actually used last time, then capped at the maximum. This has a self-reinforcing quality: - if the previous draw happened to be small, the next draw's upper bound (3x the previous) is also small, keeping the sequence in a lower range for a bit; - if the previous draw was large, the next draw's upper bound is correspondingly large, letting the sequence range higher. Over many attempts, this produces a delay sequence that trends upward with organic-looking variability, rather than being pinned to a purely deterministic, attempt-indexed ceiling with independent noise layered on top. Critically, because each draw depends on the actual realized previous value, two different clients that failed at the same instant and are both running decorrelated jitter are extremely unlikely to end up with correlated sequences even after several rounds, since their random draws diverge and then compound divergently — whereas with full jitter, two clients at the same attempt index are drawing from the identical `[0, cap_n]` range each round, so nothing structurally prevents a coincidental re-clustering at some later round even though each round is independently randomized. ## What the published comparison found In AWS's own published simulation results (from the 'Exponential Backoff and Jitter' AWS Architecture blog post), decorrelated jitter performed comparably to or slightly better than full jitter on aggregate metrics like total number of retries needed and completion time across a simulated population of clients recovering from a shared failure, while producing a visually smoother, more organically increasing delay curve per client with less of the choppiness full jitter can exhibit. This is why several AWS SDKs (and AWS's own recommended reference implementation) ship decorrelated jitter as the default retry policy rather than full jitter, even though full jitter remains extremely common and perfectly reasonable elsewhere since it only depends on the attempt index and is simpler to implement. ## The practical trade-off The practical trade-off is modest: - decorrelated jitter requires the client to carry state between attempts (the previous sleep value) rather than being able to compute delay n from the attempt index alone, a small implementation complexity cost; - it also makes the delay sequence slightly harder to reason about or bound tightly in advance, since the next delay's range depends on what was actually drawn last time rather than being knowable purely from 'this is attempt 4.' For most systems, the choice between full jitter and decorrelated jitter is a secondary tuning decision rather than a correctness-critical one — the primary, load-bearing decision is having some jitter at all versus none, since that's what prevents synchronized thundering-herd retries in the first place; the specific jitter formula chosen affects the smoothness and marginal efficiency of the spread, not whether the fundamental thundering-herd problem is solved. ## Where it shows up A concrete real-world reference is the AWS SDK for Java v2 and the AWS SDK for Go v2, both of which document `FULL_JITTER` and a decorrelated-style backoff strategy as selectable retry policies, with recent SDK major-version defaults favoring the jittered strategies over legacy fixed-delay retry to reflect the operational lessons AWS accumulated from large-scale client fleets. Engineers building custom retry logic for a high-fanout internal service — where many client instances are likely to fail together against a shared dependency — are the ones most likely to benefit from choosing decorrelated jitter deliberately over full jitter, while for lower-fanout, less-correlated failure scenarios the distinction rarely matters in practice.

  • What state does a decorrelated-jitter implementation need to maintain that a full-jitter implementation does not?
    It needs to carry forward the actual delay value used on the previous attempt, since the next delay's range (base to previous_sleep * 3) is computed from it. Full jitter, by contrast, can compute each attempt's delay purely from the attempt index and doesn't need to remember anything about prior draws.
  • Is the difference between full jitter and decorrelated jitter likely to be the deciding factor in whether a system experiences a retry storm?
    No — the presence of jitter at all versus its absence is the load-bearing factor in preventing synchronized retries; the choice between full and decorrelated jitter is a secondary refinement affecting the smoothness and marginal efficiency of the spread, not whether the basic thundering-herd problem is solved. A system with either jitter strategy correctly implemented is far better off than one with backoff but no jitter at all.
  • Why are two different clients running decorrelated jitter, having failed at the same instant, unlikely to re-synchronize their retries even after several rounds?
    Because each client's next delay depends on its own actual previous delay rather than a shared, attempt-indexed range, small random differences between the two clients' early draws compound and diverge over subsequent rounds instead of both being redrawn from an identical range each time. This makes coincidental re-clustering across rounds less likely than with full jitter, where every client at the same attempt index draws from the exact same range regardless of its own prior draws.

Full jitter is like each runner in a relay independently rolling dice each leg to decide their pace, unrelated to how fast the last runner went. Decorrelated jitter is like each runner's pace being randomly nudged relative to the previous runner's actual pace, so the overall relay's rhythm drifts naturally rather than resetting randomly at every leg.

saying these in an interview costs you the question

  • Thinks decorrelated jitter and full jitter produce identical delay distributions
  • Doesn't know decorrelated jitter depends on the previous actual delay, not the attempt index
  • Claims decorrelated jitter is stateless like full jitter
  • Presents the full-vs-decorrelated jitter choice as more important than having jitter at all
  • Can't explain what state must be carried across retries for decorrelated jitter to work

context