skip to content

Requests hitting a backend service triple within a minute while its success rate falls, and after you restart it, it saturates again within seconds. How do you confirm this is a retry storm rather than a genuine traffic surge, and what do you do to break it?

level: seniorimportance: must knowfreq 68%

answer

  1. load rising while success falls
  2. edge flat, interior tripled
  3. the trigger is gone, the state persists
  4. a cold restart has less capacity
  5. restore in steps, watch success rate

basics

~20 s

Retry storms show load rising as success falls, with internal call volume spiking while edge traffic stays flat — organic surges raise both together. Breaking one requires cutting offered load below the degraded service's reduced capacity, then restoring in steps.

solid answer

~50 s

The diagnostic signature is a correlation organic traffic never shows: offered load climbing *because* success is falling. Compare the request rate at the edge with the internal call rate — if the front door is flat while the backend sees 3x, the extra volume is being manufactured inside your system. Ratios of total attempts to unique request identifiers, or duplicate idempotency keys, confirm it. The important operational fact is that this is a metastable state: the original trigger may be long gone, but the retry load is now self-sustaining, so restarting the service just feeds it into the same wall — with less capacity than before, because its caches are cold and its connection pools empty. Breaking it means reducing offered load below that degraded capacity: shed hard at the edge, disable retries via a config push or flag, drop already-expired queued work, then restore traffic in steps — 10%, 25%, 50% — watching success rate at each step rather than reopening to 100%.

go deeper

for a junior

Know that a failing service can cause its callers to retry, and that those retries add load that keeps the service failing, so more traffic is not always more users.

for a middle

Be ready to name the evidence that separates a storm from organic growth — load rising as success falls, edge rate flat while internal rate spikes, attempts exceeding unique request identifiers — and explain why restarting alone does not clear it.

for a senior

Demonstrate the operator's sequence under time pressure: shed at the edge, kill retries by configuration, drop expired backlog, restore in steps against success rate, and explain why a cold service has less capacity than the one that just failed.

for a principal

Own the structural answer: retry budgets enforced in the shared client library, an incident lever for retry policy that exists before it is needed, amplification ratio as a standing metric, and a shed path cheap enough to absorb a storm platform-wide.

## Recognising it Organic traffic surges and retry storms both look like a spike on a request-rate graph, and telling them apart in the first two minutes changes what you do next. The distinguishing evidence: **Load rises as success falls.** Real user demand is indifferent to your success rate; it does not triple *because* you started returning errors. A tight inverse correlation between error rate and request rate is close to diagnostic on its own. **The edge is flat while the interior is not.** Compare requests arriving at the outermost layer with calls arriving at the struggling service. Organic growth moves both. A storm typically shows a flat or modestly raised edge and a large multiple internally, because the volume is being manufactured between them. **Attempts exceed unique requests.** If requests carry a request identifier, an idempotency key, or a trace identifier, the ratio of total attempts to distinct identifiers is the amplification factor, measured rather than guessed. A rise from near 1.0 to 3 or 5 is the storm quantified. **Latency at the caller is pinned at its timeout.** A caller whose latency distribution has collapsed onto its timeout value is a caller that is about to retry everything. ## Why it does not clear itself: metastability The crucial mental model is the metastable failure — a state the system stays in after the original trigger has gone, because the load the failure generates is enough to sustain the failure. Research on this (Bronson et al., HotOS 2021) named the pattern, but every on-call engineer has met it: the deploy was rolled back, the slow dependency recovered, and the service is still down. Two things make the trap deeper than people expect: **Degraded capacity is lower than nominal capacity.** A service that just restarted has cold caches, empty connection pools, un-JIT-compiled code and no warmed thread pools. If its cache hit ratio drops from 99% to 0, the load reaching the database behind it multiplies by roughly a hundred. So the load level it must be brought under to recover is well below the load it was comfortably serving an hour ago. This is precisely why "just restart it" fails and why restoring to 100% traffic at once fails again. **Well-behaved clients still amplify.** Correct exponential backoff with jitter reduces synchronisation and slows the rate of retries — it does not remove them. Every client adds attempts at exactly the moment the service has the least capacity to absorb them. Add the mechanisms that are not retries at all but behave identically in aggregate: health-checking clients failing away from an unhealthy replica onto the remaining ones, which *moves* load rather than reducing it; browsers and mobile apps whose users hit refresh; upstream schedulers re-queueing failed jobs. A candidate who says "we use jittered backoff, so we cannot have a retry storm" has missed the point. ## Breaking it The single principle: **offered load must be brought below the degraded service's actual capacity**, and no amount of restarting, scaling or tuning substitutes for that. In rough order of speed: 1. **Shed hard at the edge.** Admit a small fraction and reject the rest cheaply. If criticality tiers exist, keep only the top tier. Blunt but fast, and it is the lever that actually decides the outcome. 2. **Turn retries off.** A runtime configuration push or feature flag that sets client retry counts to zero for the affected dependency removes the amplification at its source. This is a good example of a control that must exist *before* the incident — a retry policy you can only change by editing code and deploying is not an incident lever. 3. **Drop the backlog.** Purge or deadline-expire queued work rather than serving requests whose callers vanished ten minutes ago. Serving the backlog is spending your recovered capacity on nothing. 4. **Restore in steps.** 10%, then 25%, 50%, 100%, holding at each level until success rate and latency are stable and caches have refilled. Watch success rate, not request rate — request rate looks healthy in a storm. 5. **Only then scale.** Adding instances during a storm often makes things worse: new instances arrive cold, take a share of the traffic they cannot serve, and add connection pressure on the same shared database. ## Preventing the next one The durable fixes are structural: retry budgets or caps enforced in the shared client library rather than left to each caller; a server-side shed that is cheap enough to absorb the storm; and treating the retry-amplification ratio as a first-class metric so the next storm is visible in seconds rather than inferred in minutes. ## The interview point The forced choice with a clock is whether to throw away traffic you could theoretically serve. Candidates who have actually carried a pager reach for load shedding early and restore in steps; candidates who have not reach for scaling up and for restarting, and describe an incident that lasts three hours.

  • Why can adding more instances during a retry storm make the incident worse?
    New instances start cold: empty caches, unwarmed pools, uncompiled hot paths. They immediately take a share of a load they cannot serve, so they fail too, and each one opens fresh connections against the same shared database that is already the constraint. You have added consumers of the bottleneck rather than capacity. Shed first, recover, then scale.
  • Your client library uses exponential backoff with full jitter. Why is that not sufficient protection?
    Backoff slows and desynchronises retries; it does not stop them. Each retry still lands when the service has the least capacity, and jitter does nothing about health-check failover moving load onto surviving replicas, users refreshing, or upstream schedulers re-queueing jobs. Backoff makes the storm gentler, but only a cap or budget, plus server-side shedding, bounds it.
  • Once the service is stable, why restore traffic in steps rather than all at once?
    Because the capacity you have just recovered is not the capacity you had before the incident — caches are still refilling and pools are still warming, so full traffic can push you straight back into the metastable state. Stepping up while watching success rate and latency lets you find the level the system currently supports and gives caches time to rebuild between steps.

A jammed revolving door does not clear when you stop pushing new people in from outside — the crowd already inside keeps shoving. You have to hold the entrance shut, let the interior drain, then let people back in a few at a time.

saying these in an interview costs you the question

  • Says jittered backoff makes a retry storm impossible
  • Restarts the service and expects it to recover on its own
  • Scales out during the storm before reducing offered load
  • Reads the traffic spike as genuine user demand
  • Restores to full traffic the moment errors stop

context