skip to content

Your service returns HTTP 503 with a correct Retry-After header during a partial outage, clients obey it, and the service still collapses when it comes back. What retry behaviour do you require from clients and what do you enforce server-side?

level: principalimportance: should knowfreq 38%

answer

  1. Retry-After says when, not how many
  2. one value = synchronised herd at recovery
  3. full jitter + deadline + breaker + retry budget
  4. server: randomise delay, shed cheap, ramp gradually
  5. 3 layers x 3 retries = 27 attempts

basics

~20 s

Obeying a single Retry-After synchronises every client into one burst at recovery. Require exponential backoff with jitter, capped attempts, an overall deadline and circuit breakers; server-side, randomise the advertised delay, shed load, and ramp capacity back gradually.

solid answer

~60 s

`Retry-After` tells clients *when*, not *how many*. If ten thousand clients are told to wait sixty seconds, they all return in the same second and the recovering service dies on the first breath — a thundering herd created by your own correct header. **Client contract** (documented, and shipped in your SDKs so it isn't optional): exponential backoff with **full jitter** on top of the `Retry-After` floor; a bounded attempt count *and* an overall deadline; a **circuit breaker** so a client stops attempting entirely once failures dominate; and no retries at all for permanent errors. **Server side**, because you cannot trust third-party clients: **vary the advertised delay per response** so the herd is spread by construction; keep a load shedder that rejects cheaply at the edge rather than doing work it will drop; ramp acceptance gradually on recovery instead of opening at full capacity; and consider per-caller quotas so one misbehaving client cannot consume recovery capacity. Also watch **retry amplification across layers**: three tiers each retrying three times is twenty-seven attempts at the bottom. Retry at one layer, usually the outermost.

code

http · 5 lines
http
HTTP/1.1 503 Service Unavailable
Retry-After: 47
Content-Type: application/problem+json

{"type":"https://api.example.com/problems/unavailable","title":"Temporarily unavailable","status":503,"retryable":true}

go deeper

for a junior

Recognise that many clients retrying at the same moment creates a burst, and that random jitter spreads them out.

for a middle

Describe backoff with jitter, attempt limits and deadlines, and why the server should randomise the advertised delay.

for a senior

Add circuit breakers, cheap edge-level load shedding, gradual capacity ramp on recovery, and cross-layer retry amplification.

for a principal

Own the whole loop: client obligations shipped in SDKs, server-side defences that assume uncooperative clients, retry budgets, fairness and prioritisation under scarcity, and game-day validation.

## Why correct headers still produce collapse `Retry-After` answers *when* a single client should come back. It says nothing about *how many* clients come back, or how many times. Broadcasting one identical value to every rejected caller is an act of synchronisation: a population that arrived spread out over a minute is now aligned to a single instant. The recovering service — cold caches, empty connection pools, unwarmed JIT, reconnecting database sessions — meets peak load at its weakest moment and fails again. That failure produces another synchronised `Retry-After`, and the system oscillates. ## What clients must do **Jitter, not just backoff.** Exponential backoff alone still leaves a cohort that failed together retrying together, because they share the same schedule. Randomisation is what breaks the correlation — the widely-used "full jitter" approach picks a random wait in `[0, backoff]` rather than backoff exactly. `Retry-After` sets the floor; jitter spreads the population above it. **Bound by attempts and by a deadline.** Attempt count alone doesn't bound wall-clock time under long backoffs; a deadline alone can permit many attempts. Real systems need both, and the deadline should derive from the caller's own timeout budget — retrying after the user has already given up is pure waste. **Circuit breakers.** Backoff moderates one caller's rate; it does not stop it. When a dependency is clearly down, a breaker opens and requests fail immediately without touching the network. This protects the calling service too — threads and connections that would block on a dead dependency stay free — and gives the recovering service quiet time. Half-open probing lets a small number of trial requests test recovery instead of the whole population. **Retry budgets.** A stronger control than per-request limits: cap retries as a *fraction* of overall traffic (e.g. retries may not exceed 10% of requests). Under broad failure the retry rate is capped by construction, no matter how many individual requests are failing — which per-request attempt limits cannot guarantee. **No retries on permanent errors.** Retrying a 400 or 403 adds load and never succeeds. ## Why the server cannot rely on any of that You control your own SDKs; you do not control a customer's hand-rolled `while` loop. Assume some fraction of traffic is badly behaved and defend accordingly. **Randomise the advertised delay.** Instead of `Retry-After: 60` for everyone, send values sampled from a range. The herd is spread whether or not clients implement jitter — the single highest-leverage server-side fix, and it costs nothing. **Shed load cheaply and early.** Rejecting at the edge before authentication, database access or business logic keeps rejection cost far below the cost of serving. A rejection that costs as much as a success provides no relief. **Ramp capacity back gradually.** On recovery, accept a fraction of offered load and increase it as health metrics hold. Opening at 100% into a synchronised herd is what causes the second collapse. Concurrency limiting with adaptive thresholds does this continuously rather than only at recovery. **Fairness under scarcity.** Per-caller quotas ensure recovery capacity isn't consumed entirely by whichever client retries most aggressively. Without them, the worst-behaved caller is rewarded. **Prioritise.** During recovery, serve interactive traffic before batch and background work; the retry storm is disproportionately automated traffic. ## Retry amplification across layers If every tier in a chain retries three times, the bottom tier sees up to 27 attempts per original request. This is one of the most common ways a small dependency blip becomes a total outage. The rule: **retry at one layer**, normally the outermost that still has the context to retry meaningfully, and have inner layers fail fast and propagate. Where multiple layers must retry, budget the total explicitly rather than letting each choose independently. ## Verifying it None of this can be assumed to work. Game-day exercises — kill a dependency, watch what the retry population does at recovery — are what surface the third-party client that ignores `Retry-After` entirely, or the internal service with three nested retry layers nobody had counted. Instrument retries as a distinct metric from first attempts, so you can see the storm forming rather than inferring it afterwards.

  • Why is exponential backoff without jitter insufficient?
    Clients that failed at the same moment share the same backoff schedule, so they stay synchronised and arrive together at every subsequent attempt — the doubling only changes when the burst lands, not that it is a burst. Randomising the wait decorrelates the population. Full jitter, choosing uniformly in the interval from zero to the computed backoff, is the common approach.
  • What is a retry budget and why is it stronger than a per-request attempt limit?
    A retry budget caps retries as a proportion of total traffic, for example allowing retries to be at most ten percent of requests. Per-request limits bound each individual call but say nothing about aggregate load, so during a broad outage every request retrying its permitted maximum still multiplies total traffic. A budget bounds the aggregate directly, which is the quantity that actually harms the dependency.
  • How do you decide which layer in a multi-tier call chain should retry?
    Normally the outermost layer that still has enough context for a retry to be meaningful and knows the caller's remaining deadline; inner layers fail fast and propagate. Retrying at every tier multiplies attempts geometrically, so a three-deep chain retrying three times each generates up to twenty-seven attempts at the bottom. If multiple layers genuinely must retry, budget the total explicitly rather than letting each choose independently.

saying these in an interview costs you the question

  • Assuming correct Retry-After handling by clients is sufficient to prevent recovery collapse
  • Applying exponential backoff without jitter and expecting the herd to disperse
  • Bounding retries only by attempt count, with no deadline, circuit breaker or aggregate budget
  • Retrying independently at every tier of a call chain, multiplying load geometrically
  • Restoring full capacity immediately on recovery rather than ramping acceptance gradually

context