A service starts degrading and every client begins retrying, which keeps it down. How would you set retry policy across clients — backoff, jitter, attempt caps, budgets — and what should the API communicate through the HTTP Retry-After response header?
answer
- Retries multiply load exactly when capacity is gone
- Full jitter or you build synchronised waves
- Retry budget ≈ 10% of successes caps amplification
- Retry at one layer only — layers multiply
- Retry-After on 429/503, and jitter it too
basics
~20 sRetries multiply load exactly when capacity is gone. Use exponential backoff with jitter, a small attempt cap, and a retry budget capping retries as a fraction of successful traffic. Retry at one layer only, honour Retry-After on 429 and 503, and shed load rather than queueing it.
solid answer
~60 sA retry storm is a positive feedback loop: errors cause retries, retries add load, load causes more errors. **Client policy** - Exponential backoff **with full jitter** — unjittered backoff resynchronises clients into waves. - A small attempt cap (2-3 total) rather than "retry until success". - A **retry budget**: allow retries only while they stay under roughly 10% of that client's successful request rate, and stop entirely once the ratio blows out. This is the most effective single control, because it bounds amplification regardless of the error rate. - **Retry at one layer only.** Three layers each retrying three times is 27 requests; pick the layer with the most context and let the others pass failures through. - Respect the request deadline — never retry past it. **Server side** - Return **429** or **503** with `Retry-After` when shedding, and make shedding cheap so rejection costs far less than serving. - Prioritise first attempts over retries where you can distinguish them (an attempt-count header helps). - Add a circuit breaker so a dead dependency fails fast instead of consuming timeouts.
go deeper
Know that retries add load during an outage and that you need exponential backoff, jitter, and a limited number of attempts.
Add honouring Retry-After on 429 and 503, not retrying 4xx, and bounding retries by the request deadline.
Bring in retry budgets, single-layer retry policy, circuit breakers, cheap load shedding, and the metrics that reveal a storm early.
Frame it as a metastable-failure and capacity problem: amplification factors, per-operation-class policy encoded in shared clients, prioritising first attempts over retries, and staged recovery.
## The failure mode Suppose a service running at 70% capacity loses a third of its instances. Latency rises, timeouts fire, clients retry. If every failed request produces two retries, offered load *triples* at the exact moment capacity fell. The service cannot recover even after the original trigger is gone, because retry traffic alone exceeds what it can serve. This is metastable failure: the system stays down under load it created itself, and only load shedding or turning clients off breaks the loop. ## Client-side controls, in order of value **1. Retry budget (adaptive throttling).** Track a moving ratio of retries to successful requests per client process, and permit a retry only while retries stay under roughly 10% of successes. Under healthy conditions almost nothing is refused; during an outage the budget dries up and amplification is capped near 1.1x instead of 3x. Fixed policies cannot do this — a cap of 3 attempts still triples load when *everything* fails. If you take one idea from this topic, take this one. **2. Jitter.** Exponential backoff alone converts random arrivals into synchronised waves: everyone fails at t, everyone retries at t+1s, t+3s, t+7s. Full jitter — sleeping a uniform random value in `[0, base * 2^attempt]` — spreads the load flat. Deterministic backoff is one of the most common real bugs in retry code. **3. Attempt caps and deadlines.** Two or three total attempts, always bounded by the caller's remaining deadline. Propagate that deadline downstream so nobody retries work whose answer is no longer wanted. **4. Retry at one layer.** Amplification is multiplicative across layers: SDK, service client, gateway and a job scheduler each retrying gives 3×3×3×3. Choose the layer that knows whether the operation is replay-safe and whether the deadline is alive — usually the service client — and make the others transparent. An attempt-count header propagated end to end makes this auditable and lets servers deprioritise retries. **5. Circuit breaking.** After a threshold of failures to a dependency, fail fast for a cooldown and probe with a trickle. This protects your own threads and latency as much as the callee. ## Server-side obligations - **Say when to come back.** `429 Too Many Requests` and `503 Service Unavailable` carry the `Retry-After` header, in seconds or as an HTTP date. It is the only in-band way to coordinate a fleet you do not control. Jitter it per response, or you have just scheduled a synchronised thundering herd yourself. - **Shed cheaply.** Rejection must cost far less than service, otherwise shedding does not help. Reject at the edge, before expensive auth or database work. - **Prefer shedding to queueing.** A deep queue converts overload into latency, so every request times out, gets retried, and is served after nobody wants it. Bound queues and drop early. - **Distinguish retries.** If clients send an attempt counter, drop high-attempt requests first: you preserve first-attempt throughput, which is what users experience. - **Do not retry what will not help.** Never retry 4xx (except 429 and a token-refresh 401), and only retry idempotent or explicitly replay-safe operations. ## Getting out of a storm in progress The budget usually does it automatically. If not: shed aggressively at the edge, temporarily disable client retries via remote config if you own the clients, scale out, and only then let traffic back. Recovery must be gradual — restoring full traffic to a cold, empty-cache service just re-triggers the collapse. ## What to instrument Retry rate as a share of total requests, attempt-number distribution, budget-exhaustion events, and `Retry-After` compliance. A rising retry share is the earliest warning of a storm and should page before the error rate does. ## The judgment part There is no universal setting. High-value, low-volume operations (a payment) deserve more attempts and human escalation; high-volume reads deserve near-zero retries and a cached fallback. Decide per operation class, encode it in a shared client so teams do not each invent a policy, and treat retry configuration as a capacity decision rather than a code detail.
- Why is a retry budget better than simply capping attempts at three?A fixed cap still triples offered load in the exact scenario that matters, when nearly every request is failing. A budget is proportional to success: it permits retries only while they stay a small fraction of successful calls, so during a broad outage retries dry up automatically and amplification stays near 1.1x. It adapts to conditions instead of assuming failures are rare and independent.
- Why must Retry-After values be jittered by the server?Because every client receiving the same value returns at the same instant, so the server schedules its own thundering herd for a few seconds later. Adding a random spread per response — or having clients apply jitter on top of the advertised value — distributes the returning traffic and lets the service recover gradually instead of collapsing again on the first wave.
- What is wrong with queueing overload instead of rejecting it?A deep queue turns overload into latency, so requests wait past their client's timeout and are retried while the server is still working on the abandoned original. You end up spending full capacity producing responses nobody is waiting for. Bounded queues with early drop, plus fast cheap rejection, keep the work the server does actually useful.
Everyone redialling a jammed switchboard the instant the line drops: the calls that might have connected are crowded out by redials, and the jam sustains itself long after the original fault.
saying these in an interview costs you the question
- Exponential backoff with no jitter, producing synchronised retry waves
- Retrying at several layers at once without noticing the multiplication
- Treating 'retry until success' as robustness
- Ignoring the Retry-After header, or sending an identical value to every client
- Unbounded request queues, on the theory that no work should ever be rejected