A microservices call chain is A -> B -> C -> D, and each service independently retries failed calls to its downstream up to 3 times with backoff. When D starts failing, what happens to the total request volume D's failure generates upstream, and how does a retry budget limit this?
answer
- multiplicative not additive across hops: 3x3x3=27
- budget = % of traffic, not count per request
- Envoy/gRPC/Finagle implement retry budgets
- exhausted budget -> fail fast, no retry
- per-dependency budget, not global
basics
~20 sIf every layer retries 3 times independently, one failure at the bottom can multiply into dozens of retries flowing back up through the chain — a retry storm. A retry budget caps the fraction of a service's total traffic that's allowed to be retries, so it stops amplifying once that cap is hit.
solid answer
~50 sWith per-hop retries stacked across a chain, retry counts multiply rather than add: if C retries D up to 3x, B calling C sees each of C's already-3x'd calls potentially fail and retries C up to 3x, and A retries B up to 3x — the theoretical worst case is 3x3x3=27 downstream attempts at D for a single original request from A, which is retry amplification. A retry budget caps retries as a percentage of total request volume over a rolling window (e.g., 'retries may not exceed 10% of requests') rather than allowing unlimited per-request retry counts; once the budget is exhausted, further failures are surfaced immediately without retrying. This decouples the retry limit from any single request's attempt count and ties it instead to the aggregate health of the system, preventing the multiplicative blow-up that per-hop fixed-attempt-count retries alone don't prevent.
go deeper
Should recognize, at a high level, that retries at multiple layers of a system can add up to a lot of extra traffic when something downstream fails.
Should be able to explain that per-hop retry counts compose multiplicatively across a call chain, even if they can't derive the exact worst-case number unprompted.
Should articulate the retry-budget mechanism precisely (percentage-of-traffic cap over a rolling window, fail-fast once exhausted) and why it succeeds where per-hop attempt caps alone don't.
Should discuss calibrating budget ratios against dependency capacity and organic error rate, per-dependency scoping of budgets, and reference real systems (Envoy, gRPC, Finagle) that implement this as a first-class mechanism.
## Per-hop retries multiply, they do not add In a single hop — one client calling one dependency — a retry policy with a max-attempts cap (say, 3) bounds that hop's worst-case amplification to 3x: one logical request can generate at most 3 downstream attempts. The problem in a multi-hop call chain is that this bound doesn't add across hops, it multiplies. Consider A calling B, B calling C, and C calling D, where each hop independently retries up to 3 times on failure. 1. If D starts failing, C retries its call to D up to 3 times before giving up and returning an error to B. 2. B, seeing that failure from C, retries its call to C up to 3 times — and each of those 3 calls to C can itself trigger up to 3 calls to D, so B's retries alone can generate up to 9 attempts at D. 3. A, retrying its call to B up to 3 times, can then trigger up to 3x9=27 attempts at D from a single original request at A. This compounding is called **retry amplification**, and it means a small, per-hop retry limit that looks perfectly reasonable in isolation can produce an enormous multiplier — in this example nearly two orders of magnitude — by the time it reaches the failing service at the bottom of the chain. In real systems with deeper chains or higher per-hop retry counts, the multiplier grows explosively, and this is a principal mechanism by which a single failing low-level dependency turns into a cascading, fleet-wide outage: the failing dependency is buried under amplified retry traffic from every layer above it, which keeps it failing, which keeps generating more retries, in a self-sustaining loop. ## What a retry budget limits instead A retry budget is the standard mitigation, and it works by changing what's being limited. Per-hop max-attempts caps limit retries per request; a retry budget instead limits retries as a fraction of aggregate traffic over a rolling time window, independent of how many individual requests are involved. A typical formulation (used by systems like Envoy proxy, gRPC, and Finagle) tracks, over the last N seconds, the ratio of retry attempts to original requests, and caps that ratio — for example, 'retries may not exceed 20% of the volume of original requests.' - Every retry attempt consumes a token from this budget; once the budget is exhausted, new retries are rejected outright and the failure is surfaced to the caller immediately instead of being retried. - Because the budget is shared across all requests flowing through that client, it caps the aggregate extra load a struggling dependency will see, regardless of how deep the call chain is or how many upstream layers are each independently retrying — a crucial property that per-hop attempt caps lack. A dependency that starts failing will see retry traffic rise but self-limit to a bounded percentage above its organic request rate, rather than facing an unbounded multiplicative spike. ## The trade-off, and picking the ratio The trade-off is between resilience for any single request and protection for the shared dependency. A retry budget means that during a real outage, once the budget is exhausted, requests fail fast rather than getting the 'one more chance' a per-attempt retry would have given them — a deliberate sacrifice of some individual-request success probability in exchange for preventing the aggregate storm that would otherwise make the outage worse and longer for everyone. Configuring the budget itself is a genuine design decision: - too generous a ratio (e.g., 90%) barely constrains amplification and defeats the purpose; - too strict a ratio (e.g., 1%) means even brief, easily-recoverable blips get surfaced as hard failures instead of smoothed over by a retry. Budgets are also typically implemented per downstream dependency, since a service might have a healthy budget's worth of headroom for one dependency while another is in active distress. ## Failure modes, and where it shows up Failure modes that retry budgets are specifically built to prevent show up as exactly the multiplicative pattern above: a single low-level dependency's transient failure escalating into full unavailability as retry traffic compounds through every layer of a deep call graph, often visible as a sharp, non-linear spike in request volume at the deepest, most-shared dependency (a primary database, a shared cache cluster, an auth service) that's disproportionate to the number of actual end-user requests in flight. A well-known real-world instance of this class of pattern is documented in Twitter's Finagle and Google's SRE literature: both describe retry-budget mechanisms explicitly motivated by observed cascading-failure incidents where naive per-hop retries at every layer of a deep service graph amplified a small backend blip into a much larger, longer outage, and both treat 'cap retries as a percentage of traffic, not a fixed count per request' as the fix.
- Why doesn't simply lowering each hop's max-attempts count (e.g., from 3 to 2) fully solve the amplification problem in a deep call chain?It reduces the multiplier's base but doesn't eliminate the multiplicative structure — a 4-hop chain with 2 retries per hop still multiplies to 2^4=16x in the worst case, which is still a large amplification for a single failing request. The fundamental issue is that per-hop caps compose multiplicatively across hops regardless of the exact per-hop number, whereas a shared retry budget caps the aggregate ratio directly.
- How would you decide the retry-budget percentage for a given client-to-dependency relationship in practice?It's typically calibrated against the dependency's known capacity headroom and the baseline organic error rate — a common starting point is a modest ratio like 10-20% of original request volume, then tuned based on observed incident behavior (whether budgets are getting exhausted during real transient blips, suggesting raising it, or whether amplification incidents still occur, suggesting lowering it or shortening the rolling window).
- If a service's retry budget is currently exhausted and a genuinely transient one-off failure occurs, what happens to that request?It fails fast without being retried, even though a retry might well have succeeded — this is the deliberate cost of the budget mechanism. The assumption is that during a period of sustained failures dense enough to exhaust the budget, protecting the downstream dependency from further amplified load matters more than squeezing out success for any individual request, since continued retrying at that point is more likely to prolong the outage than resolve it.
Like a rumor passed through several layers of middle management, where each manager 're-confirms' by re-asking their own team before passing it up — a single false alarm at the bottom can turn into dozens of re-confirmation calls flooding back down. A retry budget is like a rule that caps how many re-confirmation calls the whole chain is allowed to make in an hour, no matter how many separate rumors are moving through it.
saying these in an interview costs you the question
- Assumes retry counts add across hops in a call chain rather than multiply
- Doesn't recognize that a small per-hop max-attempts value can still produce large amplification in a deep chain
- Defines retry budget as just 'a max-attempts count', not distinguishing it from per-request retry limits
- Thinks a retry budget applies globally across all dependencies rather than per downstream target
- Believes exhausting a retry budget means the whole service stops serving traffic, rather than just failing fast on further retries