A service retries a failing dependency up to three times with a 2-second timeout per attempt, while its own caller expects an answer within 3 seconds. What is wrong with that configuration, and how would you fix it?
answer
- 3 x 2s > 3s budget: attempts run after the caller left
- retry inside remaining = deadline - now
- attemptTimeout = min(cap, remaining); backoff costs budget too
- never retry deadline-exceeded; only idempotent ops
- layers multiply: 3x3x3 = 27; budget + jitter + breaker
basics
~20 sAttempts multiply: worst case is over 6 seconds against a 3-second promise, so attempts 2 and 3 run after the caller has already given up — pure wasted load. Fix: retry inside the remaining deadline, set each attempt's timeout to min(cap, remaining), and skip retries when too little time is left.
solid answer
~60 sThe retry policy and the deadline were configured independently, so they contradict each other. Worst case is 3 x 2 s plus backoff — over 6 s — while the caller abandons at 3 s. Everything after that point is load with no possible benefit, and it arrives exactly when the dependency is already struggling. Fix the shape, not the numbers: - **Retry inside a budget.** Before each attempt compute `remaining = deadline - now`; if it is below the minimum useful attempt time, stop and return the failure now. - **Size each attempt** as `min(attemptCap, remaining)`, so the last attempt shrinks rather than overrunning. - **Only retry what is safe** — idempotent operations, or non-idempotent ones carrying an idempotency key — and only retriable failures, never a deadline-exceeded error. - **Cap amplification.** Nested retries multiply across layers (3 layers x 3 attempts = 27 calls), so retry at one layer, and add a retry budget — for example retries capped at a few percent of traffic — plus jittered backoff and a circuit breaker.
go deeper
Do the arithmetic: three two-second attempts cannot fit in a three-second promise, so later attempts are wasted.
Add the fix — check the remaining budget before each attempt and size the attempt timeout as the minimum of the cap and what is left.
Discuss amplification across layers, retry budgets, jittered backoff, circuit breaking, and which errors are retriable, including the ambiguity of timed-out writes.
Argue the systemic view: retries are load applied precisely when capacity is scarce, so the policy must be budget-derived and globally capped, with metastable-failure risk and idempotency guarantees treated as platform requirements.
## The arithmetic Three attempts of 2 seconds is 6 seconds of timeout, plus backoff between attempts — call it 6.5–7 s worst case. The caller's budget is 3 s. So in the worst case the first attempt uses 2 s, the caller gives up during the second, and the second and third attempts execute entirely inside a window where **no one will read the result**. The configuration cannot deliver its intent under exactly the conditions retries exist for. ## Why this is worse than merely useless Retries are correlated with dependency trouble. When the dependency slows down, every caller starts timing out, and each timeout triples the request rate against a service that is already overloaded. That positive feedback loop is a standard cause of *metastable failure*: even after the original trigger passes, the retry-amplified load keeps the system down until traffic is forcibly cut. Retries transform a partial brownout into an outage. Amplification also compounds with depth. If three layers each retry three times, one user request can become 27 leaf calls: ```text A (3 tries) -> B (3 tries) -> C (3 tries) = 27 calls at C ``` The usual rule is **retry at one layer only** — typically the one closest to the failure that can still judge idempotency — and have other layers fail fast. ## The fix: retries live inside a deadline budget ```text while true: remaining = deadline - now if remaining < minUsefulAttempt: return lastError # no time to try attemptTimeout = min(attemptCap, remaining) result = call(dependency, timeout = attemptTimeout) if result.ok or not retriable(result): return result sleep = min(backoffWithJitter(n), remaining - minUsefulAttempt) if sleep < 0: return lastError ``` Properties worth naming: total time is bounded by the deadline no matter how the retry count is configured; the final attempt is *shortened* rather than allowed to overrun; and backoff is itself charged to the budget, which people routinely forget. With a 3-second budget you might realistically get two attempts of about 1.2 s with a short jittered pause — chosen by the budget, not by a hard-coded count. ## What is retriable - **Never retry a deadline-exceeded error** by default. The budget is gone; a retry can only make things worse. - **Only retry safe operations.** Reads and genuinely idempotent writes are fine. A non-idempotent mutation may only be retried when it carries an idempotency key the server deduplicates on — remember that a timeout means *unknown*, so the first attempt may have succeeded. - **Distinguish failure classes.** Connect failures and explicit "unavailable"/"overloaded" responses are worth retrying, ideally on a different instance. Validation errors and authorization failures are deterministic — retrying them is pure waste. Honour explicit backpressure signals such as a retry-after hint. ## Additional safeguards **Backoff with jitter.** Fixed backoff synchronizes clients into waves; exponential backoff with randomization spreads them. Full jitter (a uniform random draw up to the current backoff) is the usual default. **Retry budgets.** Rather than a per-request attempt count, cap retries as a fraction of overall traffic — for example, retries may not exceed ~10% of successful requests over a rolling window. When the dependency is broadly failing, the budget exhausts and retries stop automatically, which is precisely the behaviour a fixed count gets wrong. **Circuit breaking.** After a sustained failure rate, stop calling for a cooldown and fail fast, then probe with a trickle. This protects both the dependency and the caller's own threads and connections. **Hedging is not retrying.** A retry starts after a failure; a hedge starts a second attempt while the first is still pending, to cut tail latency. They share the same budget and the same idempotency requirement, but solve different problems. ## The interview point The deep answer is not "reduce the timeout to 800 ms". Any hard-coded pair of (attempts, per-attempt timeout) is a latency contract expressed in the wrong currency: it drifts the moment anyone changes the topology or the caller's budget. Deriving attempt timeouts from the remaining deadline makes the policy self-consistent by construction, and makes the retry count an *outcome* rather than an input.
- Retrying only when time remains still means the dependency gets extra load exactly when it is struggling. How do you limit that?Cap retries globally rather than per request: a retry budget that allows retries only while they stay under a small fraction of successful traffic over a rolling window, so a broad outage silently disables them. Add exponential backoff with full jitter to prevent synchronized waves, and a circuit breaker that fails fast during sustained failure and probes with a trickle. Together these make retry load proportional to spare capacity instead of to the failure rate.
- If the dependency call is not idempotent, can you retry at all?Only with server-side deduplication. The client generates a unique key for the logical operation and sends the same key on every attempt; the server records it and returns the original outcome for repeats. Without such a mechanism, a timeout is an unknown outcome and a retry risks double-applying the change, so the safe options are to fail and let the user decide, or to query the operation's status before deciding.
It is like calling someone back three times after they have already left the building: each call is polite in isolation, and all of them ring an empty desk while the switchboard is on fire.
saying these in an interview costs you the question
- Treating the fix as "tune the numbers" instead of deriving attempt timeouts from the remaining budget.
- Retrying after a deadline-exceeded error.
- Retrying at every layer, multiplying one user request into dozens of leaf calls.
- Forgetting that backoff sleep consumes the deadline budget too.
- Assuming a timeout proves the operation did not happen, so a non-idempotent retry is safe.
- Using fixed backoff with no jitter, synchronizing clients into retry waves.