Your dispatcher resubscribes on every carrier failure without limit; what must a retry decision take into account before it resubscribes again?
answer
- unbounded retry is a loop
- ask what kind of failure
- some failures never clear
- bound by count or by deadline
- exhaustion carries the original cause
basics
~20 sA retry decision needs three inputs: the kind of failure, since some can never clear; a bound on attempts, as a count or a deadline; and some space between attempts. When the bound is reached, the failure must reach the subscriber.
solid answer
~40 sRetrying unconditionally turns a permanent failure into an infinite loop. The decision has to inspect the **failure itself** first: a rejected recipient, a malformed message or a refused authorisation will fail identically on every attempt, while a timeout or a refused connection may well clear. It then needs a **bound** — a maximum attempt count, or better a deadline, since the caller waiting on the result cares about elapsed time rather than attempt number. Between attempts it needs **space**: resubscribing immediately re-runs the same call against conditions that have not had a moment to change. Finally, when the decision says stop, the failure must continue downstream carrying the original cause, not be replaced by a bare `retries exhausted` with the cause dropped.
code
pseudocode · 10 linespipeline = send_once(message)
.retry_if(function(failure, attempt_number) {
if (failure.kind == PERMANENT) return false // never clears
if (attempt_number >= 4) return false // bound reached
wait(spacing_for(attempt_number))
return true
})
// when the decision returns false the sequence fails,
// carrying the original failure plus attempt_numbergo deeper
Know that a retry needs a limit and that not every failure is worth repeating. Be able to name one failure that can clear on its own and one that never will.
Separate the three inputs cleanly: the predicate reads the failure, the bound counts attempts or watches a deadline, and the spacing keeps attempts from becoming a tight loop. Explain why each is a different job.
Show what exhaustion must produce — the original cause with attempt context attached — and name the timeout ambiguity a predicate cannot resolve. Argue for a deadline over a count when a caller is waiting.
Decide what the platform requires before a retry may be attached at all: how failures must be typed so a predicate can read them, what bound shape is the default, and who is accountable for the duplicates retryable timeouts produce.
## Retrying is a decision, not a reflex Attaching a retry with no conditions says `repeat this forever, whatever went wrong`. For an outbound notification dispatcher that is a specific and expensive failure mode: a message addressed to a recipient the carrier will never accept is resubscribed without end, each attempt re-running every stage above the retry point, while the real fault — a bad address — never surfaces anywhere a human would see it. A retry decision has three inputs, and a design that leaves out any one of them has a known way of going wrong. ## Input one: which failure is it The predicate inspects the **failure value**, because that is the only thing that says whether the condition can clear. The line is not a fixed list, but the shape of it is stable: | Likely to clear on a later attempt | Will fail identically every time | |---|---| | a timed-out call with no answer | a recipient the carrier rejects as invalid | | a refused or dropped connection | a message the carrier refuses as malformed | | a temporary refusal by the carrier | credentials the carrier refuses to accept | | a transient failure inside the carrier | a request the carrier says it will never accept | The right-hand column is the one that matters. Repeating those is pure waste, it delays the moment the caller learns the truth, and it disguises a defect as an outage. A predicate that cannot classify the failure at all is a signal that the failure is being carried in too coarse a form — a single opaque failure type for every cause leaves the decision nothing to read. ## Input two: the bound Even a genuinely transient failure needs a bound, because `transient` is a guess. Two shapes are common, and they answer different questions: - **A maximum attempt count** is simple and makes the worst-case work explicit: attempts multiplied by the cost of the retried scope. - **A deadline** bounds the thing the caller actually feels — elapsed time. With spacing between attempts, a count bound has a worst-case duration nobody computed; a deadline states it directly and stops mid-schedule when it expires. Use the count when the retried scope is expensive and you are protecting the dependency; use the deadline when something is waiting on the answer. The predicate and the bound are separate jobs: deciding **whether** this kind of failure is repeatable is not the same as deciding **how many** times or **how long**, and a decision that only counts will happily retry a permanent failure the full number of times. ## Input three: space between attempts Resubscribing the instant the failure arrives re-runs the same call against conditions that have had no time to change, and does it at whatever rate the failure returns — which, for a fast failure like a refused connection, is a tight loop. Separating attempts in time gives the condition a chance to clear and stops the retry adding load exactly when the dependency is least able to take it. How that spacing is computed — whether it grows, how it is randomised across many independent clients, whether attempts draw on a shared budget — is a resilience-pattern subject in its own right, with its own trade-offs; the requirement here is only that the gap exists and is not zero. ## What happens when the decision says stop The retry stage stops repeating and the sequence fails. Three things matter about that moment: 1. **The original cause travels.** Replacing it with a generic exhausted-attempts signal destroys the only evidence of what actually went wrong; wrap it if you like, but carry it. 2. **Attach the context the retry knows.** How many attempts ran and how long they took are facts only the retry stage has, and they are what makes the failure diagnosable later. 3. **Do not turn exhaustion into silence.** Ending the sequence as if it had simply completed makes a failed send indistinguishable from a successful one to everything downstream. ## The ambiguity the decision cannot resolve One case is worth stating plainly because the predicate looks like it handles it and does not: a timeout says only that no answer arrived. It does not say whether the carrier accepted the request. Classifying timeouts as retryable is usually right, and it means that some retried sends will be duplicates in the world. The predicate cannot fix that. What fixes it is what the retried scope contains — whether the repeated stage is one it is safe to repeat.
- What should the failure that reaches the subscriber after the last attempt carry?The original cause, intact. Add what only the retry stage knows — how many attempts ran and over what elapsed time — as context around it. A generic exhausted signal with the cause discarded leaves nobody able to tell a bad address from a dependency outage.
- A predicate retries only timeouts. What still breaks if the carrier times out on a request it actually accepted?The retry sends a duplicate. A timeout is ambiguous about whether the effect happened, so classifying it as retryable is a decision to accept duplicates. The predicate cannot resolve that; only the contents of the retried scope can — whether the stage being repeated is safe to repeat.
- Why prefer a deadline over an attempt count when a caller is waiting on the result?A count bounds attempts, not time, and once attempts are spaced apart the worst-case duration is an emergent number nobody chose. A deadline bounds exactly what the caller feels and stops partway through the schedule when it expires.
saying these in an interview costs you the question
- Retries every failure kind, including ones that can never succeed.
- Treats a large attempt limit as handling permanent failures.
- Resubscribes immediately, with nothing changed between attempts.
- Decides from the attempt number alone, never from the failure.
- Swallows the failure once attempts run out, so nothing downstream learns.
- Replaces the original cause with a bare exhausted-attempts signal.