skip to content

Once the response is sent, a failure in the deferred work has no caller to return to — how do you design for that?

level: seniorimportance: should knowfreq 54%

answer

  1. no caller left to return to
  2. the response set the promise
  3. retry, then dead-letter
  4. alert on backlog age
  5. carry the correlation identifier forward

basics

~20 s

Replace the missing caller with three channels: automatic recovery (bounded retry, then a dead-letter place), operator visibility (failure counters, backlog age, a correlation identifier), and a user-facing status or follow-up. And never claim in the response what the deferred work has not yet done.

solid answer

~40 s

A synchronous failure reports itself: the caller gets a non-2xx status and reacts. Once work is deferred, that feedback path is gone, so it has to be rebuilt deliberately in three directions. **Recovery**: bounded retries with backoff for transient faults, then a dead-letter place with the cause recorded, because a permanently failing item must stop consuming capacity. **Operators**: counters for accepted, completed and failed items, plus the **age of the oldest unprocessed item** — a stalled consumer produces no errors at all, so an error-rate alert never fires. **Users**: a status they can check or a follow-up message, since they have no other way to learn. Underpinning all three, the response must be honest: if it says the action is done, a later failure has already become a lie.

go deeper

for a junior

Grasp the asymmetry: a failure during the request becomes an error status the caller sees, while a failure after the response reaches nobody unless something was built to report it.

for a middle

Explain the mechanics of recovery — bounded retry with backoff, dead-lettering a permanently failing item, and why replay demands idempotent effects.

for a senior

Show the operational instinct: alert on the age of the oldest unprocessed item, not only on errors, and be able to say why a stalled consumer is invisible to error-rate alerting.

for a principal

Treat the wording of the response as an architectural commitment. What an endpoint claims decides what must be durable before it answers, and that rule scales across teams better than any tooling choice.

## Why a late failure is silent While a handler is running, failure has somewhere to go. It becomes a status code, the caller sees it, a client may retry, a user may complain, and the endpoint's error rate moves. Every one of those signals comes from the fact that somebody is waiting. Move the work behind the response and all of them vanish at once. There is no connection to write an error to, no client that will retry, no user who knows to complain, and the accepting endpoint keeps reporting success — it did succeed, at the only thing it promised. Failure after the response is therefore silent **by construction**, and the whole design problem is to put deliberate channels where the caller used to be. ## The response defines what a late failure means | What the response said | What a later failure means | What the design owes the user | |---|---|---| | "Done" — the action is complete | The response was untrue and the user acted on it | Do not respond this way unless the effect is already durable | | "Accepted" — we will do this | An expected, recoverable case | A way to see the outcome, or a follow-up message | | Nothing about the work at all | Purely internal; the user is unaffected | Operator visibility only | This table is the first design decision, not a detail. Much of the pain of deferred work comes from a response that claimed completion for something that had merely been scheduled. ## Three replacements for the caller 1. **The system tells itself.** Retry transient faults with backoff and a cap, then move the item to a dead-letter place with its failure cause attached. Because retries replay work, the effect must be idempotent — a natural key, a conditional write, or a state check — or recovery becomes duplication. 2. **The system tells an operator.** Counters for items accepted, completed and failed; the size of the dead-letter place; and the **age of the oldest unprocessed item**. That last one is the signal most teams are missing. 3. **The system tells the user.** A status they can check, a notification when the work completes or fails permanently, or an inbox message. If the user can only find out by noticing the absence of something, the design is incomplete. ## What to alert on - **Backlog age**, not backlog size. Size depends on throughput; age directly measures how stale the oldest obligation is. - **Dead-letter arrivals**, because each one is an obligation that nothing else will fulfil. - **Failure rate per stage**, not just a global rate, so a single failing downstream is distinguishable from a broad fault. - **Completion rate against acceptance rate.** If acceptances exceed completions for long, the gap is real work quietly going missing. The classic gap is alerting only on errors. A consumer that has crashed, deadlocked, or is stuck retrying one poison item emits **no errors at all**; the error rate is flat, every dashboard looks healthy, and work simply stops. Only age-based signals catch that. ## Carry the thread of identity forward The request that accepted the work has a correlation identifier. Copy it into the durable record and into every log line and metric the deferred work emits. Without it, an incident that starts as "this customer's order never completed" cannot be traced past the accepting request, because nothing downstream carries anything that connects to it. With it, one identifier follows the obligation from acceptance to completion or to the dead-letter place — which is what makes reconciliation possible at all. ## Reconciliation closes the loop Even with retries and alerts, some items end up in a state nobody watched. A periodic pass that looks for records accepted long ago and still unfinished is cheap to run and catches everything the live path missed: items lost before they reached a consumer, items whose consumer died mid-flight, items stuck in a state transition. Treat that sweep as part of the design rather than as an incident tool — it is the only component that answers the question "is anything we accepted still undone?" without being told where to look. ## The summary Deferring work moves failure out of the request's blast radius and into a place with no natural reporter. Design the reporter: retry then dead-letter for the machine, age and counters for the operator, a status or a message for the user, a correlation identifier tying them together, and a reconciliation sweep that assumes all of it can still miss something.

  • Which signal catches a stalled consumer that an error-rate alert misses?
    The age of the oldest unprocessed item. A consumer that crashed, deadlocked, or is stuck on one poison item emits no new errors, so the error rate stays flat while progress stops. Backlog age begins rising immediately and keeps rising, which makes it the signal to alert on.
  • How do you decide between retrying and dead-lettering an item?
    By whether a retry could plausibly succeed. Transient faults — a timeout, a briefly unavailable dependency — deserve bounded retries with backoff. A malformed or permanently rejected item will fail identically forever, so it belongs in a dead-letter place with its cause recorded, where a person can decide what to do.
  • What must the response itself avoid claiming?
    Anything the system has not yet made durable. Saying the action is complete when it is merely scheduled turns every late failure into a falsehood the user already acted on. Saying it is accepted, and giving them a way to see the outcome, keeps the promise honest and the recovery path legitimate.

saying these in an interview costs you the question

  • Relies on someone reading error logs to notice failures.
  • Alerts on failure rate but never on backlog age.
  • Retries forever with no cap and no dead-letter place.
  • Treats an accepted response as proof the work succeeded.
  • Tells the user an action is complete before it is durable.
  • Assumes retries are safe without making the effect idempotent.