How would you decide which post-response work may stay in the serving process and which needs a durable handoff?
answer
- cost of loss, not duration
- who repairs it if lost
- record durably, then choose the runner
- a handoff is a second failure domain
- background work shares the serving budget
basics
~20 sDecide by the cost of losing an item, not by how long it takes. Best-effort, regenerable work can stay in the process behind bounds; anything the user was told happened, or that a person would have to repair, must be recorded durably before you respond.
solid answer
~50 sThe axis people reach for first — duration — is the wrong one. A two-second charge matters far more than a two-minute cache rebuild. Decide on **loss cost**: if the item disappeared during a deploy, would anyone notice, and would a person have to repair it? That question sorts work into three tiers rather than two. Best-effort work stays in-process behind a bounded pool and a timeout. Work that must survive gets its obligation recorded durably during the request; who performs it afterwards is then a separate, reversible choice. Only work needing fan-out, ordering, isolation from the serving process, or independent scaling justifies a full queue system and the second failure domain it brings. Two constraints cut across all of it: background work shares the serving process's capacity, and anything recoverable will be replayed, so effects must be idempotent.
go deeper
Learn the sorting question rather than a rule of thumb: ask what would happen if this work simply never ran, and let the answer decide where it belongs.
Be able to name the tiers and their requirements — bounded pool and counters for best-effort work, a durable record before responding for anything that must survive.
Argue the middle tier convincingly: a record with a state column plus a sweep delivers recoverability without the operational cost of a broker, and say when that is no longer enough.
Set the policy and the default path. State it as a rule about what responses may claim, give teams one supported mechanism, and add reconciliation so a wrong choice still surfaces.
## The axis that actually decides it The question is usually posed as "is this too slow to do in the request?", which is the wrong axis for the wrong reason. Duration decides whether work belongs *before* the response; it says nothing about what happens if the work never runs. The deciding question is: > If this item vanished silently during a routine deploy, who would notice, and what would it take to repair? Answers range from "nobody, and nothing" to "the customer, and a manual reconciliation against a payment provider". That range, not a stopwatch, is what separates the options. | Axis | Pushes toward in-process | Pushes toward a durable handoff | |---|---|---| | Cost of silent loss | Negligible, regenerable | A person must repair it, or a user was told it happened | | What the response claimed | Nothing, or an internal effect | The response said the action happened | | Volume and burstiness | Small, smooth, bounded | Spiky; needs a buffer between arrival and processing | | Capacity coupling | Cheap work that will not disturb serving | Heavy work that would compete with request handling | | Retry semantics | Nothing to retry, next request repeats it | Needs backoff, caps and a dead-letter place | | Operational cost | Nothing extra to run | A consumer, its alerts, and a second failure domain | ## Three tiers, not two 1. **Best-effort, in-process.** Metrics, cache warming, an optional trace export. Requirements: a bounded pool with a rejection policy, a timeout per task, counters for submitted, completed, rejected and failed, and an explicit written note that losing these is acceptable. 2. **Durable record, simple runner.** The obligation is written in the same durable store the request already touches, before the response goes out. A poller or sweeper claims unfinished records and performs them, with attempts and last error kept as columns of the record. This gives recoverability without introducing anything new to operate — often the right answer for a service that already has a database and no message infrastructure. 3. **Durable handoff to a dedicated system.** Warranted when the work needs fan-out to several consumers, ordering guarantees, isolation so heavy processing cannot disturb serving, or independent scaling of the workers. It is the most capable option and the most expensive: a second failure domain, poison-message handling, backlog alerting, and one more thing to upgrade. The middle tier is the one teams skip. They oscillate between "just spawn a task" and "stand up a broker", when a record with a state column plus a periodic sweep provides the durability that actually mattered. ## What a handoff costs, honestly - **A second failure domain.** The broker's availability becomes your endpoint's availability if you enqueue synchronously inside the request. - **Duplicate delivery.** Most systems deliver at least once, so every effect must be idempotent regardless of tier. - **Two places where the truth lives.** A record written in one store and a message sent to another can diverge; either enqueue from the record, or accept reconciliation. - **New operational surface.** Backlog alerts, dead-letter review, consumer deploys, and an on-call story for the consumer as well as the service. These are worth paying when the loss cost justifies them, and pure overhead when it does not — which is why the decision is made per obligation, not once per service. ## Capacity is the constraint people forget Work kept in the serving process shares that process's threads, memory and dependency connections with request handling. A burst of deferred tasks raises request latency in a way that looks exactly like a serving problem, and autoscaling on request metrics will happily add instances that each start doing more background work. Any in-process tier therefore needs a hard bound, so the failure mode under load is countable rejection rather than an unpredictable slowdown of the whole service. ## Making it a rule other teams can apply Technology rules age badly; promise rules do not. A durable version is: **if a response tells the user something happened, the effect must be durable before that response is written.** It needs no knowledge of which handoff mechanism is currently favoured, it is checkable in review by reading the response body, and it puts the argument where it belongs — on what the endpoint claims. Pair it with one supported, documented path for tier 2 so that the easy choice is also the correct one, and a periodic reconciliation sweep so that unfinished obligations surface even when a tier was chosen wrongly.
- What does a durable record plus a sweeper buy over a full queue system?Recoverability without a new failure domain. The obligation lives in storage the team already runs and backs up, a poller claims unfinished records, and retry state is a column rather than a broker feature. It gives up fan-out, low-latency dispatch and ordering guarantees, which many workloads never needed.
- Where does keeping the work in-process hurt capacity?Background work and request handling share one process's threads, memory and dependency connections. A burst of deferred tasks raises request latency in a way that looks like a serving fault, and scaling on request metrics adds instances that each take on more background work as well.
- How would you keep this decision consistent across many teams?State it as a rule about promises, not technology: if a response tells the user something happened, the effect must be durable before that response is written. Then provide one supported handoff path, so the easy choice is also the correct one, and review new endpoints against that single question.
saying these in an interview costs you the question
- Decides purely on task duration, ignoring the cost of loss.
- Moves everything to a queue and ignores the operational cost.
- Keeps user-visible obligations in-process because a queue feels heavy.
- Assumes a durable handoff removes the need for idempotency.
- Forgets that background work competes with request handling for capacity.
- Never considers a durable record with a simple in-process runner.