A handler responds 200 and then runs the real work on a thread it spawned itself — what breaks in production?
answer
- nothing records the work is owed
- the next deploy is the outage
- no backpressure, no retry, no dead-letter
- loss is silent, not loud
- durability from the record, not the runner
basics
~20 sNothing durable records that the work is owed, so a deploy, crash or scale-in silently loses whatever is in flight or queued in memory. There is also no backpressure, no retry, no dead-letter, and no caller left to report a failure to.
solid answer
~50 sThe pattern works right up to the first restart. Because the only record that the work exists is an in-memory task, replacing the process — a routine deploy, a scale-in, an out-of-memory kill — discards everything not yet finished, and nothing anywhere knows it was owed. The other three gaps follow from the same root: there is no backpressure, so a burst can create work faster than the process can drain it while competing with request handling for the same threads and connections; there is no retry or dead-letter, so a transient dependency failure is permanent loss; and there is no caller, so failures surface only if you counted them. The fix is not a bigger pool. It is to **record the obligation durably before responding**, and let the record — not the thread — be what makes the work recoverable.
go deeper
Take away the core fact: work started on your own thread lives only as long as the process. If the process restarts, that work is gone and nothing remembers it existed.
Explain the four gaps — durability, backpressure, retry, visibility — and why a bigger thread pool addresses none of them.
Demonstrate the diagnosis: losses that correlate with deploy times, no reconciliation list, background work competing with serving capacity — and drive the fix to a durable record written before the response.
Own the policy. Decide which classes of work may be lost at all, make the durable path the default one teams reach for, and treat any silently lossy endpoint as a design defect rather than a known quirk.
## Why it looks correct Answering the caller quickly and doing the slow part afterwards is the right instinct. The response is not waiting on an email, a thumbnail or a downstream write, so the client sees low latency and the endpoint looks healthy in every dashboard. In a demo and in most tests, the work also happens: the process stays up for the whole scenario, the task runs, the effect appears. This pattern usually survives months of production traffic before anyone notices what it costs. ## What is actually missing 1. **Durability.** The only evidence the work is owed is an object in memory. When the process is replaced — and in a normally operated system it is replaced constantly, for deploys, scaling and host maintenance — everything unfinished disappears. There is no list of what was lost, so nobody can reconcile it afterwards. 2. **Backpressure.** Spawning per request lets arrival rate set concurrency. A traffic spike creates tasks as fast as requests arrive, and those tasks contend with request handling for threads, memory and dependency connections. The failure looks like a serving problem, so scaling on request metrics adds capacity for the wrong cause. 3. **Retry and dead-lettering.** A dependency that is briefly unavailable turns a recoverable fault into permanent loss. There is nothing to hold the item while the dependency recovers and nowhere to park an item that keeps failing. 4. **Visibility.** The caller is gone, so a failure produces no error response, no client retry and no support ticket. Unless counters were added deliberately, the loss rate is unknown — and "unknown" is usually reported as zero. ## In-process task versus recorded handoff | Property | Self-spawned in-process task | Obligation recorded durably first | |---|---|---| | Survives a restart | No — lost with the process | Yes — the record outlives the process | | Knows what was lost | Nothing to reconcile against | The unfinished records are the list | | Behaviour under a burst | Unbounded growth, competes with serving | Backlog grows where it can be measured | | Transient dependency failure | Permanent loss | Retried, then dead-lettered | | Failure visibility | Only what you counted | Backlog age, failure counts, dead-letter count | | Cost to operate | Nothing extra to run | A consumer or sweeper, and its alerts | ## If it genuinely must stay in-process Sometimes a durable handoff is not worth it — the work is regenerable, best-effort, or trivially repeated by the next request. Then make the losses bounded and known: - Submit to a **fixed-size pool with a rejection policy**, never an unbounded spawn, so a burst produces countable rejections instead of resource exhaustion. - Give every task a **timeout**, so one stuck call cannot hold a worker forever. - Make the effect **idempotent**, so a client retry or a later repair pass cannot double it. - Export counters for **submitted, completed, rejected and failed**, so "we lose a few" is a measured number and not a hope. - Write down, in the code and in the runbook, **which work may be lost**, so the next person does not quietly add something that may not. ## The smallest change that fixes it properly Record the intent as part of the write the request was already making: a row, a document, an entry with a state field, committed before the response goes out. Once the obligation is durable, who performs it becomes an implementation choice — a consumer, a periodic sweep for unfinished records, or even the same process, which may now crash without losing anything. That is the real insight: **durability comes from the record, not from the runner.** A background thread with a durable record behind it is recoverable; a durable-looking queue client called from a thread that nothing recorded is not. ## How it shows up in an incident The usual report is not an outage. It is "some customers did not get their confirmation last Tuesday", and the counts never quite reconcile. Correlating the gaps with deploy timestamps is the diagnosis, because the loss window is exactly the moment the process was replaced. By then there is no record of the missing items, so the recovery is a manual reconstruction from whatever other trace the request left — which is the strongest argument for making the obligation durable at the moment it is accepted.
- What is the smallest change that makes this work survive a restart?Record the intent durably as part of the write the request already performs — a record with a state field — before responding. Then anything can perform it: a consumer, a sweep over unfinished records, even the same process. The durable record, not the thread, is what makes the work recoverable.
- The work really is allowed to be lost. What do you still owe it?Bounds and visibility. A fixed-size pool with a rejection policy so a burst cannot starve request handling, a timeout so one task cannot pin a worker, and counters for submitted, completed, rejected and failed — so the statement that some work is lost stays a measured number.
- Why does spawning per request behave differently from submitting to a bounded pool?An unbounded spawn lets arrival rate decide concurrency: a burst creates as many workers as there are requests, all competing with request handling for the same resources. A bounded pool turns overload into a visible queue or a countable rejection that can be alerted on.
Handing work to a thread you spawned is like writing a reminder on a whiteboard in a room that gets repainted on every deploy. It works fine until the painters come, and afterwards nobody can even say what was written.
saying these in an interview costs you the question
- Says it is fine because deploys are quick and infrequent.
- Assumes unbounded task spawning is harmless because tasks are short.
- Treats a 200 response as evidence the work actually happened.
- Thinks waiting at shutdown solves the durability problem.
- Believes a crash loses only the single task in flight.