skip to content

How do you decide which post-response work in a Go service may outlive the request context?

level: principalimportance: nice to knowfreq 32%

answer

  1. start from who notices if it is lost
  2. latency relief is not durability
  3. deploys, shutdown, backpressure
  4. detached goroutines are not a queue
  5. the data owner holds the veto

basics

~20 s

Split the work by whether losing it is acceptable. Best-effort work may run on a context detached from the request and bounded; anything reconciled later must be written inside the request, because a detached goroutine dies with the process.

solid answer

~60 s

The decision is about durability, not latency. I classify post-response work into two buckets. Best-effort — cache warms, opportunistic refreshes, telemetry export — may run on a context detached from the request with `context.WithoutCancel`, wrapped in its own timeout and submitted to a bounded worker rather than a fresh goroutine per request. Must-not-be-lost — audit, billing, anything a person or an auditor reconciles — does not belong off the request path at all: it goes in the same transaction as the business change, or into durable storage that a separate consumer drains. The reason is that detaching only buys you freedom from cancellation; it buys nothing against a deploy, a crash or a shutdown that waits for handlers and not for your goroutines. So the policy I set is: nothing whose absence would be a defect leaves the request path, with a documented exception process, a cap on in-flight detached work, and counters for attempted versus completed so the claim stays checkable — because the person who overrules me is the one whose rows went missing.

go deeper

for a junior

You are not expected to set this policy, but know the underlying fact: a goroutine started by a handler dies with the process, so work that must not be lost should be done before the handler returns.

for a middle

Be able to describe the mechanism you would use for best-effort work — a detached context with its own timeout, submitted to a bounded worker — and say plainly that it is best-effort rather than reliable.

for a senior

Argue the tradeoff with numbers: what keeping the write in the request costs in tail latency, what deploy frequency implies about pending work lost, and which categories you would refuse to move regardless.

for a principal

Own the classification, the default, the cap and the shutdown semantics, and be clear about who can overrule you and why. Show how the decision is enforced by the shared helper rather than by good intentions.

## Why this is a judgment call rather than an API question The mechanism is easy: `context.WithoutCancel` detaches work from the request's cancellation, and a timeout re-bounds it. The hard part is that detaching is *cheap to write and expensive to be wrong about*, and the cost lands on somebody who is not in the room — the data engineer reconciling audit rows, the finance team matching charges, the compliance owner who has to attest that every access was recorded. A service owner who moves a write off the request path is trading a few milliseconds of tail latency for a small, invisible loss rate. That is a legitimate trade for some work and an unacceptable one for other work, and the line is not technical. ## The classification I use **Best-effort.** Losing an occasional instance changes nothing anyone will ever notice: warming a cache, refreshing a derived value that will be recomputed anyway, emitting telemetry, invalidating something that also has a TTL. These may be detached. They should still be bounded and instrumented, but nobody is harmed by a gap. **Required-but-deferrable.** Somebody will eventually notice: notification emails, webhooks to customers, search index updates. These may leave the request path only if there is a durable handoff — a row committed in the same transaction as the business change, or a message on a broker — with a consumer that retries. A goroutine is not a durable handoff. **Required-and-reconciled.** Audit trails, billing events, anything an auditor samples. These are written inside the request, in the same transaction as the change they describe wherever the storage allows it. If that costs latency, the honest answer is that the latency is the price of the guarantee, and the conversation to have is about making the write cheaper, not about making it optional. ## What detaching actually does and does not buy It is worth being precise, because the fix is often mistaken for a guarantee. A detached context is never cancelled and has no deadline; it still carries the request's values. So the work survives the handler returning and the client hanging up. It does not survive: - **process exit** — deploys are the common case, and a service that deploys ten times a day loses a slice of pending work ten times a day; - **graceful shutdown**, unless you made it wait: the HTTP server's `Shutdown` drains in-flight *handlers*, and knows nothing about goroutines a handler spawned; - **backpressure** — one goroutine per request is unbounded by construction, so a slow dependency converts a traffic spike into memory growth and an ever-longer queue of work that will be lost at the next restart. Any policy that says "we detach it" without answering those three has not been finished. ## The controls that make the policy real **A default, stated as a default.** "Nothing whose absence would be a defect leaves the request path." Defaults that live in a document nobody reads are decoration; this one belongs in the shared helper — if the only easy way to schedule background work is a submission function that takes a bounded worker pool and applies the house timeout, the default enforces itself and `go f(r.Context())` stops appearing in review. **A cap and a shed.** A fixed number of workers and a bounded queue, with a defined behaviour when it is full: block the handler, drop with a counter, or spill to durable storage. Choosing is the point; discovering it under load is not. **Shutdown semantics.** Say explicitly what a rolling deploy waits for. Draining a bounded queue for a few seconds during shutdown is cheap and removes the largest loss source. **Evidence.** Counters for submitted, completed and dropped, per category, plus the error from every failed attempt. Without them, "best-effort" is unfalsifiable and the first anyone hears of a problem is a gap discovered downstream weeks later. ## Who can overrule you, and on what grounds This is the part to say out loud in an interview. The service owner sets the default because they own the latency and the capacity. But the owner of the *data* — audit, billing, compliance — can overrule any classification for their category, because they carry the consequence of a gap, and their requirement is usually not negotiable in the way a latency target is. The workable arrangement is: engineering proposes the classification and provides the numbers (what it costs in tail latency to keep it in the request), the data owner accepts or rejects, and the decision is written down next to the code so the next person does not silently re-litigate it by adding a `go` statement. ## The shape of a good answer Start from the loss model rather than the API. Distinguish latency relief from durability. Name the three things detaching does not survive. Give the controls — default, cap, shutdown, counters — and be explicit that the person who eats the cost of being wrong gets a veto.

  • A team says the audit write is too slow to keep in the request. What do you propose?
    Measure it first — it is often a synchronous round trip that batching or a local durable buffer fixes. If it truly is expensive, keep durability and move the slow part: commit a compact record in the same transaction as the business change and let a consumer do the enrichment. That preserves the guarantee while taking the cost off the response.
  • How do you stop this decision from being re-made silently by whoever writes the next handler?
    Make the sanctioned path the easiest one: a single submission helper backed by a bounded pool that applies the house timeout and the detached context, so hand-rolled goroutines stand out in review. Record the classification for each category of work next to the code, so the next person amends a decision rather than inventing one.
  • What would make you accept losing some records during deploys?
    Only if the owner of that data accepts it explicitly, with a measured loss rate rather than a hope, and there is a way to detect and backfill a gap. For telemetry that is usually fine. For anything reconciled against money or access, the answer is no, and the fix is a durable handoff rather than a longer shutdown grace period.

saying these in an interview costs you the question

  • Treats context.WithoutCancel as a durability guarantee
  • Moves audit or billing writes off the request path for latency
  • Spawns unbounded goroutines and calls it a queue
  • Ignores what a rolling deploy does to pending work
  • Sets a policy with no counters to check it against
  • Decides the data owner's loss tolerance without asking them