skip to content

A handler detaches after-response span work with context.WithoutCancel(r.Context()); the goroutine profile shows those goroutines climbing all day. What went wrong?

level: seniorimportance: nice to knowfreq 28%

answer

  1. detached from what, exactly?
  2. Done never fires on that context
  3. nothing is left to stop the work
  4. one goroutine per request, none ever exiting
  5. give it a deadline, then cap the count

basics

~20 s

context.WithoutCancel keeps the request's values but strips its cancellation and deadline, so the detached work has no stop condition left and any stuck call parks a goroutine forever. Give the detached context its own timeout and bound how many can run.

solid answer

~50 s

`context.WithoutCancel` is doing exactly what it says: the returned context still delegates `Value` to its parent, so the span survives, but `Done()` returns nil, `Deadline()` reports no deadline and `Err()` stays nil. The request context's cancellation was the only thing that would ever have stopped that work, and you removed it without replacing it. Now any call that blocks — a slow collector, a full queue, a lock — parks its goroutine permanently, and at a steady request rate the count grows linearly, which is what the goroutine profile is showing: thousands of goroutines parked at one line. The fix is two-part: wrap the detached context in `context.WithTimeout` with `defer cancel()` inside the goroutine, and stop spawning one goroutine per request — hand the work to a bounded set of background workers that drops and counts when saturated.

code

go · 13 lines
go
func handle(w http.ResponseWriter, r *http.Request) {
	w.WriteHeader(http.StatusAccepted)

	// keep the request's values (the span) but drop its cancellation...
	ctx := context.WithoutCancel(r.Context())
	// ...then give the detached work a stop condition of its own
	ctx, cancel := context.WithTimeout(ctx, 5*time.Second)

	go func() {
		defer cancel() // deferred INSIDE the goroutine, not in the handler
		finishSpan(ctx)
	}()
}

go deeper

for a junior

Know that a goroutine only ends when its function returns, and that a context which is never cancelled gives a blocked goroutine no way out. A count that climbs and never falls means work that never finishes.

for a middle

Explain the mechanics: WithoutCancel keeps Value delegation but returns a nil Done channel and no deadline, so any select on Done blocks forever and every blocking call in that path is unbounded.

for a senior

Demonstrate the diagnosis and the layered fix — read the goroutine profile, correlate the ramp with request rate, add WithTimeout with cancel deferred inside the goroutine, then replace goroutine-per-request with bounded workers that drop and count.

for a principal

Own the policy: telemetry must never apply backpressure to the serving path, so decide the overflow behaviour, the concurrency ceiling and the shutdown drain up front, and make them the default in the shared helper rather than each team's discovery.

## What WithoutCancel actually returns `context.WithoutCancel(parent)` (Go 1.21+) returns a context that is derived from `parent` for values and detached from it for everything else: - `Value(key)` delegates up the parent chain, so the request's span and ids are still reachable — that is why it was reached for here; - `Done()` returns **nil**. A receive on a nil channel blocks forever, so a `select` case on `ctx.Done()` in that work can never fire; - `Err()` returns nil, always; - `Deadline()` returns the zero time and `false`. It is deliberately, permanently un-cancellable. Used correctly that is the point: work that must outlive the response — finishing and shipping the request's span is the textbook case — cannot use `r.Context()`, because the server cancels that context when the handler returns and when the client disconnects. `WithoutCancel` keeps the trace identity while shedding a cancellation that would fire at exactly the wrong moment. ## Why goroutines then climb The mistake is treating "detached" as the end of the design. The request context was carrying two things: values and a stop condition. You kept the values and deleted the stop condition, and put nothing in its place. Every blocking operation in the detached path is now unbounded: - a network call whose client has no timeout of its own; - a send on a channel whose consumer has fallen behind or died; - a mutex held by something that is itself stuck; - a retry loop whose exit condition was `ctx.Err() != nil`, which is now permanently nil. One goroutine per request, each with no way to end, at a steady arrival rate, is linear growth for as long as the process runs. It is not a burst; it is a ramp, which is exactly the shape the goroutine profile shows. The memory cost is larger than the stacks. Each parked goroutine keeps alive everything it references: the span, the whole context chain behind it (which `WithoutCancel` keeps pinned by design), any buffers captured by the closure, and often the request payload. So the heap climbs alongside the goroutine count, and the process eventually dies of memory rather than of goroutines. ## Diagnosing it - The **goroutine profile** is the direct evidence. The count alone tells you it is a ramp; `debug=2` dumps every goroutine's stack, and a leak of this kind shows thousands of identical stacks parked on the same line, with wait durations in the hours. - `runtime.NumGoroutine()` sampled over time gives you the same ramp as a plottable number and is cheap enough to keep permanently. - Correlating the ramp's slope with request rate confirms "one per request": if the count rises by roughly the number of requests served, the work never exits. - Restarting the process "fixing" it for a few hours is another signature of an accumulation rather than a burst. ## The fix, in order of importance **1. Give the detached context a stop condition of its own.** Immediately after `WithoutCancel`, derive a timeout sized for the work — a few seconds is generous for finishing a span — and `defer cancel()`. Two details matter: the `cancel` must be deferred *inside* the goroutine, since deferring it in the handler cancels the work as soon as the handler returns, which cancels the very thing you detached; and the calls in the detached path must actually observe the context, otherwise a deadline is decoration. **2. Bound the concurrency.** Even with a timeout, `go` per request means the ceiling is "arrival rate multiplied by timeout", and under an incident that number is large. A fixed number of long-lived background workers fed by a buffered channel, or a counting semaphore around the spawn, turns an unbounded leak into a bounded queue. **3. Decide the overflow policy explicitly, and prefer dropping.** This is telemetry. When the queue is full the right behaviour is almost always to drop the work and increment a counter, not to block the request path or to grow without limit. A `select` with a `default` on the send gives you that in three lines. Losing some spans during an incident is cheaper than turning a telemetry backlog into an outage of the service being traced. **4. Handle shutdown.** Detached work is invisible to the server's graceful-shutdown path, so on `SIGTERM` you either drain the workers with a bounded wait or accept that in-flight spans are lost — but make it a decision rather than an accident. ## What a strong answer sounds like "`WithoutCancel` keeps the span but removes the only thing that ever ended that work — `Done()` is nil, so nothing in the path can time out. Add a `WithTimeout` with `defer cancel()` inside the goroutine, then replace goroutine-per-request with a bounded worker set that drops and counts when it is full. And `Background` would not have helped: it is equally un-cancellable and loses the span as well."

  • What exactly do Done, Err and Deadline return on a context from context.WithoutCancel?
    `Done()` returns nil — and a receive on a nil channel blocks forever, so a select case on it never fires. `Err()` returns nil always, and `Deadline()` returns the zero time with ok false. Only `Value` still delegates to the parent. That is the whole design: values kept, cancellation removed.
  • Would using context.Background() for the detached work have avoided the leak?
    No. `Background` is equally never cancelled, so the goroutines would pile up identically, and you would also lose the span, so the work would be unattributed as well as unbounded. The leak comes from having no stop condition, not from which root was chosen.
  • Where should defer cancel() go for the detached context, and what breaks if it is misplaced?
    Inside the goroutine that does the work. If you defer it in the handler, cancel runs as soon as the handler returns, which cancels the detached work immediately — you get no leak and no telemetry either. If you forget it entirely, the timer and the context stay alive until the deadline expires, which is a smaller but real accumulation.
  • How do you keep the fix from turning into a different outage?
    Bound it. A fixed worker set fed by a buffered channel, with a non-blocking send and a dropped-work counter, means a slow telemetry path degrades to lost spans instead of growing memory or, worse, blocking request handlers. Never let the tracing path apply backpressure to the serving path.

You cut the wire that would have pulled the plug at closing time so the cleaners could finish after the shop shut, and then never gave the cleaners a shift end. They are all still in there.

saying these in an interview costs you the question

  • Thinking WithoutCancel supplies some default timeout
  • Expecting the goroutine to exit when the handler returns
  • Selecting on ctx.Done() of a WithoutCancel context and expecting it to fire
  • Deferring cancel in the handler instead of inside the goroutine
  • Adding more goroutines because the background work is falling behind
  • Blaming the garbage collector for a rising goroutine count