skip to content

Why does Go's deadlock detector stay silent when a production service hangs with every worker blocked?

level: seniorimportance: should knowfreq 40%

answer

  1. it is a head count, not an analysis
  2. servers always have something runnable
  3. the listener is parked on the network
  4. the dump beats the detector
  5. group by wait reason and blocked duration

basics

~20 s

The runtime aborts only when nothing at all could run again. A server always has a goroutine parked on network I/O, in a syscall or on a timer, so the terminal condition never holds and the process hangs quietly. Read the goroutine stacks instead.

solid answer

~60 s

Go's deadlock report is a scheduler bookkeeping check, not a deadlock analysis: it fires only when no goroutine is running, runnable, blocked in a syscall, or waiting on a pending timer. A real service always fails that test — the listener is parked on the network poller, some goroutine is in a syscall, a ticker is pending — so even a total application-level deadlock never prints anything. The process stays alive, holds its port, and stops making progress. The diagnostic is the goroutine dump. In production I hit the `net/http/pprof` endpoint `/debug/pprof/goroutine?debug=2`, which prints every goroutine's stack with its wait reason and how long it has been blocked, without killing anything; `runtime.Stack(buf, true)` does the same in-process, and sending SIGQUIT prints the same dump but terminates the process. I look for a large group of goroutines parked on the same line with a growing blocked duration — for example thousands in `[chan send]` because their consumer died. The fix is structural: deadlines and cancellation on every blocking path, bounded queues, and a goroutine-count metric that alerts before the hang.

code

text · 11 lines
text
goroutine 1 [chan receive, 47 minutes]:
main.importAll(...)
	/app/import.go:42

goroutine 118 [chan send, 47 minutes]:
main.produce(...)
	/app/import.go:71

goroutine 9 [IO wait]:
net/http.(*conn).serve(...)
	/usr/local/go/src/net/http/server.go:2092

go deeper

for a junior

Know that Go's deadlock message only appears when the whole program is asleep, which almost never happens in a server. A hung service usually prints nothing at all.

for a middle

Explain the exact reason: a goroutine parked on network I/O, in a syscall, or on a pending timer keeps the runtime from declaring the program dead. Know that a goroutine dump lists every stack with its wait reason.

for a senior

Demonstrate the investigation: capture stacks without killing the process, group by wait reason and blocked duration, take two snapshots, and name the missing receiver. Then describe the structural fix, not just the restart.

for a principal

Own the posture: treat the detector as absent in production, require deadlines and cancellation on blocking paths, mandate goroutine-count alerting and automatic dump capture, and decide what the health check must actually exercise.

## Why the message never appears in production The runtime prints `fatal error: all goroutines are asleep - deadlock!` only when its scheduler concludes that nothing could ever run again: no goroutine running, none runnable, none blocked in a system call, and no pending timer that would wake one. That is a very strong condition, and a server violates it constantly: - the HTTP listener is parked in the network poller waiting for a connection, which counts as waiting on the outside world; - background goroutines hold timers for tickers, retries and cache expiry; - something is almost always in a syscall — reading a file, writing a log line, waiting on a socket. So a service can be **completely** deadlocked at the application level — every worker blocked on a channel nobody will ever service — and the runtime will say nothing. The process keeps its port open, health checks that only test the listener keep passing, and requests pile up. Silence from the detector is not evidence of health. ## What to look at instead The artefact you want is the **goroutine dump**: a stack trace for every live goroutine, each headed by a line naming its number, its wait reason, and how long it has been blocked. Three ways to get one: 1. **`/debug/pprof/goroutine?debug=2`**, exposed by importing `net/http/pprof`. This is the production tool: it is non-destructive, gives full stacks, and can be captured repeatedly. 2. **`runtime.Stack(buf, true)`** inside the process, useful for a watchdog that logs all stacks when a health check fails. 3. **SIGQUIT** (Ctrl-\ from a terminal), which prints the same dump — and then kills the process. Excellent locally, blunt in production, because you trade the outage's evidence for the outage's end. ## How to read it Do not read it top to bottom. Group it: - **Count by wait reason.** A few thousand goroutines all in `[chan send]` or `[chan receive]` is the shape of a hang; a handful in `[IO wait]` is normal. - **Count by the deepest application frame.** The stuck goroutines will share a line of your code. That line is the bug, or one hop from it. - **Read the durations.** `[chan receive, 47 minutes]` says this has been stuck since long before the symptom was reported, which usually points at start-up ordering or a consumer that died early rather than at the current traffic. - **Look for what is missing.** If thousands of producers are blocked sending, the interesting question is which goroutine used to receive and why it is gone — often a worker that returned on an error path, or a `range` over a channel whose loop exited. Taking **two** dumps a few minutes apart is worth far more than one: the groups whose stacks and durations advance are working, and the ones frozen at the same frame with growing durations are your deadlock. ## The fixes that actually hold - Every blocking operation gets a way out: a `context.Context` with a deadline, a `select` with a cancellation case, or a bounded queue that sheds load instead of parking producers forever. - Exactly one side owns closing a channel, and every error path still either delivers a value or cancels the context the other side is waiting on. - Instrument goroutine count and alert on sustained growth. A hang almost always shows as an unbounded goroutine count long before it shows as a full outage, and `runtime.NumGoroutine()` costs nothing to export. ## What the runtime gives you and what it does not The deadlock detector is a convenience for a program small enough that the whole thing can fall asleep — a test, a script, a one-shot importer. It is not an operational safety net, and the correct posture is to assume it will never fire in a service and to build the observability that will.

  • A dump shows 4,000 goroutines blocked in `[chan send]` on the same line. What is your reading?
    Producers are running and the consumer is not. Either the receiving goroutine returned early on an error path, or the `range` over the channel exited, or it was never started on this code path. The count also says the channel is an unbounded backlog in disguise: whatever the fix, producers need a bounded queue or a cancellation case so they shed load instead of accumulating.
  • Why is SIGQUIT a poor first choice on a hung production instance?
    It prints all goroutine stacks and then terminates the process, so you get exactly one snapshot and you destroy the state you were investigating. The pprof goroutine endpoint gives the same stacks without killing anything and lets you take a second snapshot minutes later to see which groups are actually frozen. Keep SIGQUIT for a local reproduction or a last resort.
  • How would you catch these hangs before they become an outage?
    Export `runtime.NumGoroutine()` and alert on sustained growth — a leak shows there long before requests fail. Add a liveness check that exercises the real work path rather than just the listener, keep a deadline on every blocking operation so a stuck path fails loudly, and capture a goroutine dump automatically when the check fails so the evidence exists without anyone logging in.

saying these in an interview costs you the question

  • Assumes no deadlock message means no deadlock
  • Reaches for SIGQUIT first on a live production instance
  • Reads the dump linearly instead of grouping goroutines
  • Restarts and moves on without capturing stacks
  • Expects the runtime to detect a partial deadlock