Why does Go's "all goroutines are asleep - deadlock!" report never fire in a hung production service?
answer
- the check is all-or-nothing
- the scheduler must have nothing to run, ever
- one pending timer means future work
- a listener or a ticker keeps it quiet
- silence is not evidence
basics
~20 sThat report needs the whole program stuck: no runnable goroutine, no thread in a system call, no armed timer. One open listener, one resync ticker or one sleeping goroutine is enough to keep the runtime silent forever.
solid answer
~40 sThe runtime's check is all-or-nothing and it only runs when the last thread is about to go idle. A real service always has something that disqualifies it: a goroutine parked in the network poller on an open listener or connection, a thread inside a blocking system call, or — the one that catches people — a single armed timer from `time.Sleep`, a `time.Ticker`, `time.AfterFunc` or a server's own timeouts. A pending timer means future work exists, so the runtime cannot claim permanent deadlock. There is also no partial detection: three goroutines in a lock cycle while forty others serve traffic will never be reported, because those forty are runnable. So silence is not evidence of health — a hung service is confirmed by reading goroutine stacks, not by waiting for the runtime to speak.
code
go · 7 linesfunc (r *reconciler) run() {
tick := time.NewTicker(30 * time.Second)
defer tick.Stop()
for range tick.C {
r.reconcile() // may block forever on a lock cycle
}
}go deeper
Remember that the deadlock message belongs to small programs. A service that hangs quietly can still be deadlocked, and you should not report it as healthy because nothing was printed.
Be able to list what disqualifies the check — a runnable goroutine, a thread in a system call, a goroutine in the network poller, any armed timer — and explain why a pending timer makes the proof impossible.
Demonstrate the operational consequence: a partial deadlock is invisible to the runtime, so you confirm from goroutine stacks, two snapshots and a near-zero CPU line rather than waiting for a crash.
Own the conclusion that this failure mode never self-reports, so the platform must supply the evidence: a way to pull goroutine stacks from a live process and an alert on stuck work rather than on process death.
## The check is all-or-nothing The Go runtime reports `fatal error: all goroutines are asleep - deadlock!` only when it can prove the *entire process* can never make progress. That proof needs three things at once: no thread doing anything, every non-system goroutine parked, and no timer armed anywhere. Miss any one and the runtime returns quietly and the process hangs with no output at all. Production programs fail that test structurally, not accidentally. ## The four things that keep it silent **A goroutine in the network poller.** A server with an open listener has a goroutine parked in `Accept`, and every live connection has a goroutine parked in a read. The runtime cannot prove that no byte will ever arrive, so it will not conclude. This alone disqualifies essentially every networked service. **A thread inside a system call.** Reads on files, `os/exec`, DNS resolution on some platforms, and any cgo call occupy an OS thread that is not idle. The check bails out at its first stage when any thread is busy. **An armed timer.** This is the subtle one and the one worth being able to state precisely. If any P's timer heap holds a pending timer, a goroutine is scheduled to become runnable at a known instant, so permanent deadlock is impossible by definition. A `time.Ticker` driving a resync loop, a `time.Sleep` in a retry backoff, a `time.AfterFunc` guarding an operation, a server's read and write deadlines — any one of them keeps the runtime from ever concluding. A controller that wakes every thirty seconds to reconcile desired state is permanently disqualified from the check for as long as its ticker exists, even if every goroutine it owns is wedged in a lock cycle. **Anything still runnable.** A busy-spinning goroutine, a health-check loop, a metrics reporter — anything with work to do means the last thread never goes idle and the check never runs. ## There is no partial detection Even if you removed every timer and every socket, the runtime still would not help you with the common production shape: a *subset* of goroutines deadlocked while the rest keep working. Three goroutines in a mutex cycle inside a service that still answers its health endpoint is a real, user-visible outage and an invisible one to the runtime. The check answers exactly one question — can anything at all run? — and the answer there is yes. This matters for how you read symptoms. A partial deadlock looks like: latency on some endpoints, timeouts, a rising in-flight count, goroutine count climbing as new requests pile up behind the wedged ones, and CPU near zero. The process stays alive, the liveness probe may still pass, and nothing is ever printed. ## The postmortem sentence When you write up an hour-long hang, the question you will be asked is "why did it not crash?" The honest answer is that Go's deadlock detector is a whole-program proof, not a lock monitor, and a single thirty-second resync ticker in the reconcile loop was enough to make the proof unavailable — the scheduler always had a future event to point at. That is not a bug in the runtime; the runtime genuinely could not know that the timer's callback would immediately block on the same wedged mutex. ## What you do instead Because the runtime will not volunteer anything, confirmation is your job: - **Get every goroutine's stack** from the running process and read it. The evidence you need is which goroutines are parked, on what, and for how long — goroutine headers carry a wait reason and, past a minute, a wait age. - **Take two snapshots a minute apart.** Identical goroutine numbers, identical frames and wait ages that advanced by exactly the elapsed time mean nothing moved. Different frames mean slow, not stuck — a completely different investigation. - **Check CPU.** A deadlock is near-zero CPU. A livelock or a hot retry loop burns CPU and is not this. - **Reproduce without the suppressors.** In a minimal reproduction, delete every `time.Sleep`, every ticker and every background server, so the runtime *can* reach its verdict and confirm the cycle for you. ## The rule to carry Absence of the message is not absence of a deadlock. The report is a strong positive signal and a worthless negative one — treat it as "if you see it, believe it; if you do not, you have learned nothing."
- Does a goroutine blocked on a socket read count as asleep for that check?No. It is parked in the network poller, and the runtime cannot prove that no data will ever arrive — a peer could send a byte at any moment. The same applies to a thread sitting in a blocking system call. Both leave the check unable to conclude, so the process hangs with no output.
- Forty goroutines are serving traffic and three sit in a mutex cycle. Will the runtime say anything?Never. The check asks only whether anything at all can run, and forty runnable goroutines answer yes. Partial deadlock is the normal production shape and it is entirely invisible to the runtime, which is why you confirm it from goroutine stacks and a flat CPU line instead.
- Why should a minimal reproduction of a suspected deadlock contain no time.Sleep?Because a sleeping goroutine arms a timer, and a pending timer disqualifies the runtime's check. Strip sleeps, tickers and background servers out of the reproduction and the runtime will print the verdict for you instead of leaving you to read stacks.
saying these in an interview costs you the question
- Concludes there is no deadlock because no message appeared
- Thinks the runtime detects a cycle among a subset of goroutines
- Believes a blocked socket read counts as asleep
- Expects the report from a server with an open listener
- Overlooks a ticker as the reason the process never crashed