Why doesn't Go's 'all goroutines are asleep - deadlock!' error appear when a live server deadlocks?
answer
- a scheduler check, not an analysis
- it fires only when nothing could run
- one pending timer is enough to silence it
- servers always have a goroutine in the poller
basics
~20 sGo's deadlock report fires only when every goroutine is parked with no possible wakeup. A live server always has something wakeable - a goroutine in the network poller, a blocked syscall, a pending timer - so a real deadlock hangs silently instead.
solid answer
~50 sThe runtime's deadlock report is not a deadlock analyser; it is a last-resort check the scheduler makes when it finds nothing left to run. If every goroutine is parked on a channel operation or a mutex and nothing could possibly wake any of them, the runtime throws `fatal error: all goroutines are asleep - deadlock!`. It deliberately stays quiet whenever something might still make progress: a goroutine blocked in a syscall, one waiting on network I/O through the poller, or any pending timer anywhere in the process. A server accepting connections has at least one of those permanently, so a lock cycle between two request handlers produces a hang with idle CPU and a climbing goroutine count, not a crash. The report is therefore a teaching tool for small programs and useless as production monitoring; you diagnose the real thing by dumping every goroutine's stack and reading the long-parked lock and channel waits.
code
go · 9 linesfunc main() {
go func() {
time.Sleep(time.Hour) // a pending timer
}()
var mu sync.Mutex
mu.Lock()
mu.Lock() // blocks forever; the runtime reports nothing
}go deeper
Know that the message means every goroutine is blocked with no way to be woken, and that you usually see it in small programs with a channel nobody sends on or a WaitGroup that is never released.
Explain why a pending timer, a blocked syscall or a goroutine waiting on network I/O suppresses the report, and why that makes it a development aid rather than a production signal.
An interviewer expects the operational picture: the signature of a real deadlock is flat CPU with a climbing goroutine count and no logs, and you can describe capturing and reading a full goroutine dump to find the lock cycle.
Own the detection strategy: which signals alert on a stuck service, how a full goroutine dump is captured from production safely, and what lock-ordering rules the codebase enforces so cycles cannot be introduced casually.
## What the check actually is Go's scheduler keeps track of how many goroutines exist and how many are in a state that could still become runnable. When the last thread goes idle, the runtime performs a sanity check: is there anything at all that could run again? If the answer is no, the situation cannot improve, and the runtime reports it as `fatal error: all goroutines are asleep - deadlock!` and exits. Crucially, that is a scheduler bookkeeping check, not an analysis of your program. It does not look at lock ordering, it does not build a wait-for graph, and it does not know that two of your handlers are waiting on each other. It only knows whether the whole process is definitively stuck. ## What makes it stay quiet The runtime treats several states as 'could still be woken by something outside the scheduler', and any one of them suppresses the report: - **A goroutine blocked in a syscall.** The kernel may return at any moment. - **A goroutine waiting on network I/O.** Anything parked in the network poller - an `Accept`, a socket read - is waiting on the outside world. - **Any pending timer.** A single `time.Sleep`, `time.After`, `time.Ticker` or `context.WithTimeout` anywhere in the process means a wakeup is scheduled, so the runtime waits for it rather than declaring deadlock. - **Library builds.** In `c-shared` and `c-archive` builds a foreign thread may still call in, so the check does not apply. Read that list against any real service. An HTTP server has a goroutine in `Accept` at all times. A worker has a ticker. A client has request timeouts. The report is structurally impossible in production code, which is why almost everyone who has seen it saw it in a toy program or a small test. ## What a real deadlock looks like instead A lock cycle between two request handlers - handler A takes the account mutex then the ledger mutex, handler B takes them in the other order - produces this signature: - Requests stop completing; latency goes to infinity rather than to an error. - CPU is idle. Nothing is spinning; everything is parked. - The goroutine count climbs steadily, because every new request spawns a handler that joins the pile-up on the same lock and never returns. - Memory climbs with it, since each stuck goroutine holds its stack and whatever it had allocated. - Nothing whatsoever is logged. The runtime will not tell you. That combination - flat CPU, climbing goroutines, no errors - is the alert you actually want, because it is the only signal the process emits. ## Diagnosing it You need every goroutine's stack, taken while the process is hung. Sending `SIGQUIT` to a Go process makes the runtime dump goroutine stacks and exit; running the service with `GOTRACEBACK=all` guarantees the dump includes every user goroutine, and `system` additionally includes runtime frames and runtime-created goroutines. Once you have the dump, read it like this: - Goroutines parked for a long time are annotated with how long they have been waiting, for example `[sync.Mutex.Lock, 14 minutes]` or `[chan receive, 9 minutes]`. Long waits are your candidates. - Group the stacks. A hundred goroutines in the same helper waiting on the same lock are victims, not the cause. - Find the small number of goroutines that hold a lock while waiting for something else. Those are the cycle. The fix is usually structural rather than clever: establish a global lock order and hold to it, shrink critical sections so a lock is never held across a call you do not control, or replace nested locks with a single owning goroutine that serialises the state. ## When you will see the report Small programs and focused tests, where the whole process really does consist of the goroutines under study: a receive from a channel nobody sends on, a `WaitGroup` waited on after an `Add` whose matching `Done` never happens, a double `Lock` on a non-reentrant mutex. Those are exactly the cases the check was designed to make obvious, and it does that job well. Just do not carry the expectation into production. ## The takeaway Treat the deadlock report as a development convenience with a very narrow trigger condition. Production deadlock detection is your own: an alert on goroutine count and on request completion, plus the discipline of being able to capture a full goroutine dump from a hung process on demand.
- Which conditions stop the runtime from reporting a deadlock?Anything the runtime believes could still be woken from outside: a goroutine blocked in a syscall, one waiting on network I/O in the poller, or any pending timer anywhere in the process. Library builds such as c-shared are excluded too, since a foreign thread may call in. Only a fully self-contained parked program triggers the report.
- How do you find the cycle once the process is already hung?Capture every goroutine's stack. Running with GOTRACEBACK=all and sending SIGQUIT makes the runtime dump all user goroutines and exit. Look for goroutines annotated with long wait times on sync.Mutex.Lock or a channel operation, discard the crowd waiting on one lock, and find the few that hold a lock while waiting for another.
- Why does a hung service's goroutine count keep climbing?Because the listener is still accepting. Each new request spawns a handler that blocks on the same lock or channel and never returns, so goroutines and their stacks accumulate. Rising goroutine count with flat CPU and no errors is the signature to alert on, since the runtime itself will never say anything.
- Would a test that deadlocks reliably produce the report?Only if nothing else in the test binary is wakeable. A test whose subject uses a ticker, a context timeout, or any network listener will hang instead, and the harness's own timeout is what eventually ends it. Do not rely on the runtime report as the failure mode a test detects.
It is a night watchman who only raises the alarm when every single door is bolted and every light is off. Leave one window open and he says nothing, however badly the building is jammed inside.
saying these in an interview costs you the question
- Calls it a general deadlock detector for production
- Expects a deadlocked HTTP service to crash with a message
- Thinks idle CPU plus a hang guarantees a runtime report
- Believes goroutines blocked on I/O count as asleep here
- Assumes the runtime inspects lock ordering