skip to content

Confirming a Deadlock

A Go process at zero CPU that serves nothing: the fatal all-goroutines-are-asleep report fires only when every goroutine is blocked, so a partial deadlock leaves a SIGQUIT dump and no lock graph.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

questions

4

What has the Go runtime detected when it prints "fatal error: all goroutines are asleep - deadlock!"?

level: juniorimportance: must knowfreq 60%

answer

  1. a claim about the whole program
  2. the scheduler ran out of anything to run
  3. nothing runnable, no timer, no syscall
  4. it is a proof, not a lock analysis

basics

~20 s

Every user goroutine is parked waiting on another goroutine, with nothing runnable, no armed timer and no thread in a system call. Progress is provably impossible, so the runtime dumps every goroutine stack and exits.

solid answer

~40 s

It is the scheduler's proof that the whole program can never make progress again. When the last thread is about to go idle, the runtime checks every non-system goroutine; if all of them are parked on something only another goroutine could satisfy — a channel send or receive, a `select` with no ready case, a mutex, a `sync.WaitGroup.Wait` — and no timer is armed and no thread sits in a system call, nothing can ever wake anything, so it reports the fatal error, prints all goroutine stacks and exits with status 2. It is a no-progress claim about the whole program, not a lock analysis: the runtime keeps no lock graph and never says which goroutine holds what. The classic trigger is a send on an unbuffered channel that no one receives.

code

go · 8 lines
go
func main() {
	// unbuffered channel: this send needs a receiver on another goroutine
	ch := make(chan int)
	ch <- 1
}

// fatal error: all goroutines are asleep - deadlock!
// exit status 2

go deeper

for a junior

Be ready to say what the message proves — nothing in the program can ever run again — and to write the two-line channel program that produces it. Knowing it ends the process rather than raising something catchable is expected.

for a middle

Explain the three conditions the check needs: no busy thread, every user goroutine parked, no armed timer. An interviewer will want you to name a wait that counts and a wait that does not.

for a senior

Show that you treat the report as a whole-program no-progress proof and never as evidence about locks. The valuable half is knowing when it cannot fire, so you do not read silence as health.

for a principal

Frame it as a signal available only in constrained environments — tests and minimal reproductions — and set the team's expectation that production hangs need stack evidence instead. That framing decides what your diagnostics runbook is built on.

## What the message is `fatal error: all goroutines are asleep - deadlock!` comes from the Go runtime's scheduler. No library and no tool produces it — it is emitted by the same code that decides which goroutine runs on which thread, at the moment that code concludes it has nothing left to schedule, forever. ## How the check actually runs The scheduler multiplexes goroutines onto OS threads. When a thread is about to park because it found no runnable goroutine, the runtime runs an internal deadlock check before letting the process go quiet. The check happens in three stages. 1. **Is any thread still busy?** Threads running Go code, threads blocked inside a system call, and threads locked to a goroutine all count as busy. If even one is, the check returns immediately: something may still come back and produce work. 2. **Is any goroutine runnable?** The runtime walks every goroutine, skipping *system* goroutines — the garbage collector's mark workers, the scavenger, the finalizer goroutine and similar runtime-owned goroutines are not counted, because they exist to serve user code rather than to be user work. Every remaining goroutine must be in a waiting state. 3. **Is any timer armed?** The runtime looks at the timer heaps of every P (the scheduler's per-processor contexts). A single pending timer means a goroutine will become runnable at a known future instant, so no claim of permanent deadlock is possible and the check returns. Only if all three stages pass does the runtime have a proof: no event can ever occur that makes any goroutine runnable. It then prints the fatal error, dumps the stack of every goroutine, and exits with status 2. ## What counts as "asleep" Waits that count toward the check are the ones that can only be satisfied by another goroutine inside this process: - a send or a receive on a channel, including one that will never have a counterpart; - a `select` with no ready case and no `default`; - `sync.Mutex.Lock`, `sync.RWMutex.RLock`, `sync.WaitGroup.Wait`, `sync.Cond.Wait` — all of which park through the runtime's semaphore; - `<-ctx.Done()` on a `context.Context` that nothing will ever cancel. Waits that do **not** count are the ones the runtime cannot prove will never end: a goroutine sleeping on a timer, a goroutine blocked reading a file or a socket, a goroutine spinning in a loop, and anything sitting in a cgo call. Any one of those leaves the runtime unable to conclude, and the program simply hangs in silence. ## It is a fatal error, not a panic The output begins with `fatal error:`, not `panic:`. The runtime does not unwind the stack for it, so this is not something application code observes or handles; the process ends there. ## What the message does not mean This is the part that misleads people, and it is worth being precise about. - **It is not lock-cycle detection.** The runtime holds no graph of which goroutine owns which mutex. It never reports "goroutine 7 holds the lock goroutine 9 wants". It only knows that nobody can run. - **It is whole-program.** A cycle among three goroutines while forty others keep serving traffic will never trigger it, because those forty are runnable. - **Its absence proves nothing.** A hung program that prints nothing may be perfectly deadlocked; the check simply could not conclude. In practice the message is a development-time signal. You see it in toy programs, in small command-line tools, and in minimal reproductions written deliberately with no timers and no I/O. Long-running services almost never produce it. ## The sibling message The same check has a second outcome. If every user goroutine is gone — for example the main goroutine called `runtime.Goexit` and left nothing behind — the runtime prints `fatal error: no goroutines (main called runtime.Goexit) - deadlock!` instead. Same proof, different reason: there is nothing left to run rather than nothing able to run. ## Producing it deliberately The smallest reproduction is a send on an unbuffered channel in `main` with no other goroutine. Two goroutines each waiting to receive what the other will only send after receiving will do it too, as will a `sync.WaitGroup` whose counter never reaches zero. If you want the runtime to give you this verdict on a suspected cycle, strip the reproduction down until nothing sleeps, nothing ticks and nothing touches the network — every one of those hides the answer.

  • Does that report mean the program's locks form a cycle?
    No. It is a scheduler-level proof that nothing can run, and it inspects no lock ordering at all. In small programs the usual cause is a channel operation with no counterpart, or a `sync.WaitGroup` counter that never reaches zero — a mutex cycle is only one of the ways to get there, and the message will not tell you which one you hit.
  • Are the garbage collector's background goroutines counted by the check?
    No. The runtime skips its own system goroutines — mark workers, the scavenger, the finalizer goroutine — when it walks the goroutine list. Otherwise the check could never conclude, because those goroutines are always parked waiting for work. Only goroutines started by user code count.
  • What does the runtime print if the main goroutine calls runtime.Goexit and nothing else remains?
    `fatal error: no goroutines (main called runtime.Goexit) - deadlock!`. It is the same check reaching a different conclusion: rather than every goroutine being blocked, there are no user goroutines left at all, so the process can never make progress and exits.

saying these in an interview costs you the question

  • Thinks the runtime detects lock-ordering cycles
  • Assumes the message appears whenever a program hangs
  • Believes a sleeping goroutine still counts as asleep
  • Says only channel misuse can trigger it
  • Expects the runtime to name the blocking mutex
open as a page

Why does Go's "all goroutines are asleep - deadlock!" report never fire in a hung production service?

level: middleimportance: should knowfreq 50%

basics

~20 s

That report needs the whole program stuck: no runnable goroutine, no thread in a system call, no armed timer. One open listener, one resync ticker or one sleeping goroutine is enough to keep the runtime silent forever.

open as a page

In a Go goroutine dump, how do you find which goroutine holds the sync.Mutex the others are waiting on?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Go records no mutex owner, so you infer it. Group blocked goroutines by the hex receiver printed for their sync.(*Mutex).Lock frame; the holder is the goroutine already past a Lock on that same address, parked on something else.

open as a page

Why does a hung `go test` run report a test timeout instead of Go's deadlock error?

level: middleimportance: nice to knowfreq 32%

basics

~10 s

The test harness arms its own timeout timer, and any pending timer stops the runtime from claiming a deadlock. When it fires, the harness panics and dumps every goroutine's stack — the same evidence.

open as a page