skip to content

What does Go's `fatal error: all goroutines are asleep - deadlock!` actually mean?

level: juniorimportance: must knowfreq 60%

answer

  1. the scheduler ran out of things to run
  2. a head count, not a wait-for graph
  3. no timeout is involved at all
  4. one blocked goroutine is never enough
  5. unbuffered send with nobody receiving

basics

~10 s

Go's scheduler found nothing left to run: every goroutine is parked on a channel operation, a lock or a similar wait, so none can ever wake another. The runtime aborts the whole program immediately.

solid answer

~50 s

It is the Go runtime reporting that its scheduler has no goroutine it could ever run again. When the last thread executing Go code is about to sleep, the runtime checks whether anything is runnable, running, blocked in a system call, or waiting on a pending timer; if none of those exist, the program can never make progress, so the runtime aborts. The everyday causes are a send or receive on a channel that no other goroutine will ever service, or locking a `sync.Mutex` that is already held by the same goroutine. Two important limits follow from how the check works: it only fires when *every* goroutine is asleep, so a real deadlock between two goroutines while a third keeps working is never reported; and it is a fatal error, not a panic, so no deferred function runs and `recover` cannot see it. The runtime prints the banner, dumps every goroutine's stack, and exits with status 2.

code

go · 7 lines
go
func main() {
	records := make(chan string) // unbuffered: a send needs a waiting receiver
	for _, r := range []string{"row-1", "row-2"} {
		records <- r // no other goroutine exists, so this blocks forever
	}
}
// fatal error: all goroutines are asleep - deadlock!

go deeper

for a junior

Be ready to say it in one line: nothing can run any more, so the runtime gives up. Know the classic cause, an unbuffered channel send or receive with no partner, and be able to point at it in a ten-line program.

for a middle

Explain the mechanism: the scheduler checks whether anything is runnable, in a syscall, or waiting on a timer, and aborts only if nothing is. Be ready to say why a sleeping timer or a network read suppresses it.

for a senior

Show that you do not rely on it. Explain why hung services almost never print it, and describe what you look at instead: the live goroutine stacks, grouped by wait reason and blocked duration.

for a principal

Frame it as a guarantee your team should not build on. Argue for deadlines and cancellation on every blocking path and for goroutine-count alerting, so a hang surfaces as a signal you own rather than as a message the runtime may never print.

## What the message is `fatal error: all goroutines are asleep - deadlock!` is emitted by the Go **runtime**, at run time. It is not a compiler diagnostic and not a vet warning; the program started, did some work, and then reached a state the runtime could prove was terminal. A *goroutine* is Go's unit of concurrent execution, created with the `go` keyword. The runtime multiplexes goroutines onto operating-system threads and keeps counts of how many are runnable, how many are running, and how many are blocked and for what reason. ## How the runtime decides When the last thread that was executing Go code is about to go to sleep because it has nothing to do, the runtime runs an internal check and asks one narrow question: is there **any** goroutine that is running, runnable, or blocked in a system call, and is there **any** pending timer that would make one runnable later? If the answer is no to all of those, nothing that happens next can change the situation, because the only things that could unblock a parked goroutine are other goroutines and timers, and there are none. The runtime therefore aborts. Notice what this check is **not**: - It is **not a timeout**. There is no waiting period; the abort happens the instant the last runnable goroutine parks. - It is **not an analysis** of who is waiting for whom. The runtime never builds a wait-for graph and never reasons about cycles. It counts. - It is **not a race or lock-order check**. It says nothing about lock ordering; it only notices that the machine has stopped. ## Consequences of that definition 1. **Partial deadlocks are invisible.** If two goroutines are stuck on each other but a third is still serving requests, the scheduler always has something to run, so the check never fires. This is why real servers essentially never print this message even when they are badly hung. 2. **A pending timer suppresses it.** `func main() { time.Sleep(time.Hour) }` does not report a deadlock, because a timer will eventually make the main goroutine runnable. 3. **Blocking on network or file I/O suppresses it.** A goroutine parked on a socket read is waiting on the outside world, which the runtime cannot rule out as a source of progress. So the detector is best understood as a **development-time convenience for small programs**, not a production safety net. ## The classic triggers - Sending on an unbuffered channel when no other goroutine is receiving, including the case where the receiver has not been started yet. - Receiving from a channel that nobody will ever send to and nobody will ever close. - Locking a `sync.Mutex` twice from the same goroutine — Go's mutex is not reentrant. - The main goroutine waiting on a result channel while the only worker returned early on an error path without writing to it. ## It is a fatal error, not a panic A panic unwinds the goroutine's stack, running each deferred call, and a `recover` in a deferred function can stop it. A fatal error does none of that. The runtime freezes the other goroutines, prints the banner, prints a stack trace for **every** goroutine with the reason each is parked (for example `goroutine 1 [chan send]:`), and exits with status 2. Deferred cleanup does not run, buffered writers are not flushed, and no amount of `recover` will catch it. That distinction is deliberate: the condition is a property of the whole program, so letting one goroutine handle it locally would be meaningless. ## How you fix it Read the dump rather than guessing: the bracketed reason on each `goroutine N [...]` line tells you whether it is stuck sending, receiving, selecting or locking, and the frames below it give the exact line. The usual fixes are structural — start the consumer before the producer blocks, give the channel a buffer only when the bound is genuinely known, make sure exactly one side closes the channel so receivers can finish, and make every error path still deliver or cancel.

  • Two goroutines deadlock on each other while a third keeps serving traffic. Will the runtime report it?
    No. The check fires only when the scheduler has nothing at all to run. As long as one goroutine is runnable, running, in a syscall, or waiting on a timer, the runtime assumes progress is still possible and stays quiet. That is why a hung production service normally shows no fatal error — it just stops responding.
  • Why does a program whose only goroutine is in `time.Sleep(time.Hour)` not trigger it?
    Because a pending timer counts as a future source of progress. The runtime sees that a goroutine will become runnable when the timer fires, so the terminal condition does not hold. The program simply sleeps for an hour and then continues.
  • How do you fix the case where main sends on an unbuffered channel before any receiver exists?
    Start the consumer first and let main produce, or move the production into a goroutine and have main range over the channel until the producer closes it. Buffering the channel only hides the problem unless you genuinely know the bound; the structural fix is that some goroutine is always ready to receive, and exactly one side closes.

It is a night watchman, not a detective. He does not work out who is waiting for whom; he just notices that nobody in the building is moving any more and closes it down.

saying these in an interview costs you the question

  • Calls it a compile-time or vet check
  • Thinks recover or a deferred handler can catch it
  • Expects it to catch any two-goroutine deadlock
  • Believes the runtime waits for a timeout first
  • Confuses it with a race detector report