skip to content

Symptoms in Production

Working from a symptom back to the Go mechanism behind it: a goroutine count that only climbs, an RSS far above the live heap, or too many open files once traffic arrives.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

22

What has the Go runtime detected when it prints "fatal error: all goroutines are asleep - deadlock!"?

level: juniorimportance: must knowfreq 60%

answer

  1. a claim about the whole program
  2. the scheduler ran out of anything to run
  3. nothing runnable, no timer, no syscall
  4. it is a proof, not a lock analysis

basics

~20 s

Every user goroutine is parked waiting on another goroutine, with nothing runnable, no armed timer and no thread in a system call. Progress is provably impossible, so the runtime dumps every goroutine stack and exits.

solid answer

~40 s

It is the scheduler's proof that the whole program can never make progress again. When the last thread is about to go idle, the runtime checks every non-system goroutine; if all of them are parked on something only another goroutine could satisfy — a channel send or receive, a `select` with no ready case, a mutex, a `sync.WaitGroup.Wait` — and no timer is armed and no thread sits in a system call, nothing can ever wake anything, so it reports the fatal error, prints all goroutine stacks and exits with status 2. It is a no-progress claim about the whole program, not a lock analysis: the runtime keeps no lock graph and never says which goroutine holds what. The classic trigger is a send on an unbuffered channel that no one receives.

code

go · 8 lines
go
func main() {
	// unbuffered channel: this send needs a receiver on another goroutine
	ch := make(chan int)
	ch <- 1
}

// fatal error: all goroutines are asleep - deadlock!
// exit status 2

go deeper

for a junior

Be ready to say what the message proves — nothing in the program can ever run again — and to write the two-line channel program that produces it. Knowing it ends the process rather than raising something catchable is expected.

for a middle

Explain the three conditions the check needs: no busy thread, every user goroutine parked, no armed timer. An interviewer will want you to name a wait that counts and a wait that does not.

for a senior

Show that you treat the report as a whole-program no-progress proof and never as evidence about locks. The valuable half is knowing when it cannot fire, so you do not read silence as health.

for a principal

Frame it as a signal available only in constrained environments — tests and minimal reproductions — and set the team's expectation that production hangs need stack evidence instead. That framing decides what your diagnostics runbook is built on.

## What the message is `fatal error: all goroutines are asleep - deadlock!` comes from the Go runtime's scheduler. No library and no tool produces it — it is emitted by the same code that decides which goroutine runs on which thread, at the moment that code concludes it has nothing left to schedule, forever. ## How the check actually runs The scheduler multiplexes goroutines onto OS threads. When a thread is about to park because it found no runnable goroutine, the runtime runs an internal deadlock check before letting the process go quiet. The check happens in three stages. 1. **Is any thread still busy?** Threads running Go code, threads blocked inside a system call, and threads locked to a goroutine all count as busy. If even one is, the check returns immediately: something may still come back and produce work. 2. **Is any goroutine runnable?** The runtime walks every goroutine, skipping *system* goroutines — the garbage collector's mark workers, the scavenger, the finalizer goroutine and similar runtime-owned goroutines are not counted, because they exist to serve user code rather than to be user work. Every remaining goroutine must be in a waiting state. 3. **Is any timer armed?** The runtime looks at the timer heaps of every P (the scheduler's per-processor contexts). A single pending timer means a goroutine will become runnable at a known future instant, so no claim of permanent deadlock is possible and the check returns. Only if all three stages pass does the runtime have a proof: no event can ever occur that makes any goroutine runnable. It then prints the fatal error, dumps the stack of every goroutine, and exits with status 2. ## What counts as "asleep" Waits that count toward the check are the ones that can only be satisfied by another goroutine inside this process: - a send or a receive on a channel, including one that will never have a counterpart; - a `select` with no ready case and no `default`; - `sync.Mutex.Lock`, `sync.RWMutex.RLock`, `sync.WaitGroup.Wait`, `sync.Cond.Wait` — all of which park through the runtime's semaphore; - `<-ctx.Done()` on a `context.Context` that nothing will ever cancel. Waits that do **not** count are the ones the runtime cannot prove will never end: a goroutine sleeping on a timer, a goroutine blocked reading a file or a socket, a goroutine spinning in a loop, and anything sitting in a cgo call. Any one of those leaves the runtime unable to conclude, and the program simply hangs in silence. ## It is a fatal error, not a panic The output begins with `fatal error:`, not `panic:`. The runtime does not unwind the stack for it, so this is not something application code observes or handles; the process ends there. ## What the message does not mean This is the part that misleads people, and it is worth being precise about. - **It is not lock-cycle detection.** The runtime holds no graph of which goroutine owns which mutex. It never reports "goroutine 7 holds the lock goroutine 9 wants". It only knows that nobody can run. - **It is whole-program.** A cycle among three goroutines while forty others keep serving traffic will never trigger it, because those forty are runnable. - **Its absence proves nothing.** A hung program that prints nothing may be perfectly deadlocked; the check simply could not conclude. In practice the message is a development-time signal. You see it in toy programs, in small command-line tools, and in minimal reproductions written deliberately with no timers and no I/O. Long-running services almost never produce it. ## The sibling message The same check has a second outcome. If every user goroutine is gone — for example the main goroutine called `runtime.Goexit` and left nothing behind — the runtime prints `fatal error: no goroutines (main called runtime.Goexit) - deadlock!` instead. Same proof, different reason: there is nothing left to run rather than nothing able to run. ## Producing it deliberately The smallest reproduction is a send on an unbuffered channel in `main` with no other goroutine. Two goroutines each waiting to receive what the other will only send after receiving will do it too, as will a `sync.WaitGroup` whose counter never reaches zero. If you want the runtime to give you this verdict on a suspected cycle, strip the reproduction down until nothing sleeps, nothing ticks and nothing touches the network — every one of those hides the answer.

  • Does that report mean the program's locks form a cycle?
    No. It is a scheduler-level proof that nothing can run, and it inspects no lock ordering at all. In small programs the usual cause is a channel operation with no counterpart, or a `sync.WaitGroup` counter that never reaches zero — a mutex cycle is only one of the ways to get there, and the message will not tell you which one you hit.
  • Are the garbage collector's background goroutines counted by the check?
    No. The runtime skips its own system goroutines — mark workers, the scavenger, the finalizer goroutine — when it walks the goroutine list. Otherwise the check could never conclude, because those goroutines are always parked waiting for work. Only goroutines started by user code count.
  • What does the runtime print if the main goroutine calls runtime.Goexit and nothing else remains?
    `fatal error: no goroutines (main called runtime.Goexit) - deadlock!`. It is the same check reaching a different conclusion: rather than every goroutine being blocked, there are no user goroutines left at all, so the process can never make progress and exits.

saying these in an interview costs you the question

  • Thinks the runtime detects lock-ordering cycles
  • Assumes the message appears whenever a program hangs
  • Believes a sleeping goroutine still counts as asleep
  • Says only channel misuse can trigger it
  • Expects the runtime to name the blocking mutex
open as a page

What does a "too many open files" error from net.Dial or os.Open tell you?

level: juniorimportance: must knowfreq 55%

basics

~20 s

The process has hit its RLIMIT_NOFILE ceiling on open file descriptors. Every socket, file and open HTTP response body costs one, so the cause is almost always a resource opened on some code path and never closed.

open as a page

How do you dump every goroutine's stack from a running Go service over HTTP?

level: juniorimportance: must knowfreq 60%

basics

~10 s

Import net/http/pprof for its side effect and serve HTTP on a debug port, then fetch /debug/pprof/goroutine?debug=2. That prints the full current stack of every goroutine in the process, one block each.

open as a page

For a Go process, how does an inuse_space heap profile differ from the RSS the operating system reports?

level: juniorimportance: should knowfreq 35%

basics

~20 s

An inuse_space heap profile counts only live Go heap objects, sampled as of the last completed garbage collection. RSS counts every physical page the kernel has given the whole process, including freed-but-retained heap spans, goroutine stacks and runtime structures.

open as a page

How do you find out how many OS threads a running Go process has, given runtime.NumGoroutine counts goroutines?

level: juniorimportance: should knowfreq 40%

basics

~10 s

runtime.NumGoroutine counts goroutines only, never OS threads. For the thread count read the /sched/threads:threads gauge from runtime/metrics, run the process with GODEBUG=schedtrace=1000, or ask the operating system, for example /proc/<pid>/status on Linux.

open as a page

Why does Go's "all goroutines are asleep - deadlock!" report never fire in a hung production service?

level: middleimportance: should knowfreq 50%

basics

~20 s

That report needs the whole program stuck: no runnable goroutine, no thread in a system call, no armed timer. One open listener, one resync ticker or one sleeping goroutine is enough to keep the runtime silent forever.

open as a page

How do http.Transport's MaxIdleConnsPerHost and MaxConnsPerHost decide how many sockets a client holds?

level: middleimportance: should knowfreq 50%

basics

~20 s

MaxConnsPerHost caps the total connections to one host, including in-flight ones, and defaults to 0, meaning unlimited. MaxIdleConnsPerHost caps only how many finished connections are kept for reuse, and defaults to 2; surplus ones are closed rather than pooled.

open as a page

In a goroutine dump, what does the header `goroutine 4821 [chan send, 47 minutes]:` tell you?

level: middleimportance: should knowfreq 42%

basics

~20 s

4821 is that goroutine's runtime id, chan send is the reason the runtime parked it, and 47 minutes is how long it has been blocked there — a figure printed only after roughly a minute of waiting.

open as a page

What does runtime/debug.FreeOSMemory do that an ordinary Go garbage collection cycle does not?

level: middleimportance: should knowfreq 32%

basics

~20 s

FreeOSMemory forces a garbage collection and then tries to hand as much free heap memory back to the operating system as it can, at once. An ordinary cycle frees objects but leaves a background task to return pages gradually.

open as a page

In runtime.MemStats, what do HeapAlloc, HeapIdle, HeapReleased and Sys each measure?

level: middleimportance: should knowfreq 42%

basics

~20 s

HeapAlloc is bytes in live heap objects. HeapIdle is bytes in spans holding no objects. HeapReleased is the part of that already handed back to the OS. Sys is all virtual address space the runtime obtained from the OS.

open as a page

What does runtime/debug.SetMaxThreads do, and what happens when a Go program hits that thread limit?

level: middleimportance: should knowfreq 30%

basics

~20 s

debug.SetMaxThreads caps how many OS threads the Go runtime may create and returns the previous ceiling; the initial one is 10,000. Crossing it is not a panic: the process dies with a fatal thread-exhaustion error that recover cannot catch.

open as a page

In a Go goroutine dump, how do you find which goroutine holds the sync.Mutex the others are waiting on?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Go records no mutex owner, so you infer it. Group blocked goroutines by the hex receiver printed for their sync.(*Mutex).Lock frame; the holder is the goroutine already past a Lock on that same address, parked on something else.

open as a page

Your Go reverse proxy hits "too many open files" an hour into peak. How do you find the cause?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Chart the process's open descriptor count against request rate and against the descriptor limit it actually has. A count that rides concurrency means no ceiling on outbound connections; a count that only climbs means a response body on some path is never closed.

open as a page

A Go streaming service's goroutine profile total climbs for days while its heap stays flat. How do you confirm a leak and find the call site?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Take two goroutine profiles minutes apart from the same process and diff them: a leak is the one stack whose count keeps growing. Then read the full dump and find the call site those goroutines are parked at. A flat heap proves nothing.

open as a page

An index service rebuilds a snapshot hourly; its heap profile halves each time but container memory never drops. Leak or retention?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Read runtime.MemStats beside the profile. If HeapAlloc really fell and HeapIdle minus HeapReleased grew, nothing leaked: the runtime kept the pages from the rebuild peak. If HeapAlloc stayed high, the old snapshot is still reachable and it is a leak.

open as a page

A Go transcoder's OS thread count climbs past a thousand while its goroutine count stays flat — how do you diagnose that?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A flat goroutine count with rising OS threads means goroutines are stuck in long blocking calls into C that hold their thread. Confirm with the threadcreate and goroutine profiles, then bound how many such calls run at once.

open as a page

Why does a hung `go test` run report a test timeout instead of Go's deadlock error?

level: middleimportance: nice to knowfreq 32%

basics

~10 s

The test harness arms its own timeout timer, and any pending timer stops the runtime from claiming a deadlock. When it fires, the harness panics and dumps every goroutine's stack — the same evidence.

open as a page

What does Go's /debug/pprof/threadcreate profile record, and how do you use it to find what creates threads?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

Go's threadcreate profile records the stack that led to each new OS thread, aggregated by call site, so it shows which code path keeps producing threads. It counts creation events, not live threads, and is known to under-report.

open as a page

What does Go's goroutineleak profile report that the ordinary goroutine profile cannot?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

The ordinary goroutine profile lists every blocked goroutine and leaves the judgment to you. The goroutineleak profile reports only goroutines the runtime can prove will never resume, so a hit is proof — but an empty result is not.

open as a page

When the Go runtime returns heap pages to the Linux kernel, which madvise mode does it use by default?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

On Linux the Go runtime uses MADV_DONTNEED by default, so RSS falls as soon as pages are released. Setting GODEBUG=madvdontneed=0 switches to MADV_FREE, which is cheaper but leaves RSS high until the kernel needs the memory.

open as a page

When a Go service exhausts descriptors, how do you decide between raising RLIMIT_NOFILE and capping connections?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Ask whether the descriptor count is bounded by design. If the code has an explicit ceiling that is legitimately above the limit, raise the limit. If the count grows with load and nothing caps it, raising the limit only buys hours and moves the failure somewhere worse.

open as a page

Your Go service needs a blocking C codec — how do you decide between keeping it in-process with a thread cap and moving it out?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Decide from measured numbers: concurrent calls times call duration is the OS thread demand, and each thread costs a stack. Keep the codec in-process only behind an owned concurrency bound; move it out when the memory or blast radius is unacceptable.

open as a page