A crashed queue worker's GOTRACEBACK=all dump shows 300 goroutines parked in `chan receive` for 40 minutes — how do you read it?
answer
- the crowd is not the crash
- collapse identical stacks into one fact
- created by names the go statement
- look for the send side of that channel
- duration plus a missing sender
basics
~20 sStart with the goroutine marked running under the crash message; the parked ones are context. Then group identical stacks, read their created by line to find the one go statement that spawned them, and ask which goroutine was supposed to be sending.
solid answer
~50 sFirst separate the failure from the crowd: the `panic:` or `fatal error:` line and the goroutine in state `running` under it are the crash — `goroutine 1` is simply the one running `main`, not necessarily the culprit. Then collapse the 300 into one fact by grouping identical stack signatures, and read the group's `created by` line, which names the `go` statement that made them by file and line plus the parent goroutine's id; near-consecutive ids mean one loop spawned them at once. Then ask the real question: which goroutine owns the send side of that channel? If no goroutine in the dump can send, the producer returned without closing and the workers will wait forever — the bug is on the producer's exit path. The 40 minutes is what turns "blocked" into a finding; blocking on a receive is what an idle worker pool looks like.
code
text · 17 linesgoroutine 1 [running]:
main.(*Consumer).shutdown(0xc0000ac000)
/srv/worker/consume.go:118 +0x51
main.main()
/srv/worker/main.go:34 +0x1a5
goroutine 84 [chan receive, 40 minutes]:
main.(*Consumer).worker(0xc0000ac000)
/srv/worker/consume.go:74 +0x8c
created by main.(*Consumer).start in goroutine 1
/srv/worker/consume.go:66 +0x11d
goroutine 85 [chan receive, 40 minutes]:
main.(*Consumer).worker(0xc0000ac000)
/srv/worker/consume.go:74 +0x8c
created by main.(*Consumer).start in goroutine 1
/srv/worker/consume.go:66 +0x11dgo deeper
You are unlikely to be handed a 300-goroutine dump, but know that the crash is at the top and that the rest of the output is other goroutines, not other errors.
Be able to walk the mechanics: find the running goroutine, group identical stacks, and use a created by line to reach the go statement that produced a group.
Show the judgment: that a wait reason plus a duration plus a missing counterpart is the finding, and that you would look for the send side before blaming the receivers.
Own the conditions that make a crash readable at all: whether dumps are captured and retained, what traceback level production runs at given the output volume, and whether goroutines carry labels that survive into a dump.
## The dump is not a list of suspects A `GOTRACEBACK=all` crash dump prints every goroutine in the process. In a queue-consuming worker that is mostly a wall of goroutines parked in `chan receive` — which is exactly what an idle worker pool looks like. The reading job is to separate the one goroutine that failed from the several hundred that are merely waiting, and then to turn the waiting ones into a statement about the code. ## Step 1 — find the crash, not the crowd The failure is at the top of the output: the `panic:` or `fatal error:` line, followed by the goroutine that produced it, which is in state `running` (for a deadlock detected by the runtime, there is no running goroutine at all — that is itself the finding). Everything after that first goroutine is context. Two misreadings to avoid: `goroutine 1` is the goroutine running `main`, not necessarily the one that crashed; and the goroutine printed last has no special meaning. ## Step 2 — collapse identical stacks Three hundred goroutines with the same frames are one fact, not three hundred. Group them by stack signature — the ordered list of function names, ignoring the hex words and the goroutine ids. What survives is usually a handful of distinct groups: the workers, an HTTP or gRPC-shaped listener, a metrics ticker, the runtime's own housekeeping. Now each group is one question. ## Step 3 — read the `created by` line Each non-main goroutine's stack ends with ``` created by main.(*Consumer).start in goroutine 1 /srv/worker/consume.go:66 +0x11d ``` which is the `go` statement that made it, by file and line, plus the id of the parent goroutine. That line is what takes you from "a group of stuck goroutines" to a specific loop in a specific function — often the only link, because nothing in the stack itself mentions the code that set the work up. If all three hundred share one `created by` line and near-consecutive ids, they were spawned by one loop; if their ids are spread across the whole range, they accumulated over the process's lifetime, which is a very different story. ## Step 4 — ask who was supposed to be on the other end A goroutine parked in `chan receive` is waiting for a sender. So search the dump for one. Read every group's frames and ask which goroutine owns the send side of that channel. Three outcomes: - **There is a sender, also blocked** — you have a cycle; read its state and its channel too. - **There is a sender, running or in `IO wait`** — the pipeline is alive and slow, not stuck; the duration is your evidence for which. - **There is no sender anywhere** — the producer returned, and because the channel was never closed the receivers will wait forever. This is the common shape behind "the worker went quiet". The fix is on the producer's exit path, not on the workers. The duration field turns this from a guess into a claim: forty minutes on a queue that should deliver continuously says the sender has been gone for forty minutes, and the ids and `created by` lines say when and by whom the receivers were created. ## Step 5 — know what the dump is hiding At the default level the runtime elides its own frames, so a parked goroutine's top frame is your function, not the `runtime.gopark`/`runtime.chanrecv` pair underneath it. Re-reading the crash with the traceback level raised to `system` restores those frames and adds the runtime's own goroutines, which tells you precisely which runtime primitive each goroutine parked in — useful when your top frame is ambiguous, noisy otherwise. Deep stacks are truncated with `...additional frames elided...`. A goroutine executing on another thread may print `goroutine running on other thread; stack unavailable` instead of frames. And a dump is one instant: it shows where goroutines *are*, never where they have been, so a fast pipeline and a stopped one can look similar for anything under a minute of waiting. ## What raises the answer above a procedure The judgment an interviewer is listening for is that **blocking is not a bug**. The finding is never "goroutines are blocked on a channel"; it is "these workers, created by this line, have been waiting this long, and no goroutine in this process is capable of sending to them." Everything above is how you get to that sentence from a wall of text. One modern aid is worth knowing: in Go 1.27 tracebacks carry a goroutine's `runtime/pprof` labels for modules that declare `go 1.27` or later, so labels attached with `pprof.Do` — a queue name, a tenant, a message kind — appear beside the stack. In a dump of three hundred otherwise identical workers, that is often the difference between "they are all stuck" and "every one of them is stuck on the same queue".
- The dump contains no goroutine anywhere with a send on that channel in its stack. What does that tell you?That the producer has already returned. Since the channel was never closed, the receivers can never be woken and will block for the life of the process. The bug is on the producer's exit path — an early return, a swallowed error, or a missing close — not in the workers, and no amount of worker tuning will move it.
- Why might you re-read the same crash with the traceback level raised to system?The default level elides runtime frames, so a parked goroutine's top frame is your function with nothing beneath it. At system level you also see the runtime frames it parked in and the runtime's own goroutines, which pins down exactly which primitive each goroutine is waiting on. The cost is a far longer dump, so use it when your own top frame is ambiguous.
- Is "300 goroutines are blocked on a channel receive" a bug report?No. That is what a healthy idle worker pool looks like at any instant. The report needs the other three parts: how long they have been waiting, the single `go` statement that created them, and the fact that nothing in the process can send to them. State alone describes a location; the finding is the combination.
saying these in an interview costs you the question
- Blames the goroutines in chan receive without checking who sends
- Assumes goroutine 1 is always the one that crashed
- Counts goroutines and calls it a leak with no stack grouping
- Ignores the created by lines that name the spawn site
- Reads the dump as a live view rather than one instant