skip to content

schedtrace and Stack Dumps

The signals the runtime emits about the scheduler: schedtrace lines, a live goroutine count, and the stack dump a SIGQUIT forces. Interviewers ask what you read first when the count only climbs.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

questions

4

What happens when you send SIGQUIT (Ctrl-\) to a running Go program?

level: juniorimportance: should knowfreq 44%

answer

  1. no rebuild, no flag, no endpoint
  2. a signal, not a function call
  3. the terminal key beside Enter
  4. every goroutine's stack, then the process is gone

basics

~20 s

By default the Go runtime prints a stack trace for every goroutine to standard error and then dies from the signal. It is a one-shot diagnostic: the process is killed, and deferred functions never run.

solid answer

~50 s

SIGQUIT is the runtime's built-in "show me everything" switch. Unless the program has claimed the signal itself, the runtime catches it, walks every live goroutine, prints each one's stack to stderr under a header like `goroutine 4711 [chan receive, 143 minutes]:`, and then re-raises the signal so the process is terminated by it. Two things follow. It is fatal, so you get exactly one dump per process and no graceful shutdown — deferred functions and cleanup do not run and in-flight work is dropped. And it is free: no rebuild, no build flag, no debug endpoint, which makes it the first tool for a wedged binary in a terminal (`Ctrl-\`) or on a box (`kill -QUIT <pid>`). If the program calls `signal.Notify` for SIGQUIT, the default dump is disabled and the signal is delivered to the program instead.

code

text · 2 lines
text
goroutine 1 [select]:
goroutine 4711 [chan receive, 143 minutes]:

go deeper

for a junior

Be ready to name the signal or key combination you would use on a hung Go binary and to say that every goroutine's stack lands on standard error. Knowing the process does not survive it is the half most candidates leave out.

for a middle

Explain the mechanism: the runtime installs a handler for SIGQUIT, walks every live goroutine printing its state, wait reason and stack, then re-raises the signal so the process is killed rather than exiting cleanly.

for a senior

Show that you would choose the target instance deliberately and check where stderr goes first, and that you know what defeats the dump in a real deployment: a program that claimed SIGQUIT with signal.Notify, or a supervisor that never forwards it.

for a principal

The call to own is what SIGQUIT means across your services. Claiming it for graceful shutdown trades away the runtime's best zero-setup diagnostic, so decide which signal carries which meaning, write it down, and make sure stderr is captured everywhere.

## The signal and what the runtime does with it SIGQUIT is a POSIX signal whose default operating-system behaviour is to terminate the process and write a core file. The Go runtime installs its own handler for it, and that handler is why the signal matters here: **before the process dies, the runtime walks every live goroutine and prints its state and stack to standard error.** There are three ordinary ways to send it: - press `Ctrl-\` in the terminal running the program in the foreground; - `kill -QUIT <pid>` from another shell (`kill -3 <pid>` is the same signal by number); - have a supervisor or orchestrator send it to the process. Nothing has to be prepared in advance. There is no build tag, no `-gcflags`, no imported package and no HTTP endpoint involved — any Go binary, including one you did not compile, responds this way unless it has deliberately taken the signal over. ## What the output looks like The dump is a sequence of blocks, one per goroutine, each introduced by a header line: ``` goroutine 1 [select]: goroutine 4711 [chan receive, 143 minutes]: ``` The header carries three things: the runtime's goroutine id, the goroutine's state — and, when it is blocked, the *reason* it is blocked, such as `chan receive`, `chan send`, `select`, `sleep`, `IO wait` or `semacquire` — and, when it has been waiting for a minute or longer, how long it has been waiting, rounded down to whole minutes. The runtime records the moment a goroutine parked, so that duration is real elapsed blocked time, not CPU time. Those headers are usually where the answer is. In a process with tens of thousands of goroutines you do not read the dump linearly; you scan the headers for repeated states and for large ages. A thousand goroutines all reading `[chan receive, 200+ minutes]` is a different story from a thousand reading `[IO wait]` with no age at all. ## The cost: it is fatal This is the half candidates forget. After printing, the runtime restores the default disposition for SIGQUIT and re-raises it, so the process is *killed by the signal*, not exited cleanly. Concretely: - deferred functions do not run; - no shutdown hooks, flushes or connection drains happen; - every in-flight request on that instance is lost; - you get **one** dump from that process, ever. So SIGQUIT is not something you sprinkle over a fleet. You choose an instance you can afford to lose — a canary, a replica behind a load balancer that will be restarted anyway, a job that will be retried — and you take the dump deliberately, knowing it is the last thing that instance does. ## What can stop it working - **The program claimed the signal.** If it calls `signal.Notify(ch, syscall.SIGQUIT)` — or `signal.Ignore` — the runtime hands SIGQUIT to the program's own channel and its default dump-and-die path never runs. Daemons that treat SIGQUIT as "finish current work and stop" do exactly this, and in exchange they give up the runtime's best free diagnostic. - **Nothing forwards the signal.** A shell wrapper, an init process or a container runtime may absorb `Ctrl-\` or send the signal to a PID 1 shim rather than to your binary. - **stderr goes nowhere.** The dump is written to file descriptor 2. If the process was started with stderr pointing at `/dev/null`, or your logging setup only captures stdout, the dump is produced and thrown away. Check where stderr lands *before* you fire the signal, because you do not get a second try. - How much detail the traceback contains is further governed by the `GOTRACEBACK` setting, which is the knob to reach for when the default output is not enough — or is too much. ## Where it fits among the other runtime diagnostics Think of the three cheap, no-rebuild sources as answering different questions. `runtime.NumGoroutine()` answers *how many goroutines exist* — a number, over time, from inside the program. `GODEBUG=schedtrace=1000` answers *what is the scheduler doing* — how many Ps are idle, how deep the run queues are — printed periodically while the process keeps running. The SIGQUIT dump answers *what is each individual goroutine waiting on, and for how long* — once, at the cost of the process. Used in that order, they compose: the counter tells you something is wrong, schedtrace tells you whether the goroutines are competing for CPU or parked, and the dump tells you precisely what they are parked on.

  • Why might Ctrl-\ produce no dump at all on a service you did not write?
    Something took the signal first. If the program calls `signal.Notify` or `signal.Ignore` for SIGQUIT, the runtime delivers it to the program instead of running its default dump-and-die path — common in daemons that treat SIGQUIT as a graceful-stop request. A shell wrapper or container shim can also absorb the key, and if stderr is discarded the dump is written into nothing.
  • What does the `143 minutes` mean in `goroutine 4711 [chan receive, 143 minutes]:`?
    The runtime records when the goroutine last parked, and prints the elapsed wait in the header once it reaches a minute. It means this goroutine has been blocked on a channel receive for over two hours of wall-clock time — not CPU time — which is strong evidence it will never resume, and it lets you sort a huge dump by age.
  • Is it safe to take a SIGQUIT dump from a healthy production instance?
    Only if you are willing to lose it. Producing the dump is cheap, but the runtime then re-raises the signal, so the process is killed without running deferred cleanup and every in-flight request on it is dropped. Pick an instance behind a load balancer that you were going to restart anyway, and make sure stderr is being captured first.

It is a group photograph taken as the building is evacuated: everyone's exact position is captured, and afterwards the building is empty.

saying these in an interview costs you the question

  • Claims SIGQUIT lets the program shut down gracefully
  • Thinks deferred functions run before the process dies
  • Says only the current goroutine's stack is printed
  • Believes a special debug build or flag is required first
  • Expects to take repeated dumps from the same process
open as a page

What does runtime.NumGoroutine() count, and what does it deliberately leave out?

level: middleimportance: should knowfreq 52%

basics

~20 s

runtime.NumGoroutine returns how many goroutines exist right now in any state: running, runnable, sleeping or blocked. It excludes goroutines that have already finished and the runtime's own system goroutines, and the value is a snapshot that can be stale the moment it returns.

open as a page

A daemon's goroutine count climbs every tick and never falls. What do GODEBUG=schedtrace and a SIGQUIT dump each tell you?

level: seniorimportance: should knowfreq 36%

basics

~20 s

schedtrace separates busy from stuck: idle Ps with empty run queues prove the extra goroutines are parked, not starved for CPU. Adding scheddetail=1 gives a line per goroutine with its wait reason, and a SIGQUIT dump gives each one's stack and how long it has been blocked.

open as a page

In a GODEBUG=schedtrace=1000 line, what do `runqueue=12` and the trailing `[3 0 2 0]` mean?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

runqueue is the length of the scheduler's single global run queue, and the bracketed list gives each P's local run queue length, one entry per P. Both count goroutines that are runnable but not currently running; blocked goroutines appear in neither.

open as a page