A daemon's goroutine count climbs every tick and never falls. What do GODEBUG=schedtrace and a SIGQUIT dump each tell you?
answer
- existing is not the same as runnable
- the summary line never counts goroutines
- one dump, and then the process is gone
- the ages are spaced like the ticks
basics
~20 sschedtrace separates busy from stuck: idle Ps with empty run queues prove the extra goroutines are parked, not starved for CPU. Adding scheddetail=1 gives a line per goroutine with its wait reason, and a SIGQUIT dump gives each one's stack and how long it has been blocked.
solid answer
~50 sFirst separate "too much work" from "work that never finishes". Run one restarted instance with `GODEBUG=schedtrace=1000`: if it shows `idleprocs` equal to `gomaxprocs` with the global and per-P run queues at zero while the count keeps climbing, the accumulated goroutines are not competing for CPU — they are parked, so this is accumulation by blocking, not a scheduling problem. The summary line never reports how many goroutines exist, so add `scheddetail=1`, which prints a line per goroutine with its status and wait reason; you should see the population grow by one per tick with the same reason repeated, say a channel receive. Then, on an instance you can afford to lose, send SIGQUIT: the dump gives each goroutine's stack plus a blocked-for-N-minutes age in its header, and ages spaced exactly one tick apart pin the accumulation to one step of the tick cycle.
code
text · 1 lineSCHED 300010ms: gomaxprocs=4 idleprocs=4 threads=11 spinningthreads=0 needspinning=0 idlethreads=8 runqueue=0 [0 0 0 0]go deeper
You would not be asked to run this investigation, but know the two halves: a goroutine that exists is not necessarily running, and a stack dump shows what each one is waiting on.
Explain what each instrument contributes — schedtrace for runnable versus parked, scheddetail for per-goroutine wait reasons and a growth rate, the dump for stacks and ages — and why the summary line alone can never confirm accumulation.
Demonstrate sequencing under production constraints: state what would falsify the leak hypothesis before you look, pick a sacrificial instance and check where stderr goes, and refuse capacity changes when the evidence says blocked rather than busy.
Own what is available before an incident: whether a canary can be restarted with GODEBUG at all, whether SIGQUIT still belongs to the runtime rather than to a shutdown handler, and what an on-call engineer is authorised to kill at three in the morning.
## The situation A cron-like daemon wakes on a fixed interval and starts a goroutine per tick to handle that tick's job record. A gauge fed by `runtime.NumGoroutine()` has been climbing all week in a straight line, and the process has not fallen over yet. Memory is up but not alarming. Nothing else looks wrong. The count tells you *that* something accumulates. It cannot tell you whether the daemon is producing work faster than it can run it, or producing work that never finishes. Those have opposite fixes — capacity versus cancellation — so resolving them first is the whole job. ## Step one: is the population runnable or parked? Restart one instance with `GODEBUG=schedtrace=1000`. Every second you get a line summarising Go's goroutine scheduler. Two shapes matter: - **`idleprocs=0`, `runqueue` growing.** Every P is busy and the backlog of *runnable* goroutines is deepening. This is saturation: the daemon really is producing more work than it can execute, and the count climbs because starts outpace completions. - **`idleprocs` equal to `gomaxprocs`, global and per-P run queues at zero.** Nothing is runnable. The Ps are idle because there is nothing for them to do, and yet the population keeps growing. The goroutines are parked, blocked on something that is not going to complete. The second shape is the one this daemon shows, and it is decisive. It immediately rules out the reflex fixes — raising `GOMAXPROCS`, adding instances, adding concurrency — because none of them help work that is not waiting for CPU. Spending your one change budget there is the expensive mistake this measurement prevents. Note what schedtrace's summary line deliberately does *not* give you: a goroutine count. Every number on it is about runnable work, so the line from a healthy idle process and the line from this one look the same. ## Step two: how many, and waiting on what? Add `scheddetail=1`, so the setting reads `GODEBUG=schedtrace=1000,scheddetail=1`. The one-liner is replaced by a multi-line dump each interval: a line per P, a line per M, and a line per goroutine giving its numeric status and, for a waiting goroutine, its wait reason in parentheses — `chan receive`, `select`, `semacquire`, `sleep` and so on. Now you can count. Compare two dumps an interval apart: - the number of goroutine lines grows by exactly one per tick — a rate that matches the daemon's schedule, not the request rate or anything else; - the new lines all carry the *same* wait reason. That is the confirmation: one goroutine per tick parks on the same kind of wait and never leaves it. A leak has a rate and a signature, and detail mode gives you both without touching the code. Detail mode is not free. It locks the scheduler while it walks every P, M and goroutine, and it prints a line per goroutine per interval — with tens of thousands of goroutines that is megabytes of stderr per minute and a real slowdown. Use a slow interval, one instance, and turn it off when you have the answer. ## Step three: where exactly, and since when? Schedtrace tells you the wait reason but not the code. For that you need stacks, and the zero-setup way to get them is SIGQUIT: `kill -QUIT <pid>` makes the runtime print every goroutine's stack to stderr and then kills the process. Budget for that before you fire it: - choose an instance you can afford to lose — the process will not survive, deferred cleanup will not run, and in-flight work on it is dropped; - confirm stderr is actually captured, or the dump goes nowhere and you have spent the instance for nothing; - if the daemon has claimed SIGQUIT with `signal.Notify` for its own shutdown handling, the runtime's dump never runs and you need another route. What the dump adds is the *age* in each goroutine's header: ``` goroutine 5122 [chan receive, 61 minutes]: goroutine 5123 [chan receive, 60 minutes]: goroutine 5124 [chan receive, 59 minutes]: ``` Consecutive goroutines whose blocked times are spaced one tick apart, running back to the age of the process, are an accumulation you can date. The oldest is as old as the process, which is what distinguishes a leak from a slow backlog. And the stacks show which wait in the tick's code path is the one nobody ever satisfies. ## What would have falsified the hypothesis Decide this before you look, or you will find what you expect: - ages all in seconds rather than tens of minutes — a legitimate in-flight backlog that drains, i.e. slow, not leaking; - a population that plateaus instead of tracking uptime — bounded concurrency working as designed; - `idleprocs=0` with a growing run queue — saturation, an entirely different fix. ## The shape of the fix Once the wait is identified, the remedy belongs to the code path, not the runtime: give the wait a deadline or a cancellation it must respect, and make the tick's goroutine unable to outlive the tick. The runtime's diagnostics have done their job when they have told you which wait, at what rate, since when.
- What evidence would have told you this is not a leak?Ages that stay small. If every parked goroutine's header shows seconds rather than tens of minutes, and the population plateaus instead of tracking process uptime, you are watching a backlog of real in-flight work that drains — slow, not leaking. A leak's oldest members are as old as the process itself.
- Why not simply raise GOMAXPROCS or add instances when the goroutine count climbs?Because schedtrace already showed every P idle and both run queues empty. Nothing is waiting for CPU, so more parallelism gives more capacity to run work that is not runnable; the count keeps climbing and the change budget is spent. Capacity is the answer to the other shape — busy Ps and a deepening global run queue.
- You can only restart one instance with GODEBUG set. What do you set, and why not fleet-wide?`GODEBUG=schedtrace=1000,scheddetail=1` on a single canary. Detail mode locks the scheduler while it walks every P, M and goroutine and prints a line per goroutine per interval, so with tens of thousands of goroutines it produces megabytes of stderr a minute and slows the process. The plain one-line form is cheap enough to leave running if you only need the runnable-versus-parked distinction.
- The daemon handles SIGQUIT itself for graceful shutdown, so no dump appears. What now?The runtime's default dump is bypassed once the program claims the signal, so you lean harder on what is left: scheddetail=1 already gives you a per-goroutine status and wait reason and a growth rate, which localises the wait even without stacks. Longer term, treat "SIGQUIT means shutdown" as a decision to revisit, since it costs the service its cheapest stack dump.
saying these in an interview costs you the question
- Concludes the scheduler is starved without checking idleprocs
- Raises GOMAXPROCS in response to a climbing goroutine count
- Expects the schedtrace summary line to report a goroutine count
- Sends SIGQUIT to a healthy instance that is serving traffic
- Reads a large goroutine population as proof of CPU load