Your diagnostics helper dumps every goroutine's stack on demand in a service holding 200,000 goroutines — what does that cost, and what should the helper hand back?
answer
- one of the two options pauses the whole program
- counts beat stacks when the stacks repeat
- let the caller supply the destination
- a compact profile versus megabytes of text
- the trigger belongs to an operator, not a timer
basics
~20 sruntime.Stack with all set to true stops the world for the whole walk and formats megabytes of text, so pause and output both scale with goroutine count. Prefer pprof.Lookup("goroutine").WriteTo, take an io.Writer, and rate-limit the trigger.
solid answer
~50 sTwo costs dominate: `runtime.Stack(buf, true)` keeps the world stopped while it walks *and formats* every goroutine, and the text it produces at 200,000 goroutines is measured in megabytes. Building that into a string or an HTTP response body means a multi-megabyte allocation on top of the pause. The better default in a helper other teams import is `pprof.Lookup("goroutine").WriteTo(w, debug)`: `debug` 0 writes the compact binary profile that `go tool pprof` can aggregate and diff, 1 writes counted text, 2 writes full stacks in the familiar panic-dump layout. Take an `io.Writer` and return an `error` so the caller chooses the destination and you never materialise the dump in memory. Then make it operational: single-flight concurrent requests, rate-limit, never trigger it from a health check or a timer, and expose `Profile.Count()` when the caller only wants the number.
code
go · 9 lines// DumpGoroutines writes a goroutine profile to w.
// debug: 0 = binary pprof, 1 = counted text, 2 = full stacks as text.
func DumpGoroutines(w io.Writer, debug int) error {
p := pprof.Lookup("goroutine")
if p == nil {
return errors.New("goroutine profile unavailable")
}
return p.WriteTo(w, debug)
}go deeper
Know that there are two ways to capture every goroutine from inside a process — runtime.Stack with all true, and the runtime/pprof goroutine profile — and that neither is free.
Explain the difference in encoding and aggregation: the profile can collapse identical stacks and be written in a binary form for tooling, while runtime.Stack always renders full text.
Show the operational reasoning: pause and output scale with goroutine count, so the trigger is explicit and rate-limited, the API streams to a writer, and a cheap count is published continuously instead.
Own the interface other teams will depend on. Fix the signature, the encoding choice and the guardrails before adoption, since a diagnostic that can pause any importer's service is a decision you cannot quietly revise later.
## The two mechanisms, and what each really costs There are two ways to capture the state of every goroutine from inside a running process. **`runtime.Stack(buf, true)`** formats the calling goroutine's stack followed by every other goroutine's, into a buffer you supply. To take a coherent snapshot the runtime stops the world, and it stays stopped for the whole operation — walking frames *and* rendering them into text. Both the pause and the output grow with the number of goroutines. At a few hundred goroutines nobody notices. At 200,000, you are pausing every in-flight request for a long time on human scales, and producing an output measured in megabytes. **`runtime/pprof`'s goroutine profile**, reached with `pprof.Lookup("goroutine")`, produces the same information through the profiling machinery instead. `Profile.WriteTo(w, debug)` streams it to a writer, and the `debug` argument picks the encoding: - `0` — the binary pprof format, compact and meant for `go tool pprof`, which can aggregate identical stacks, diff two captures and render call graphs. - `1` — legacy text: each distinct stack once, with a count of how many goroutines share it. - `2` — every goroutine's full stack as text, in the same layout the runtime prints when a program dies from an unrecovered panic. That aggregation is the real argument for the profile at scale. Two hundred thousand goroutines in a service are rarely 200,000 distinct stacks; they are a handful of stacks with enormous counts. Format 0 or 1 collapses them, which turns megabytes of repetition into something a human or a tool can act on. Format 2 does not collapse anything and is as large as the equivalent `runtime.Stack` text. And if the question is only "how many", `Profile.Count()` answers it without rendering anything at all. ## Designing the helper's API The chair here is the library author: other teams import this package and wire it into binaries you do not operate. Three decisions follow. **Take an `io.Writer`, return an `error`.** A `func Dump() string` forces the entire multi-megabyte dump to be built in memory before the caller sees a single byte, and it gives the caller no way to stream it to a file, a socket or a compressor. The standard library already models the right shape in `Profile.WriteTo`, and matching it means callers can compose you with everything that already accepts a writer. **Expose the encoding, do not choose it silently.** A caller collecting for tooling wants the binary form; an on-call engineer reading a terminal at 3am wants text they recognise. Surfacing the `debug` value as a parameter, documented in your own words, avoids both a wrong default and a fork of the function per format. **Do not make it implicit.** A helper that dumps on every error, on a ticker, or as a side effect of a metrics scrape turns a diagnostic into an outage source. The trigger belongs to the operator: a signal handler, an authenticated admin endpoint, a one-shot flag. ## Operating it safely Several guardrails belong in the helper itself rather than in each caller's head. - **Single-flight it.** Two dumps racing means two stop-the-world pauses back to back, and monitoring reacting to the first pause is exactly what triggers the second. Collapse concurrent requests onto one capture. - **Rate-limit it.** A minimum interval between dumps, enforced inside the helper, bounds the damage a retry loop can do. - **Bound the destination.** Writing straight to a writer avoids the memory spike, but the writer itself can be slow. A slow consumer does not extend the stop-the-world pause for the profile path, since the collection and the writing are separate phases, but it does hold the buffer alive. - **Prefer the count first.** In a lot of incidents the actionable fact is that the goroutine count is 200,000 and climbing. Publish that continuously and cheaply, and reserve the full dump for the moment somebody actually needs stacks. - **Treat the output as sensitive.** Stack traces carry function arguments as raw words, so they can leak values from your process into wherever the dump lands. That shapes who is allowed to trigger it and where the bytes are permitted to go. ## When runtime.Stack is still the right answer It is the right answer when you need the text and cannot depend on tooling: a last-gasp handler writing to standard error, an embedded environment with no way to fetch a binary profile off the box, or a test harness that wants a human-readable snapshot in its failure output. It is also the simplest thing that works for the single-goroutine case, where `all` is false, there is no pause at all, and `debug.Stack()` already wraps it. The scale argument only bites when `all` is true.
- Why is func Dump() string the wrong signature for this helper?Because it forces the whole dump — megabytes at high goroutine counts — to be assembled in memory before the caller sees anything, and it removes every streaming destination. `func(io.Writer, int) error` matches `Profile.WriteTo`, lets the caller stream to a file, socket or compressor, and reports failures properly.
- How do you stop the dump from becoming its own incident?Make the trigger explicit and operator-owned, single-flight concurrent requests so two pauses cannot stack up, and enforce a minimum interval inside the helper. Never wire it to a ticker, a health check or an error path, where a retry loop turns one pause into a sustained one.
- When is the text form still the better choice than the binary profile?When a human has to read it now and there is no tooling within reach — a last-gasp handler writing to standard error, a box you cannot copy files off, or a test harness embedding the snapshot in its failure output. debug 2 gives the panic-dump layout engineers already know how to scan.
- What can you publish continuously instead of dumping stacks?The goroutine count. `Profile.Count()` on the goroutine profile reports how many entries it holds without rendering any stacks, so it is cheap enough to sample regularly. In many incidents the count and its slope are the actionable signal, and the full dump is only needed once.
saying these in an interview costs you the question
- Calling runtime.Stack(buf, true) a cheap memory read
- Returning the whole dump as a string from the helper
- Collecting a full dump on a timer so it is always fresh
- Believing a bigger buffer removes the stop-the-world pause
- Assuming a profile is only useful when served over HTTP
- Ignoring that stack output can carry argument values off the box