skip to content

How do you make a Go binary always dump every goroutine on a panic without GOTRACEBACK?

level: seniorimportance: should knowfreq 30%

answer

  1. the environment is not the only knob
  2. same level names, called from code
  3. one line at the top of main
  4. it can raise but never lower
  5. and a second place to write the report

basics

~20 s

Call runtime/debug.SetTraceback("all") at the top of main. It takes the same level names as GOTRACEBACK and applies them from inside the process, so a binary running where you cannot edit the environment still prints every goroutine when it dies.

solid answer

~40 s

`debug.SetTraceback` in `runtime/debug` takes the same level strings as GOTRACEBACK - none, single, all, system, crash - and applies them to the running process, so `debug.SetTraceback("all")` at the top of `main` bakes full dumps into the binary. One rule matters: a call that would **lower** the level below what GOTRACEBACK already asked for is ignored, so an operator can still raise a shipped agent to system, and you can never accidentally silence what they asked for. Pair it with `debug.SetCrashOutput`, which registers an additional file the runtime writes the whole crash report to, in addition to standard error - that is what saves you when a supervisor restarts the process and throws stderr away. Both have to run before the crash, so open that file during start-up.

code

go · 11 lines
go
func main() {
	debug.SetTraceback("all")
	f, err := os.OpenFile("/var/lib/agent/crash.log", os.O_WRONLY|os.O_CREATE|os.O_APPEND, 0o600)
	if err == nil {
		if err := debug.SetCrashOutput(f, debug.CrashOptions{}); err != nil {
			log.Printf("crash capture disabled: %v", err)
		}
		f.Close() // the runtime duplicated the descriptor
	}
	run()
}

go deeper

for a junior

Know that the crash verbosity GOTRACEBACK controls can also be set from inside the program, by a call in the runtime/debug package, and that it belongs at the very top of main.

for a middle

Explain the precedence rule - a level set in code can raise the environment's but never lower it - and that SetCrashOutput adds a destination for the crash report rather than replacing standard error.

for a senior

Design for the machine you cannot log in to: full dumps compiled in, the report written to a file the next upload can carry, and an honest list of the failures that still produce no Go output at all.

for a principal

Own what leaves the customer's machine. Crash reports carry function arguments, file paths and goroutine counts, so decide who may read them, how long they are retained, and whether support may ship a more verbose build into the field at all.

## The problem this solves GOTRACEBACK is an environment variable, and setting an environment variable assumes you can reach the place where the process is started. For an agent shipped to customer-operated hardware - launched by whatever supervisor that machine runs, on a box you will never get a shell on - you cannot. The only artefact of a crash is whatever the supervisor captured from standard error before restarting the process, and if the binary shipped with the default level, that artefact is one goroutine deep. The stack that mattered was never printed, and the next chance to learn anything is the next crash. So the crash posture has to be a property of the binary, decided at build time, not of the machine. ## `debug.SetTraceback` ```go debug.SetTraceback("all") ``` One call, taking the same level strings as the environment variable: `none`, `single`, `all`, `system`, `crash`. From that point on the runtime prints at that level when the program fails. The rule that makes it safe to ship: **if the level requested is lower than the one the environment already asked for, the call is ignored.** Code can raise the verbosity above the environment's, never below it. Two consequences follow, and both are worth stating in an interview: - A library or a framework you import cannot quietly turn your crash output down. - An operator who sets `GOTRACEBACK=system` on a machine they do control still gets `system`, even though your `main` asks for `all`. Your baked-in level is a floor, not a ceiling. Put the call first in `main`, before any goroutine is started and before any work that can fail, because it only governs failures that happen after it runs. ## `debug.SetCrashOutput` Raising the level is only half the problem. The report still goes to standard error, and on a machine you do not operate you have no guarantee anyone is keeping standard error. `debug.SetCrashOutput(f, debug.CrashOptions{})` registers **one additional file** that the runtime writes the whole crash report to, *in addition to* standard error - it is not a redirect. Its semantics are worth knowing precisely: - The runtime **duplicates** the file descriptor, so you may close your own `*os.File` as soon as the call returns. - There is exactly one such destination; a later call replaces the earlier one, and passing `nil` disables it. - It has to be set up **before** the crash, which means opening the file during start-up while everything still works. For a shipped agent that already buffers telemetry records on local disk between uploads, the natural destination is a small append-only file next to that buffer, uploaded with the next batch. That turns "the customer says it restarted overnight" into a report you can read at your desk. ```go func main() { debug.SetTraceback("all") if f, err := os.OpenFile("/var/lib/agent/crash.log", os.O_WRONLY|os.O_CREATE|os.O_APPEND, 0o600); err == nil { if err := debug.SetCrashOutput(f, debug.CrashOptions{}); err != nil { log.Printf("crash capture disabled: %v", err) } f.Close() // the runtime duplicated the descriptor } run() } ``` ## What this still will not capture Be explicit about the limits, because a senior answer that promises total coverage is wrong: - **Recovered panics.** If a frame recovers, the runtime prints nothing - there is no failure. Whatever you want recorded, your own code has to record. - **Kills from outside.** A SIGKILL, an out-of-memory kill by the kernel, a power cut or a watchdog reset produce no Go output at all, because the runtime never runs again. - **Failures before your call.** Anything that dies during package initialisation, before `main` reaches `SetTraceback`, prints at whatever the environment says. - **Growth.** The runtime appends the whole report; it does not rotate or bound the file. If crashes are frequent, that file is your problem to cap. ## Choosing the level to ship `all` is almost always the right baked-in level for an agent: every user goroutine, no run-time noise, and a report small enough to upload. `system` is for a support build handed to one customer chasing a suspected runtime or cgo problem - it roughly doubles the noise for frames most readers cannot use. `crash` is a different decision altogether, because it aborts the process to produce a core file, and on hardware you do not own you will never collect it.

  • What happens if the program calls debug.SetTraceback("none") while GOTRACEBACK=all is set?
    Nothing - the call is ignored. SetTraceback cannot lower the level below what the environment asked for, so an operator who turned crash output up keeps it. Code can only make the runtime's failure output more verbose than the environment, never quieter.
  • What exactly does debug.SetCrashOutput give you that standard error does not?
    A second destination, chosen by the program, that the runtime writes the whole crash report to: a file on the machine's own disk, or a pipe to a small monitoring process. The descriptor is duplicated, so you may close your copy; a later call replaces the destination and nil disables it. Standard error still receives the report.
  • Which failures will this setup still leave you with nothing to read?
    A panic that some frame recovers, since there is no runtime failure at all; a SIGKILL or an out-of-memory kill by the kernel, where the runtime never gets to print; a crash during package initialisation before your call runs; and power loss. It also does not rotate or size-bound the file for you.

saying these in an interview costs you the question

  • Thinks debug.SetTraceback can silence dumps an operator turned on
  • Believes SetCrashOutput redirects the report away from standard error
  • Plans to open the crash file lazily inside the failing path
  • Assumes recovered panics still produce a runtime crash report
  • Expects a shipped agent to pick up GOTRACEBACK from a shell it never has