skip to content

When is GOTRACEBACK=crash worth running in production, and what does it cost?

level: seniorimportance: nice to knowfreq 20%

answer

  1. one rung above system
  2. the operating system gets involved
  3. a file the size of your memory
  4. SIGABRT, and everything you held

basics

~20 s

GOTRACEBACK=crash prints the most verbose Go traceback and then aborts the process with SIGABRT so the operating system can write a core dump. Use it only where cores are configured and collectable, and a memory image is acceptable.

solid answer

~50 s

`crash` sits at the top of the GOTRACEBACK ladder: everything `system` prints - user goroutines, run-time frames, the runtime's own goroutines - and then, instead of exiting cleanly, the process aborts in an operating-system-specific way; on Unix it raises SIGABRT so the kernel can write a core. The payoff is a complete post-mortem: every stack plus the heap, inspectable offline. The costs are real. The operating system must actually be configured to write cores or you get nothing extra; a core is roughly the size of the process's memory and lands on local disk; it contains everything in memory, including credentials and customer data; and the process dies by signal, which some supervisors treat differently from a clean exit. On hardware you do not own and cannot pull gigabyte files off, `all` plus a captured crash report is the better trade.

go deeper

for a junior

Know that crash is the loudest GOTRACEBACK setting and that, unlike the others, it deliberately aborts the process so the operating system can write a core dump instead of exiting normally.

for a middle

Explain the mechanism: crash prints what system prints and then raises SIGABRT rather than exiting, and whether a core file actually appears is decided by the host's core size limit and core pattern.

for a senior

Show that you weigh it: disk on the crashing host, whether you can retrieve the file at all, and the fact that a core is a full image of memory. Say what you would run fleet-wide instead and on which single instance you would enable crash.

for a principal

Treat this as a data-handling decision as much as a debugging one. Decide who may enable core dumps, where cores may be stored and for how long, and what evidence a team must bring before asking for them on hardware a customer operates.

## What `crash` actually does `GOTRACEBACK=crash` is `system` plus one extra step. The runtime prints the full report - every user goroutine, run-time function frames, the goroutines the runtime created for itself - and then, rather than calling exit with the usual failure status, it crashes the process in an operating-system-specific way. On Unix that means raising SIGABRT, whose default action is to terminate the process and let the kernel write a core. That last part is the whole point, and it is also where most of the misunderstanding lives: **Go does not write the core file.** All Go does is die in a way that asks the operating system to. Whether a file appears is entirely the host's decision. ## What you get for it A core is a snapshot of the process's memory at the moment it died. With the matching binary you can inspect not just the stacks - which the traceback already gave you - but the actual values: the contents of a buffer, what was in a map, which struct field was unexpectedly zero, how many records were queued. For a class of bugs that the stacks alone cannot settle - memory corruption, a suspected runtime or cgo problem, a value that is wrong in a way nobody can reproduce - that is the difference between a theory and an answer. ## What it costs **It may silently produce nothing.** The core size limit for the process may be zero, the kernel's core pattern may point at a handler that drops it, the destination may not be writable. The Go traceback still prints in full, which is exactly why people believe the setting worked while no core was ever written. Verify on the target host, not on your laptop. **Size.** A core is roughly resident memory. On an agent sharing a small local disk with buffered telemetry records waiting to be uploaded, a multi-gigabyte file dropped at the worst possible moment can fill the disk and take out the very buffer you need - a diagnostic that destroys evidence. **Data exposure.** A core is an image of everything the process held: credentials, tokens, decrypted customer data, request bodies. It is not debug output, it is a copy of production memory sitting outside every control you normally apply to that data. Where it may be written, who may read it, how it travels and when it is deleted are questions to answer *before* you enable it, not after somebody attaches one to a ticket. **Termination semantics.** The process dies by signal rather than exiting with the usual failure status. Supervisors that distinguish the two may back off restarts, alert differently, or classify it as a crash loop. **Retrieval.** On customer-operated hardware you rarely have a way to pull a large binary file off the machine, and you may not be permitted to. A setting whose only artefact you cannot collect is not a diagnostic. ## When it is right - On hardware **you** own - a canary instance, a soak-test fleet, one node deliberately configured for it. - For a bug you are **actively chasing**, where stacks have already failed to settle the question, most often a suspected runtime, memory-corruption or cgo problem. - With the host prepared in advance: a checked core limit, a known destination with room, and a documented handling and deletion path for the file. ## When it is wrong As a fleet-wide default "because it is the most detailed level". You will pay disk, restart-behaviour and data-handling costs on every host, mostly for cores nobody opens. ## What to run instead For a service, ship `all`: every user goroutine, no run-time noise, output small enough for a log pipeline to keep whole. Raise to `system` when the suspicion moves below your code. Add a second destination for fatal output, from inside the program, so the report survives a supervisor that discards standard error - that combination answers the majority of production crashes without ever producing a memory image. ## What to say in an interview Name the mechanism (SIGABRT, and the OS writes the core, not Go), name the prerequisite (host core configuration, or you get nothing), and name the two costs that make it a decision rather than a setting: a file the size of your memory on the host's disk, and that file containing everything the process knew.

  • You set GOTRACEBACK=crash and got a full traceback but no core file. Why?
    Because writing the core is the operating system's job, not Go's. If the process's core size limit is zero, the kernel's core pattern points at a handler that drops it, or the destination is not writable, the abort produces no file. The Go-side traceback prints either way, which is why the setting looks like it worked.
  • What is the argument against leaving GOTRACEBACK=crash enabled across a fleet?
    Every crash then writes a memory-sized file to local disk on hosts that may be tight on space, and each file is a full copy of process memory - tokens, keys, customer data - outside your normal data controls. It also changes how the process terminates, which some supervisors treat as a different class of failure.
  • Which level would you ship as a service's default instead, and why?
    all: a stack for every user goroutine, no run-time noise, no core file, and a report a log pipeline can keep whole. Raise to system when you suspect the runtime or cgo, and turn on crash only on an instance you own and can collect a core from.

saying these in an interview costs you the question

  • Thinks the Go runtime itself writes the core file
  • Assumes a core appears without any host configuration
  • Treats a core file as ordinary debug output, not a memory image
  • Enables crash fleet-wide because it is the most detailed level
  • Expects the usual clean failure exit status after an abort