Why not leave GODEBUG=gctrace=1,schedtrace=1000 permanently enabled across a fleet of Go services?
answer
- unstructured stderr, not a log record
- the release notes may change the format
- the diagnostic is not free to run
- start-up only, so it costs a restart either way
- a bounded run, not a telemetry contract
basics
~20 sGODEBUG tracer output is unstructured text on standard error whose format Go releases may change, it costs real run-time work, and it can only be switched at process start. It is a bounded diagnostic run, not telemetry.
solid answer
~50 sFour reasons. The output is written by the runtime's own printer to standard error - no timestamps you control, no levels, no fields, no route through `log/slog` - so it interleaves with real logs and is lost entirely if the job runner captures only stdout. The line formats are documented as subject to change between releases, so anything parsing them into dashboards is one toolchain upgrade from breaking silently. It is not free: `schedtrace` keeps the runtime's monitor thread out of deep sleep and prints under the global scheduler lock, `scheddetail=1` walks every P, M and goroutine per tick, and `gctrace`'s volume is unbounded on an allocation-heavy service. And these settings are read only at start-up, so "leave it on so we needn't restart later" is really a gap in your launch tooling - fix that instead.
code
text · 4 lines# relaunch a single instance for the investigation, then relaunch without it
GODEBUG=gctrace=1,schedtrace=5000 ./nightly-import -config /etc/import.toml \
1>>/var/log/import.log \
2>/var/log/import-godebug.loggo deeper
Know that these settings are for a deliberate diagnostic run rather than something you leave switched on, and that their output goes to standard error, not into the application's logs.
Explain the concrete costs: an unstructured stderr stream, line formats the Go release notes may change, and real run-time work such as a monitor thread kept awake to print scheduler lines.
Own the run end to end — the setting, the period, where stderr goes, how long it stays on and how it comes back off — and be able to say why standing runtime numbers belong in a supported metrics surface rather than in parsed trace lines.
Set the platform rule: which diagnostic environment variables a service may be relaunched with and by whom, how the launch path supports that without a code change, and why a dashboard built on a format documented as unstable is a fleet-scale liability.
## The temptation Somebody discovers `GODEBUG=gctrace=1`, gets a useful answer from it during an incident, and proposes making it standard across every service so the data is always there. The instinct is right — you do want continuous visibility — but the mechanism is wrong, and an interviewer asking this is checking whether you can say why in operational terms rather than reciting that it is "for debugging". ## 1. The output is not a log, and not addressed to a machine GODEBUG tracer output is produced by the runtime's own low-level printing routine, straight to file descriptor 2. It does not pass through `log`, `log/slog`, or whatever handler your service configures. There is no severity, no timestamp in your format, no request id, no JSON, and no correlation with anything. It interleaves, unsynchronised, with anything else the process writes to stderr, including panic output. In a job that runs under a scheduler capturing only stdout into the job log, it is simply discarded and the operator concludes the setting did nothing. ## 2. The format is explicitly unstable The runtime documentation says of each of these lines that the format is subject to change. That is not boilerplate: fields have been added and reworked across releases. A dashboard, alert or log-parsing rule built on the shape of a `gctrace` line is a liability that fails quietly on the next toolchain bump — you keep getting a green dashboard fed by a regex that no longer matches. Standing, machine-consumed runtime numbers have a supported home; `runtime/metrics` is that home, and it exists precisely so nobody has to parse trace lines. ## 3. It costs something, and the cost varies by setting - `schedtrace=N` keeps the runtime's monitor thread out of its deep sleep so it can print on schedule, and each line is emitted while the global scheduler lock is held. - `scheddetail=1` turns each of those ticks into a walk over every P, every M and every goroutine, under that same lock. On a goroutine-heavy service this is the difference between a diagnostic and an outage. - `inittrace=1` puts the allocator on its accounting path during start-up. - `gctrace=1` is comparatively cheap per collection, but on an allocation-heavy service the collections are frequent and the output volume is unbounded — it is a steady, pointless write load on stderr for the life of the process. None of these is catastrophic on its own, which is exactly why they end up permanent. Multiply by a fleet and you are paying continuously for output nobody reads. ## 4. One setting in this family is not a tracer at all `asyncpreemptoff=1` reports nothing. It disables signal-based asynchronous goroutine preemption, so a tight loop containing no function calls can run un-preempted for a long time, delaying garbage collection and the scheduling of other goroutines. It also disables the conservative stack scanning the collector uses for asynchronously preempted goroutines. It exists so you can bisect a suspected preemption or stack-scanning problem for one run. Anything that reaches production with it set has changed the runtime's behaviour, not its verbosity, and the symptoms — long pauses, starved goroutines — look like an application bug. ## 5. Start-up only, which is the real complaint underneath These settings are captured into runtime variables during start-up; setting the environment variable from inside a running process does not enable them. So "leave it on permanently" is usually a proxy for "we cannot get a diagnostic build of this process without a painful restart". That is the problem to solve. A launch path where an operator can relaunch one instance with a chosen `GODEBUG` value, with stderr captured to its own file, turns the argument off entirely. ## What good practice looks like Treat a GODEBUG run like a controlled experiment: decide the setting and the period, relaunch one instance with stderr redirected to a dedicated file, record the toolchain version next to the output because the format is release-specific, gather for a bounded window, then relaunch without it. Keep the ready-to-use line in the job definition, commented out, so the next person on call does not have to rediscover it. Continuous numbers come from a supported metrics surface; GODEBUG answers one question, once.
- Which setting in this family changes behaviour rather than only reporting, and when would you use it?`asyncpreemptoff=1`. It disables signal-based asynchronous goroutine preemption, so a tight loop without function calls can run un-preempted for a long time, delaying garbage collection and scheduling, and it also disables the conservative stack scanning used for asynchronously preempted goroutines. It is a bisection tool for one run when you suspect preemption is involved in a crash — never a standing setting.
- Your scheduler captures the job's stdout into the job log, and you see no gctrace lines. What is wrong?Nothing in Go. The runtime writes GODEBUG output to standard error with its own printer, so a runner capturing only file descriptor 1 discards it. Redirect stderr in the launch command or job definition. Also check the variable actually reaches the process: a wrapper script that sanitises the environment before exec will strip it.
- How do you make a one-off diagnostic run repeatable for the next engineer on call?Keep the exact GODEBUG line in the job or unit definition, commented out and ready to enable; send stderr to its own file rather than the application log; record the Go toolchain version alongside the output, since the line formats are release-specific; and write down what you were looking for, so the next person knows whether the setting is still the right one.
saying these in an interview costs you the question
- Treats gctrace output as a stable format to parse into dashboards
- Leaves scheddetail=1 enabled because it is 'only logging'
- Sets asyncpreemptoff=1 believing it reduces scheduling overhead
- Assumes GODEBUG output reaches the application's log pipeline
- Plans to toggle a runtime tracer during an incident without a restart