skip to content

How do you prove a service still needs its GODEBUG opt-out after a Go upgrade?

level: seniorimportance: should knowfreq 32%

answer

  1. do not decide from code review
  2. the runtime counts it for you
  3. counts behaviour, not the setting
  4. watch a full business cycle
  5. zero over months is the evidence

basics

~20 s

Read the runtime's /godebug/non-default-behavior counter for that setting: it increments only when the program actually took the old code path. Zero across a full traffic cycle, including rare batch paths, is the evidence to remove the setting.

solid answer

~50 s

Do not guess from code review — measure. For every compatibility setting the runtime exposes a counter named `/godebug/non-default-behavior/<setting>:events`, and it increments only when the program's behaviour actually differed from the default because that setting was non-default. Export it, then watch it across a full traffic cycle: a week of production, plus the month-end batch and the paths that only fire on retries, because those are where the dependence usually hides. A non-zero counter names the code you still have to fix and gives you a rate to judge urgency by. A flat zero, over a long enough window, is the argument for deleting the setting. Back it with an A/B: flip the setting off on part of the fleet and compare. Do this on a schedule rather than on demand — compatibility settings are removed after a documented window, and when the release that removes yours lands, the setting simply stops having an effect.

code

text · 2 lines
text
/godebug/non-default-behavior/panicnil:events
/godebug/non-default-behavior/asynctimerchan:events

go deeper

for a junior

Know that the Go runtime exposes a per-setting counter of how often the old behaviour was actually taken, so this is a measurement question rather than a code-reading one.

for a middle

Explain what the counter counts and why that differs from whether the setting is present, and describe how the environment variable makes a no-rebuild A/B possible.

for a senior

Show the operational judgment: which rare paths you insist on covering before trusting a zero, and that removal is scheduled against the release that deletes the setting rather than against your backlog.

for a principal

Own the standing process — every opt-out instrumented and dashboarded at the moment it lands, reviewed at each toolchain upgrade, with a named owner for the fix when the counter is not zero.

## Why this needs measuring at all A GODEBUG opt-out is invisible by construction. It restores the behaviour your code already expected, so nothing in the service misbehaves, no log line appears, and no test fails. The result is that opt-outs outlive their reason by years: the person who set it has moved on, nobody can name the code path that needed it, and the team's honest answer to "can we remove this?" is "nobody knows". The Go runtime anticipated exactly this and instrumented it. ## The counter For each compatibility setting there is a metric named: ``` /godebug/non-default-behavior/panicnil:events /godebug/non-default-behavior/asynctimerchan:events ``` The important word is *behavior*. It does not count how many times the setting was read or whether it is set — it counts occasions on which the program actually behaved the old way **because** the setting was non-default. If your service sets `panicnil=1` and never panics with a nil value, that counter stays at zero and the setting is doing nothing for you. That makes it a decision instrument rather than a curiosity: - **Non-zero, steady:** you genuinely depend on the old behaviour. The rate tells you how hot the path is, and the counter's existence tells you exactly which behaviour to go and fix. - **Non-zero, rare and spiky:** the dependence lives on an unusual path — an error branch, a retry, a scheduled job. This is the case that would have made a code-review-only removal an outage. - **Zero over a full cycle:** the strongest evidence you can get that removing the setting is safe. ## What a full cycle means This is where teams get it wrong. A day of steady traffic is not a cycle. Include the month-end or quarter-end batch, the nightly reconciliation, the deploy path, the failover drill, and the error paths that only run when a dependency is down. A leaked goroutine or a timer edge case may only be reached when something else is already failing, and "we watched it for 24 hours and it was zero" is precisely how that gets missed. Aim for at least one full business cycle, and prefer instrumenting once and leaving the counter on a dashboard for months over running a one-off check. ## Two supporting techniques **A/B the fleet.** Because the environment variable overrides whatever the module declares, you can start a fraction of instances with the setting flipped to the new behaviour while the rest stay on the old one, and compare error rates and the affected behaviour directly. It is a cheap, reversible experiment: no rebuild, and rolling back is a restart. Keep the two halves comparable — same traffic mix, same version otherwise — or the comparison proves nothing. **Pin the exit date to the toolchain, not to a sprint.** Compatibility settings are kept for a documented window and then deleted along with the old code path. When that happens the setting stops having any effect, and it does so quietly rather than by failing the build or the boot. `asynctimerchan`, the opt-out for the Go 1.23 timer-channel change, was removed in Go 1.27; a fleet still leaning on it simply got the new behaviour the day it upgraded. So the removal ticket should carry the release that deletes the setting, and the upgrade checklist should ask whether any setting in play is on its last release. ## Putting it together as a routine 1. Every compatibility setting you add is recorded in the module and paired with its counter on a dashboard. 2. The counter is reviewed at each toolchain upgrade, not on demand. 3. Zero across a full cycle plus a clean A/B is the standing bar for deleting the setting; a test that pins the new behaviour goes in with the deletion so nothing quietly reintroduces the dependence. 4. Non-zero means the fix is scheduled with a named owner, because the deadline is set upstream and will arrive whether or not the work does. The habit worth carrying out of this is the general one: an opt-out with no measurement attached is not a decision, it is a deferral with no end date.

  • What exactly increments the /godebug/non-default-behavior counter for a setting?
    An occasion on which the program actually took the old code path because that setting was non-default — not the act of reading the setting, and not merely having it set. That distinction is what makes the counter useful: a service that sets an opt-out but never reaches the affected behaviour reports zero, which is the signal that the setting can go.
  • The counter is zero after 24 hours. Is that enough to remove the setting?
    No. Dependence on old behaviour usually hides on paths that do not run every hour: month-end batches, retry and error branches, failover drills, the deploy path itself. Watch across a full business cycle, and prefer leaving the counter on a dashboard for months. Twenty-four hours of steady traffic mostly proves the happy path does not need it.
  • What happens on the day the Go release that removes your setting lands?
    The old code path is gone, so the setting stops having any effect and the program behaves the new way. Nothing fails the build to warn you. That is why a removal ticket should name the release that deletes the setting rather than a sprint, and why the toolchain upgrade checklist should ask which settings are on their last release.
  • How would you A/B a candidate removal without rebuilding?
    Start a fraction of the fleet with the GODEBUG environment variable set to the new behaviour; the run-time value overrides whatever the module declares, so no rebuild is needed and rollback is a restart. Keep the two halves otherwise identical in version and traffic mix, and compare error rates and the affected behaviour directly rather than only overall latency.

saying these in an interview costs you the question

  • Decides from code reading that the setting is unused
  • Thinks the counter increments merely because the setting is set
  • Judges safety from a single day of traffic
  • Assumes a compatibility setting is supported indefinitely
  • Expects a loud failure when a removed setting is still passed