skip to content

Would you leave Go's block and mutex profilers enabled permanently in production, and at what rates?

level: principalimportance: should knowfreq 28%

answer

  1. two owners, one hot path
  2. cost per event, not per second
  3. the two profilers are not symmetric
  4. make it adjustable, not permanent
  5. canary the expensive one, document the number

basics

~20 s

Usually yes for the mutex profiler at a coarse fraction, and off or very coarse for the block profiler. Cost scales with how often goroutines block, so measure it per service and make both rates adjustable at runtime.

solid answer

~50 s

This is a standing tradeoff between evidence at 3am and throughput every other hour, and it should be decided with numbers. The cost of both profilers scales with **event frequency**, not uptime: a compute-bound service pays nearly nothing, while a chatty channel-heavy pipeline can pay several percent. So benchmark the real hot path at candidate rates and set the default from that. Typically the mutex profiler stays on permanently at a coarse fraction, because contention events are rare and it is the profile that names a bottleneck; block profiling defaults to off or very coarse, because it fires on every channel operation that parks. What resolves the argument is the mechanism, not the number: expose both rates behind an authenticated admin endpoint so on-call can raise them on one instance mid-incident, and record the measured overhead beside whatever default you ship.

go deeper

for a junior

You are not expected to set fleet policy, but know that these profilers are off by default for a reason and that turning either on in production is a decision someone owns rather than a free switch.

for a middle

Be able to explain that the cost is charged per blocking or contention event, so the same rate is cheap in one service and expensive in another, and that the two profilers differ sharply in how often they fire.

for a senior

Show the operational instinct: measure the overhead on the real hot path, prefer a coarse standing rate over full fidelity, and build a runtime control so an incident never needs a redeploy to get data.

for a principal

Own the tradeoff explicitly. Name who decides and who can overrule, quantify the overhead in the same currency as the SLO, price the cost of diagnosing incidents without data, and leave a documented default with a trigger to revisit it.

## What is actually being decided Two people want different things and both are right. The person carrying the pager wants the block and mutex profilers on, always, at a useful rate — because the moment a service serialises at 3am is the worst possible moment to discover that the profile is empty and the fix is a redeploy under pressure. The person who owns the latency or throughput SLO sees a permanent tax on the hot path with no incident to justify it, paid on every request forever. The decision is a posture, and it belongs to whoever owns the service's performance budget, with the on-call rotation's requirements as a hard input. It is not a per-engineer preference and it should not be re-litigated in each pull request. ## The cost model you must be able to state Both profilers are **event-driven**, not interval-driven. That is the fact the whole decision turns on. - Block profiling: when enabled, the runtime times blocking operations and captures a stack for each sampled event. The bill is charged per *blocking event*. - Mutex profiling: an event is recorded only when a lock release actually unblocks a waiter — that is, only on real contention. The bill is charged per *contention event*. So the overhead does not depend on uptime, heap size or core count; it depends on how concurrent and how chatty the program is. Two consequences follow: 1. **The same rate has wildly different costs in different services.** A batch job that computes for milliseconds between channel sends barely notices block profiling at a moderate rate. A pipeline doing a small handoff per item can pay a very visible percentage. 2. **The asymmetry between the two profilers is structural.** Contention events are rare in a healthy service, so mutex profiling at a coarse fraction is close to free — and it is the profile that names a bottleneck rather than merely describing waiting. Blocking events are common even in a healthy service, so block profiling is the expensive one. That asymmetry is why the honest default for most services is *mutex on, block off or very coarse*, rather than "both on" or "both off". ## How to decide, with evidence 1. **Measure, do not argue.** Benchmark the real hot path with the profilers off, then at candidate rates, on your own workload. A number from someone else's service is not evidence about yours. 2. **Express the cost in the same currency as the SLO.** "1.2% throughput at p99-neutral" is a sentence the SLO owner can accept or reject. "It is basically free" is not. 3. **Price the other side too.** The cost of *not* having the data is a repeat incident diagnosed by guesswork, plus the change-management time to enable profiling on a burning service. If your last three contention incidents each cost an hour of blind bisecting, that belongs in the comparison. 4. **Pick the coarsest rate that still ranks the waits correctly.** Sampling is weighted towards long waits, so coarse settings usually preserve the ordering you care about. Full-rate settings belong in a local reproduction, not in the fleet. ## The mechanism that dissolves the disagreement Most of the heat in this argument comes from treating it as binary and permanent. It is neither, because both rates are ordinary runtime calls that take effect immediately. - Put both behind an authenticated admin route so on-call can raise them on **one instance** during an incident and restore them afterwards — use the previous value returned by the mutex setter to restore cleanly. - If you want richer standing data, enable the expensive profiler on a small canary fraction of the fleet rather than everywhere. You get real production stacks and pay for them on a slice of traffic. - Make the current setting visible in a status endpoint. A profiler nobody can see the state of will be silently left at the wrong value for a year. With that in place, "always on at full rate" stops being the only way to guarantee data at 3am, and the SLO owner's objection stops being existential. ## Writing the decision down Whatever you land on, record three things next to the default: the measured overhead at that rate on this service, the reason for the value, and a trigger to revisit it — a major version upgrade, a significant change in the service's concurrency shape, or the next contention incident. Undocumented profiling settings drift: someone turns a rate up during an incident and never turns it back, or a new service copies a value chosen for a completely different workload. ## What separates a strong answer A strong answer refuses the binary. It states the event-driven cost model, distinguishes the two profilers on that basis rather than treating them as a pair, produces a number from a benchmark rather than an opinion, and then designs the disagreement away by making the setting adjustable at runtime, observable, canaried and documented — while being explicit about who owns the call and who can overrule it.

  • The SLO owner rejects any standing overhead at all. What do you propose?
    Keep both profilers off by default but ship the admin control and prove it works in a game day, so enabling costs seconds rather than a deploy. Then run the expensive profiler on a canary slice to keep some standing evidence, and bring the measured cost and the diagnosis time of past incidents back for review.
  • How do you stop a rate raised during an incident from staying raised for a year?
    Restore it explicitly using the previous value the mutex setter returns, expose the current setting in a status endpoint so drift is visible, and have the admin control apply to one instance that will be replaced on the next deploy rather than to a persisted configuration value.
  • Why not simply set both to full rate on one canary instance and be done?
    Because a canary at full rate stops being representative: its latency and throughput no longer match the fleet it is meant to represent, so it distorts both the profile and any canary-based release check. A coarse rate on the canary keeps the instance comparable while still yielding usable stacks.
  • What makes the mutex profiler the easier one to leave on permanently?
    It records only genuine contention events, which are rare in a healthy service, so the standing cost is small. It also attributes delay to the lock holder, so it names a bottleneck directly rather than producing a large body of waiting that has to be interpreted.

saying these in an interview costs you the question

  • Answers on instinct with no benchmark on the real workload
  • Treats the two profilers as a single on-or-off pair
  • Sets both to full rate fleet-wide for safety
  • Assumes overhead scales with uptime rather than blocking frequency
  • Requires a redeploy to change the setting during an incident
  • Leaves the chosen rate undocumented and unowned