skip to content

What do the arguments to runtime.SetBlockProfileRate and SetMutexProfileFraction mean?

level: middleimportance: nice to knowfreq 22%

answer

  1. the two arguments are different units
  2. one is time, one is a count
  3. one of them has a getter
  4. negative means different things to each
  5. nanoseconds blocked versus one event in N

basics

~20 s

They are not the same unit. runtime.SetBlockProfileRate takes nanoseconds: it aims for one sampled event per that many nanoseconds spent blocked, and 1 samples everything. runtime.SetMutexProfileFraction takes a fraction: on average one contention event in N is reported.

solid answer

~50 s

`runtime.SetBlockProfileRate(rate)` is time-based: the profiler aims to sample one blocking event per `rate` nanoseconds of blocked time, so `1` records essentially every event, `10000` samples roughly one per 10µs blocked, and anything at or below zero turns it off. Long blocks are always sampled and short ones are sampled probabilistically, with the recorded delay scaled up so the totals stay a fair estimate. `runtime.SetMutexProfileFraction(n)` is event-based: on average one contention event in `n` is reported, `1` reports all of them, and `0` turns it off. It also returns the previous value, and a **negative** argument means "tell me the current setting without changing it" — where a negative argument to the block setter would simply disable profiling. Mixing the two units up is how people end up with a rate a thousand times more expensive than they intended.

code

go · 9 lines
go
// time-based: aim for one sample per 100µs spent blocked
runtime.SetBlockProfileRate(100000)

// event-based: report about one contention event in ten
runtime.SetMutexProfileFraction(10)

// a negative argument reads the mutex fraction without changing it
cur := runtime.SetMutexProfileFraction(-1)
_ = cur

go deeper

for a junior

Know that both take an int but mean different things: one is nanoseconds of blocked time per sample, the other is one contention event in N. Passing zero or less to the block setter turns it off.

for a middle

Explain the sampling weighting: long blocks are always captured, short ones probabilistically and then scaled, so aggregate delay stays a fair estimate while the raw event count becomes approximate.

for a senior

Show how you would pick a value — benchmark the hot path at a candidate rate, take the coarsest setting that still ranks the waits correctly, and use the mutex setter's return value to save and restore around a diagnostic window.

for a principal

Standardise the values across services rather than leaving each team to guess, and require the measured overhead behind whatever number you make the default.

## Same shape, different units The two switches look symmetric and are not. Getting them the wrong way round is the classic mistake, because both take a bare `int` and neither will complain. ### runtime.SetBlockProfileRate: nanoseconds of blocked time ```go runtime.SetBlockProfileRate(rate int) ``` The documented contract is that the profiler *aims to sample an average of one blocking event per `rate` nanoseconds spent blocked*. So the argument is a **time budget between samples**, not a divisor of events. How that plays out: - `rate = 1` records essentially every blocking event. Maximum fidelity, maximum cost. - `rate = 10000` targets one sample per 10 microseconds of blocked time across the program. - `rate = 1000000` targets one sample per millisecond blocked — coarse and cheap. - `rate <= 0` disables block profiling entirely. The sampling is biased towards long waits on purpose: an event that blocked for longer than the rate is always recorded, while shorter events are recorded with a probability proportional to how long they blocked. The delay of a sampled short event is then scaled up, so aggregate totals remain a reasonable estimate of the real blocked time even though only a fraction of events were captured. This means a coarse rate does not hide a genuinely slow wait; it hides the very short ones, which are exactly the ones you can afford to lose statistically — though it also means the *contentions count* at a coarse rate is an estimate, not a tally. ### runtime.SetMutexProfileFraction: one event in N ```go runtime.SetMutexProfileFraction(rate int) int ``` Here the argument is a plain fraction of **events**: on average one contention event in `rate` is reported. There is no time dimension at all. - `rate = 1` reports every contention event. - `rate = 100` reports about one in a hundred. - `rate = 0` turns mutex profiling off. - `rate < 0` **leaves the setting unchanged** and simply returns the current value. The return value is the previous fraction, which makes the negative case a read and gives you a clean save-and-restore idiom for a temporary measurement: ```go prev := runtime.SetMutexProfileFraction(5) // turn it up, keep the old value defer runtime.SetMutexProfileFraction(prev) ``` ### The asymmetry, stated plainly | | block | mutex | |---|---|---| | argument means | nanoseconds blocked per sample | one event sampled in N | | every event | `1` | `1` | | off | `<= 0` | `0` | | read current value | not available | negative argument | | returns | nothing | previous fraction | Two consequences follow directly. First, `SetMutexProfileFraction(1000)` is a *cheap* setting while `SetBlockProfileRate(1000)` is a *fairly expensive* one — one microsecond between samples is fine-grained. Second, a habit of "pass -1 to read it" works for mutex and silently disables block profiling. ### Choosing values The cost of both profilers scales with how often the underlying event happens, not with wall-clock time, so the right value depends entirely on the workload. A useful method: 1. Benchmark the hot path with the profiler off to get a baseline. 2. Re-run with a candidate rate and measure the throughput difference on your own workload — not on someone's blog post. 3. Pick the coarsest rate that still shows the stacks you need, and record why. Coarse rates preserve the *ranking* of waits far better than people expect, because the sampling is weighted by time blocked. You almost never need `1` outside a local reproduction. ### Reading back what is set There is no getter for the block rate, so if you expose these through an admin endpoint you have to remember what you set. For the mutex fraction, `runtime.SetMutexProfileFraction(-1)` is the getter, and it is worth using in a status handler so on-call can see the current posture without guessing.

  • At a coarse block profile rate, is a long wait at risk of being missed?
    No. Events that blocked for longer than the rate are always sampled; shorter ones are sampled with probability proportional to their duration and then scaled up. So coarse rates lose fine-grained short waits and make the contentions count an estimate, while preserving the ranking of the slow ones.
  • How would you temporarily raise the mutex profile fraction and restore it afterwards?
    Use the return value: `prev := runtime.SetMutexProfileFraction(5)` then `defer runtime.SetMutexProfileFraction(prev)`. The setter returns the previous fraction, so a save-and-restore pair keeps a diagnostic window from permanently changing the service's profiling posture.
  • Why is there no equivalent getter for the block profile rate?
    The block setter returns nothing and treats any non-positive argument as "off", so there is no spare value left to mean "read only". If a service needs to report its current block rate, it has to remember the value it set — typically in the same config that an admin endpoint writes.

saying these in an interview costs you the question

  • Reads SetBlockProfileRate's argument as one event in N
  • Reads SetMutexProfileFraction's argument as nanoseconds
  • Passes -1 to SetBlockProfileRate expecting to read it
  • Assumes a coarse block rate drops long waits entirely
  • Ignores the previous value returned by SetMutexProfileFraction