skip to content

A containerized Go service reports a much lower GOMAXPROCS after only a toolchain upgrade. What changed, and how do you confirm it?

level: seniorimportance: should knowfreq 38%

answer

  1. same source, different toolchain
  2. the host's cores versus the container's quota
  3. one release taught the runtime about cgroups
  4. print both numbers on the first log line
  5. a GODEBUG restores the older default

basics

~20 s

Since Go 1.25 on Linux the default GOMAXPROCS is the lower of the usable CPU count and the container's cgroup CPU limit, so a rebuild on a newer toolchain gets far fewer Ps. Confirm by logging runtime.GOMAXPROCS(0) beside runtime.NumCPU().

solid answer

~50 s

This is a version skew, not a code change. Before Go 1.25 the default `GOMAXPROCS` came from the CPUs the process could see, so a container with a 2-CPU quota on a 64-core node started 64 Ps. Since **Go 1.25** on Linux the runtime also reads the cgroup CPU bandwidth limit and defaults to the lower of the two, and it re-reads that periodically. Confirm it in three cheap steps: log `runtime.GOMAXPROCS(0)` and `runtime.NumCPU()` at startup — in a quota'd container they now disagree; read the container's CPU limit and check it matches the new value; and A/B the old behaviour with `GODEBUG=containermaxprocs=0`, which makes the runtime ignore the cgroup limit. Then decide deliberately: keep the derived default, or pin the value with the environment variable — which also opts out of the periodic updating.

code

go · 3 lines
go
log.Printf("GOMAXPROCS=%d NumCPU=%d", runtime.GOMAXPROCS(0), runtime.NumCPU())
// 64-core host, 2-CPU cgroup limit, Go 1.25+: GOMAXPROCS=2 NumCPU=64
// same host and container, older toolchain:  GOMAXPROCS=64 NumCPU=64

go deeper

for a junior

Know that GOMAXPROCS has a default the runtime derives from its environment rather than a fixed number, and that runtime.GOMAXPROCS(0) shows you what you actually got.

for a middle

Be able to state the derivation: usable CPUs, and on recent Go on Linux the container's CPU limit, whichever is lower — and to explain why the host's core count is the wrong input inside a container.

for a senior

Demonstrate the investigation: log the two numbers, compare against the container's quota, A/B with a GODEBUG rather than a rebuild, and rule out an explicit setting before theorising about the runtime.

for a principal

Own the fleet consequence: a toolchain bump silently changes the CPU-parallelism budget of every containerized service, so the rollout needs a canary plan and a way to see the value on every process.

## The change that explains it Up to and including Go 1.24, the default `GOMAXPROCS` was the number of logical CPUs usable by the process. On Linux that already respected the process's CPU affinity mask, but it knew nothing about **cgroup CPU bandwidth limits** — the mechanism a container runtime uses to say "this process may use two CPUs' worth of time". So a service in a 2-CPU container on a 64-core node started with 64 Ps: 64 Go-code execution slots against an entitlement of two CPUs. **Go 1.25** changed the default on Linux in two ways: 1. The runtime considers the CPU bandwidth limit of the cgroup containing the process. If that limit is lower than the usable CPU count, the limit wins. 2. The runtime periodically re-reads both inputs and updates `GOMAXPROCS` if they change while the program is running. Both behaviours apply only when the value was *not* set explicitly. Setting the `GOMAXPROCS` environment variable, or calling `runtime.GOMAXPROCS(n)`, pins the value and turns off the automatic updating; `runtime.SetDefaultGOMAXPROCS()` hands control back to the runtime. So: same source, same image layout, new toolchain, and the P count collapses from the node's core count to roughly the quota. That is the whole mystery. ## Confirming it, cheaply and in order **1. Make the two numbers visible.** Log `runtime.GOMAXPROCS(0)` next to `runtime.NumCPU()` on the first line the process writes. `NumCPU` reports CPUs the process may use; `GOMAXPROCS` reports Ps. Before the change they matched in a container; now they diverge, and the divergence *is* the diagnosis. **2. Check the quota.** Read the container's CPU bandwidth limit (on cgroup v2 that is the `cpu.max` entry for the process's cgroup). If the new `GOMAXPROCS` tracks that limit rather than the node's core count, you have your answer. **3. A/B without a new build.** Run one instance with `GODEBUG=containermaxprocs=0`, which tells the runtime to ignore the cgroup limit and default the old way. If behaviour reverts, the cause is settled. `GODEBUG=updatemaxprocs=0` separately disables the periodic re-reading, which is worth trying if the value appears to change mid-life. **4. Rule out an explicit setting.** Grep the deployment for a `GOMAXPROCS` environment variable and the code for `runtime.GOMAXPROCS(` calls. Either one overrides everything above and also disables updating, and a leftover from a previous incident is a common find. ## What actually changes in behaviour Fewer Ps is usually the *correct* configuration and often an improvement: with 64 Ps against a two-CPU entitlement, the runtime keeps trying to run 64 goroutines' worth of Go code, the kernel throttles the cgroup when the quota is exhausted, and the result is bursty scheduling and ugly tail latency. Matching the P count to the entitlement reduces that. But it is a real behaviour change and some things genuinely regress: - **Anything sized from the old number.** Code that captured `runtime.GOMAXPROCS(0)` once at init to size sharded counters or per-slot buffers now sizes differently — and, because the runtime may update the value later, a captured number can also go stale. - **Code that assumed `NumCPU() == GOMAXPROCS()`.** That assumption is now false in every CPU-limited container. - **Workloads that relied on wide oversubscription** to overlap latency get less overlap and may show lower throughput even as latency improves. - **Local development diverges from production**, because the container-aware behaviour is a Linux-and-cgroup story; on a developer laptop the default is still the machine's CPU count. ## Then make a decision, not a reflex The two honest outcomes are: accept the derived default because it matches what the platform actually grants the process, or pin `GOMAXPROCS` explicitly because you have a specific reason — and record that reason, because pinning also means the runtime will no longer follow a later change to the container's CPU limit. What you should not do is leave `GODEBUG=containermaxprocs=0` in place permanently: it is a rollback lever for an investigation, not a configuration.

  • Which GODEBUG settings let you test the old behaviour without shipping a different build?
    `GODEBUG=containermaxprocs=0` makes the runtime ignore the cgroup CPU limit, restoring the pre-1.25 derivation from the usable CPU count. `GODEBUG=updatemaxprocs=0` separately stops the runtime re-reading its inputs and revising GOMAXPROCS while the process runs. Both are investigation levers; neither belongs in a permanent deployment.
  • What in application code can silently break when GOMAXPROCS changes?
    Anything that read `runtime.GOMAXPROCS(0)` once at init and sized a structure from it — sharded counters, per-slot buffers, stripe arrays. Those keep the old width, and on Go 1.25+ the value they captured can also become stale if the runtime later revises it. Code asserting `runtime.NumCPU() == runtime.GOMAXPROCS(0)` breaks outright in a CPU-limited container.
  • Does setting the GOMAXPROCS environment variable turn anything else off?
    Yes. An explicit value — environment variable or a `runtime.GOMAXPROCS(n)` call — opts the process out of the runtime's periodic re-derivation, so a later change to the container's CPU limit is no longer picked up. `runtime.SetDefaultGOMAXPROCS()` returns the process to the runtime-derived default and its updating.
  • Why might latency improve while throughput on one stage falls after the change?
    A P count far above the CPU entitlement oversubscribes the quota: the runtime keeps more Go code in flight than the cgroup will pay for, so work is throttled in bursts and tail latency suffers. Sizing the P count to the entitlement smooths that, but a stage that leaned on wide oversubscription to overlap waiting loses some of that overlap.

saying these in an interview costs you the question

  • Blames application code when only the toolchain changed
  • Assumes runtime.NumCPU() reports the container's CPU quota
  • Thinks the cgroup limit was always honoured by default
  • Believes GOMAXPROCS cannot change during a process's life
  • Leaves a GODEBUG opt-out in place as permanent configuration