What does GOMAXPROCS control, and how does a container's CPU limit affect its default?
answer
- how much Go runs at the same instant
- one per scheduling context, not per goroutine
- in a container, CPU count misleads
- Go 1.25 also reads the cgroup CPU limit
basics
~20 sGOMAXPROCS caps how many goroutines run Go code simultaneously. Since Go 1.25 on Linux its default is the lower of the machine's usable logical CPUs and the container's cgroup CPU limit; older releases looked only at the CPU count.
solid answer
~40 sGOMAXPROCS is the number of scheduling contexts (Ps) the Go runtime keeps, so it bounds how many goroutines can be executing Go code at the same instant. It is not a cap on how many goroutines you may create, and it is not a cap on OS threads. Historically the default was `runtime.NumCPU()` — the logical CPUs the process can see, honouring its CPU affinity mask but not the cgroup CPU bandwidth quota. That is why a container limited to half a CPU on a 64-CPU node used to start with 64 Ps and get throttled hard. Since Go 1.25 on Linux the default also takes the cgroup CPU limit into account and uses the smaller value. You can still override it with the `GOMAXPROCS` environment variable or `runtime.GOMAXPROCS(n)`.
code
go · 3 linesfmt.Println("logical CPUs visible:", runtime.NumCPU())
// passing 0 queries the setting without changing it
fmt.Println("GOMAXPROCS:", runtime.GOMAXPROCS(0))go deeper
Be ready to say in one sentence what GOMAXPROCS bounds — simultaneous execution of Go code — and to name the two functions that show you the numbers: runtime.NumCPU and runtime.GOMAXPROCS(0).
Explain why a visible CPU count and a cgroup CPU quota are different things, and what the runtime does with each when it picks the default on a modern toolchain.
Show that you check this on real deployments: read the chosen value from inside the container, compare it with the CPU limit the manifest asked for, and treat a mismatch as a first-class performance bug.
Own the position that the runtime's own default, not a hand-set number, should be the fleet norm, and be able to justify why a hard-coded value is a liability the day someone resizes a service.
## What the setting actually is The Go runtime multiplexes goroutines onto OS threads through an intermediate resource the runtime calls a **P** — a scheduling context that owns a run queue and the per-CPU caches a thread needs in order to run Go code. A thread must hold a P to execute Go code, and `GOMAXPROCS` is simply how many Ps exist. So the honest one-line definition is: **GOMAXPROCS is the maximum number of goroutines that can be executing Go code simultaneously.** Three things it is *not*: - **Not a limit on goroutines.** You can create a million goroutines with `GOMAXPROCS=1`; they just take turns. - **Not a limit on OS threads.** A goroutine that blocks in a system call or in cgo hands its P to another thread, so the process can hold many more threads than Ps. - **Not a limit the kernel enforces.** The container's CPU quota is enforced by the kernel; GOMAXPROCS is only the runtime's own idea of how much parallelism to attempt. ## Where the default comes from For most of Go's history the default was `runtime.NumCPU()`. That function reports the number of logical CPUs **usable by the process**, which means it respects a CPU affinity mask (a `taskset`-style pinning) — but it has never respected a cgroup CPU *bandwidth* limit, because a bandwidth quota does not remove CPUs from view. A process in a container that is allowed 0.5 CPU on a 64-core node still sees 64 CPUs. That mismatch is the classic container problem. The runtime starts 64 Ps, happily runs 64 goroutines in parallel, burns the cgroup's slice of the quota period in a fraction of the period, and then the whole cgroup is frozen by the kernel until the next period begins. The program is not slow because a function is slow; it is slow because it asked for far more parallelism than it was allowed to consume. **Go 1.25 changed the default on Linux.** The runtime now derives GOMAXPROCS from the smaller of the logical CPUs it can use and the cgroup CPU bandwidth limit, rounding a fractional limit up to a whole number, and it does not reduce the value below 2 on account of the limit. On a node with 64 CPUs and a 0.5-CPU quota, a Go 1.25 program therefore starts small instead of starting at 64. ## Seeing it from inside the container The two numbers you want are `runtime.NumCPU()` (what the process can see) and `runtime.GOMAXPROCS(0)` (what the runtime chose — passing 0 queries without setting). Logging both at start-up, next to the CPU limit the deployment asked for, turns an argument into a fact. If the CPU count is large, GOMAXPROCS equals it, and the CPU limit is a fraction of a core, you have found a real problem. ## Overriding it The value is still yours to set: - the `GOMAXPROCS` environment variable, read once at start-up; - `runtime.GOMAXPROCS(n)` with `n >= 1`, which sets the value and returns the previous one. Either of those pins the number. Prefer leaving it alone on a modern toolchain: the runtime's own default tracks the limit the platform actually gave you, and a hard-coded number goes stale the moment somebody edits the deployment. ## The everyday intuition For CPU-bound work, throughput rises with GOMAXPROCS only up to the CPU time you are actually allowed to consume. Beyond that you are not adding capacity — you are adding contention, extra parallel garbage-collector workers (the collector targets roughly a quarter of GOMAXPROCS in dedicated mark workers), and more per-P memory, all paid for out of the same quota. Matching the runtime to the limit is nearly always faster than exceeding it.
- Does GOMAXPROCS limit how many OS threads the process creates?No. It fixes the number of Ps, which bounds simultaneous execution of Go code. A goroutine that blocks in a syscall or in cgo releases its P, and the runtime may start another thread to keep the P busy, so the thread count can be far higher than GOMAXPROCS.
- Before Go 1.25, did the default respect a CPU affinity mask?Yes. `runtime.NumCPU()` reports the logical CPUs usable by the process, so pinning it to a subset of cores with an affinity mask did lower the default. A cgroup CPU *bandwidth* quota does not remove CPUs from view, so it was ignored — which is exactly the gap Go 1.25 closed.
- If a container is limited to 0.5 CPU, why does the Go 1.25 default not choose 1?The runtime rounds a fractional limit up and does not go below 2 on the strength of the cgroup limit, so a sub-core quota still yields a small but non-degenerate value. One P would serialise everything, including the collector's work, and hurt latency more than the mild over-subscription does.
A CPU quota is a monthly data allowance; GOMAXPROCS is how many devices you let stream at once. Ten devices do not give you more data — they just exhaust the allowance sooner and leave you cut off.
saying these in an interview costs you the question
- Says GOMAXPROCS caps how many goroutines you can create
- Thinks GOMAXPROCS is a hard limit on OS threads
- Assumes the default always equals the host's CPU count
- Believes raising GOMAXPROCS above the CPU quota adds real parallelism
- Confuses GOMAXPROCS with the kernel-enforced CPU limit itself