skip to content

As platform owner, how do you set a fleet-wide GOMAXPROCS policy for Go services in containers?

level: principalimportance: nice to knowfreq 25%

answer

  1. three mechanisms, one owner
  2. implicit and self-correcting versus explicit and stale
  3. the go.mod line is the real prerequisite
  4. make the chosen value visible before changing it
  5. exceptions need a name and a metric

basics

~20 s

Make the runtime's container-aware default the fleet norm, enforce it with a minimum go.mod language version and a ban on code overrides, require a start-up log of the chosen value, and let a service override only with measurements and an owner.

solid answer

~50 s

There are three candidate mechanisms and they are not equal. Letting the runtime's container-aware default decide is the strongest default: one mechanism, no per-service configuration, and it tracks the CPU limit at run time including live resizes. Injecting `GOMAXPROCS` from the deployment template is explicit and auditable, but it freezes the value, goes stale the moment somebody resizes the workload, and quietly disables the runtime's own updates. Pinning in code is the worst — it is invisible to whoever operates the service. So I standardise on the runtime default and spend the policy budget on making it apply: a minimum `go` directive in `go.mod` (GODEBUG defaults follow it), a lint rule against `runtime.GOMAXPROCS` in libraries, and a mandatory start-up log of `NumCPU`, `GOMAXPROCS(0)` and the requested CPU limit. Overrides stay possible but become a decision with a name on it.

go deeper

for a junior

You are not expected to set fleet policy, but know that the value can come from the runtime, from an environment variable, or from code, and that the last of those is the hardest to see.

for a middle

Be able to compare the three mechanisms on their merits, especially that an injected environment variable is a frozen snapshot while the runtime default keeps tracking the real limit.

for a senior

Argue for a default with evidence: what you would measure across the fleet first, how you would sequence a rollout, and which services you would touch last.

for a principal

Own the policy end to end — the enforcement mechanism, the observability contract, the exception process and who carries the cost — and be ready to say what would make you reverse it.

## The decision, stated honestly Somebody has to answer this once for every Go service in the fleet: **who decides GOMAXPROCS?** The runtime, the deployment template, or the code. It is a real ownership question, because the platform team can mandate an answer and a service owner who wants a different one has to win that argument. ## Option A — the runtime's container-aware default On a recent toolchain the runtime derives the value from the smaller of the visible CPUs and the cgroup CPU limit, and re-reads the limit as the process runs. **For:** one mechanism for the whole fleet; no per-service configuration to drift; correct automatically when a limit changes, including a live resize; a new service is right on day one with nobody thinking about it. **Against:** it only applies where the module's `go` directive is recent enough, since GODEBUG defaults follow that line, so the policy is really a *version* policy in disguise. It is also implicit — nothing in the manifest says what the value will be, which frustrates anyone reading the deployment to predict behaviour. ## Option B — inject the environment variable from the deployment Compute `GOMAXPROCS` from the CPU limit in a shared template and set it on every pod. **For:** explicit and reviewable in the manifest; works uniformly across toolchain versions; easy to reason about and to audit centrally. **Against:** it is a snapshot. Resize the workload and the injected value is now wrong, silently, until somebody re-renders and redeploys. Worse, because it is an explicit setting it *disables* the runtime's own updating, so it does not merely duplicate option A — it replaces a self-correcting mechanism with one that can go stale. It also requires plumbing the limit into the pod's environment, which is one more template that must be right. ## Option C — pin it in application code **For:** essentially nothing, at fleet scale. **Against:** invisible to operators, unfixable without a code change and a release, and prone to the legacy `runtime.GOMAXPROCS(runtime.NumCPU())` idiom, which actively reinstates the wrong number inside a container. If the fleet has these, finding and deleting them is the highest-value part of the whole exercise. ## What I would mandate 1. **Default: the runtime decides.** No `GOMAXPROCS` environment variable, no code call. 2. **Enforce the prerequisite, not the intent.** A minimum `go` directive in `go.mod` across the fleet, checked in CI. Without it the policy is a wish. 3. **Ban the override in libraries outright**, and confine it in applications to the start-up path, enforced by a lint rule rather than a convention. 4. **Make the value observable.** Every service logs `runtime.NumCPU()`, `runtime.GOMAXPROCS(0)` and the CPU limit it requested, in one line, at start-up. This is the single cheapest thing on the list and it converts an incident-time investigation into a glance. 5. **Export throttling.** The cgroup's throttled-period ratio as a metric per service, so the fleet can see who is losing time to over-subscription without anybody filing a ticket first. ## The escape hatch, and who pays for it A blanket rule with no exceptions gets routed around. Two legitimate exceptions exist: - **Deliberate over-subscription.** A batch or asynchronous tier that would rather burst above its quota and absorb throttling than leave burst capacity unused. That is a defensible position for throughput-oriented, latency-insensitive work, and never for a request path. - **A latency-sensitive service that has measured a better value.** Fine — with the numbers, and with the owner accepting that they now maintain that value against every future resize. Both exceptions carry the same obligations: written justification, the value logged at start-up, and the throttling metric watched by the team that asked for the exception. The point of the policy is not that nobody may deviate; it is that a deviation is a decision somebody made, rather than a default nobody chose. ## Sequencing the rollout Do not flip everything at once. Ship the start-up log first and collect the current values across the fleet — that alone usually surfaces a handful of services running with an absurd P count. Then bump `go.mod` versions in waves, watching the throttling metric and tail latency of the services that change. Delete code overrides last, once you can see what deleting one does. The order matters because the first two steps are observation, and only the third changes behaviour.

  • A service owner insists on injecting GOMAXPROCS from the CPU limit for auditability. What is your counter?
    That auditability is better served by logging the value at start-up than by freezing it in a manifest. An injected value is an explicit setting, so it also switches off the runtime's tracking of the real limit — you gain a line someone can read and lose correctness after the next resize.
  • How do you roll this out without breaking the workloads that were quietly relying on over-subscription?
    Observe before changing: ship the start-up log, collect current values fleet-wide, and identify the services whose effective parallelism will drop. Change those last, in waves, with their owners watching tail latency and throughput. A service that genuinely wants over-subscription then documents it as an exception rather than losing it by surprise.
  • What would make you reverse the policy and pin values centrally instead?
    A fleet that cannot move to a recent language version, or an environment where the cgroup limit is not a truthful statement of what the workload may use — a deliberately over-committed node pool, say. Then the runtime's input is wrong, and an explicit number derived from capacity planning beats a correct reading of a misleading limit.

saying these in an interview costs you the question

  • Mandates a single fixed GOMAXPROCS number for every service
  • Treats injecting the environment variable as equivalent to the runtime default
  • Ignores that the go.mod language version gates the behaviour
  • Changes fleet defaults before making the current values observable
  • Allows exceptions with no owner, no measurement and no metric