skip to content

Should a fleet pin GOMAXPROCS per service or rely on the container-aware default? How do you decide?

level: principalimportance: nice to knowfreq 30%

answer

  1. one rule inherited beats many pinned numbers
  2. the default derives from limits, not requests
  3. pins opt out of automatic updating
  4. every pin needs an owner and a trigger
  5. make the value observable per process

basics

~20 s

Default to the runtime-derived value fleet-wide and treat pins as documented exceptions. The default is only trustworthy if every service builds on a toolchain that reads the container CPU limit and every workload actually has such a limit; otherwise pin.

solid answer

~50 s

Make the runtime-derived default the fleet policy and pins the exception, because one rule everybody inherits beats dozens of numbers nobody revisits. But the default is only correct under two preconditions you have to own: every service builds on a Go version whose default reads the cgroup CPU limit, and every workload actually carries such a limit — with generous limits or none at all, the runtime falls back to the host's CPU count and a small container on a big node gets a wildly oversubscribed P count. Where those preconditions fail, pinning to the entitlement the platform really grants is the honest choice. Any pin needs a named owner, a written reason, and a trigger that forces a re-look when the container's CPU allocation changes — and it must be recorded that pinning also opts the process out of the runtime's automatic updates.

go deeper

for a junior

You are not expected to set this policy, but know that the value your service runs with comes either from the deployment manifest or from the runtime's own derivation, and that you can print it at startup.

for a middle

Be able to explain why a container-aware default is usually better than a hand-picked number, and what the runtime reads to derive it.

for a senior

Show that you would verify the preconditions before trusting the default — toolchain floor, whether workloads actually carry CPU limits — and that you would make the value observable before changing it anywhere.

for a principal

Own the posture and its governance: one fleet-wide rule, pins as exceptions with owner and expiry, a migration plan with canaries, and a willingness to say the real defect is the platform's CPU limit policy when the derived number is wrong.

## What the decision actually is `GOMAXPROCS` is the number of Go-code execution slots a process gives itself. Choosing between a pinned value and the runtime-derived default is therefore a **capacity posture for every service on the platform**, not a per-service tuning trick. The platform team that owns the base images and deployment defaults owns this call, and a service owner with a measured argument can overrule it for their service — that is the shape of the decision, and stating it that way is most of the answer. ## The case for the derived default The runtime derives its default from the environment it finds. On recent Go on Linux that means the lower of the CPUs the process may use and the container's cgroup CPU bandwidth limit, re-read periodically. That is exactly the number you would compute by hand from the deployment manifest, and it follows the manifest automatically when someone resizes the container. One rule, inherited by every service, with nothing to keep in sync — which is why it should be the default posture. ## The preconditions that make it trustworthy The default is only as good as its inputs, and there are three ways it silently is not: **1. Mixed toolchains.** The container-aware default arrived in a specific Go release. A fleet spanning older and newer toolchains has two different behaviours for the same manifest, which is the worst of both worlds. The right first move is to raise the toolchain floor in the base image and the module's `go` directive, not to sprinkle environment variables over the services that behave oddly. **2. Limits versus requests.** The runtime reads the CPU *bandwidth limit*. If your platform sets tight requests and very generous limits — or no limits at all — the derived number reflects an entitlement your scheduler never actually intends to give the pod. With no limit, the fallback is the host's CPU count, so a small container on a 96-core node starts with 96 Ps. That is precisely the case where pinning near the request is defensible, and it is worth noticing that the real fix may be to your limit policy rather than to Go. **3. Not everywhere is Linux under a cgroup.** Developer machines and some environments have no such limit, so local behaviour will differ from production. Do not add a pin whose only purpose is to make a laptop match a container. ## The cost of pinning A pin is a number in a manifest, and numbers in manifests rot. Specifically: - It **opts out of the runtime's automatic updating**, so a later change to the container's CPU allocation is no longer followed. - It becomes stale the moment somebody resizes the workload, and nothing forces the two to be changed together. - Multiplied across a fleet, pins become a collection of unexplained constants that the next team dares not touch. So the governance matters more than the value: every pin gets an owner, a one-line reason tied to a measurement, and a review trigger — any change to the container's CPU limit. Set it next to the CPU limit in the deployment manifest, never baked into an image, so the two are visible together. ## Making the posture observable Whatever you choose, make the value auditable. Have every service log `runtime.GOMAXPROCS(0)` alongside `runtime.NumCPU()` at startup, and ideally expose it as a metric. Then "what P count is this fleet actually running with?" is a query rather than an archaeology project, and a toolchain bump that changes the answer shows up immediately instead of as a mysterious latency shift a week later. ## Rolling it out Treat the move to the derived default as a migration: raise the toolchain floor, add the startup log everywhere, canary one service per tier and watch CPU throttling and tail latency, and keep the runtime's opt-out `GODEBUG` as a per-service rollback lever with an expiry date rather than a fleet-wide setting. A permanent opt-out is a decision to keep the old behaviour forever without ever saying so. ## The answer that lands "Default everywhere, pins as documented exceptions, preconditions owned by the platform, value observable on every process." Then name the one condition that flips it: if the platform's CPU limits do not represent what a workload actually gets, the derived default is deriving from the wrong number, and pinning is the honest response until the limit policy is fixed.

  • What makes a per-service pin acceptable rather than a smell?
    Three things: a named owner, a written reason tied to a measurement rather than a hunch, and a trigger that forces a re-look — any change to the workload's CPU allocation. Put it in the deployment manifest next to the CPU limit rather than in the image, so a person resizing the container sees both numbers at once.
  • Your platform sets CPU requests but no CPU limits. What does the derived default give you?
    With no CPU bandwidth limit there is nothing to cap the derivation, so the runtime falls back to the CPUs the process may use — on a large node, the node's core count. A workload entitled to a fraction of that starts with a badly oversubscribed P count. That is the strongest case for pinning, and also a signal that the limit policy itself needs attention.
  • How would you roll this out across dozens of services without a fleet-wide incident?
    Raise the toolchain floor in the base image first so every service behaves the same way, ship the startup log or metric everywhere so the value is visible, then canary one service per workload tier and watch CPU throttling and tail latency before proceeding. Keep the runtime's opt-out GODEBUG as a per-service rollback with an expiry date, never as a default.
  • A service owner wants their pin kept after the migration. What do you ask for?
    The measurement that motivated it, on the current toolchain rather than the old one, and the condition under which it would be removed. Many pins predate the container-aware default and exist only to work around it; those should go. A pin that survives that conversation is a legitimate exception and should be recorded as one.

saying these in an interview costs you the question

  • Picks one global GOMAXPROCS number for every service
  • Pins the value without recording an owner or a reason
  • Assumes the default derives from CPU requests, not limits
  • Keeps a GODEBUG opt-out as permanent fleet configuration
  • Treats it purely as a per-service knob with no fleet policy
  • Ignores that a mixed toolchain fleet has two behaviours