A Go sidecar with GOMAXPROCS=64 runs under a half-core quota and stalls in bursts. How do you confirm the cause?
answer
- two numbers from inside the container
- compare the P count to the grant
- the kernel counts its own freezes
- stalls line up with periods, not requests
- idle during the stall, not busy
basics
~20 sCompare the runtime's P count with the CPU quota: read runtime.GOMAXPROCS(0) and runtime.NumCPU() from inside the container and check the cgroup's throttle counters. Stalls that align with quota periods, not with a hot function, mean the runtime is over-subscribing its allowance.
solid answer
~50 sStart from inside the container rather than from the dashboard. Log `runtime.GOMAXPROCS(0)` and `runtime.NumCPU()` at start-up and compare both against the CPU limit the deployment actually asked for — 64 Ps against half a core is the whole story. Then confirm the mechanism on the kernel side: the cgroup's `cpu.stat` exposes `nr_periods`, `nr_throttled` and `throttled_usec`, and if `nr_throttled` is a large fraction of `nr_periods` and climbs in step with the latency spikes, the process is being frozen for the tail of most periods. The signature to look for is stalls that follow the quota period rather than a particular request or code path: the sidecar's hashing work fans out across 64 Ps, spends the whole allowance in a few milliseconds, and then does nothing until the period rolls over. The fix is to match the runtime's parallelism to the grant: a toolchain and `go.mod` version whose default is container-aware, with no code override reinstating the CPU count.
code
text · 4 lines$ cat /sys/fs/cgroup/cpu.stat
nr_periods 41822
nr_throttled 41630
throttled_usec 2988421000go deeper
Know that a container has a CPU allowance the Go runtime does not automatically respect on older versions, and that the first thing to print is what the runtime chose versus what the container was given.
Explain the mechanism end to end: many Ps drain a per-period CPU budget early, the kernel freezes the group, and the stall shows up as tail latency rather than as high average CPU.
Drive the diagnosis: compare the P count against the grant from inside the container, correlate the throttled-period ratio with the latency spikes, rule out hot code and contention, then fix the runtime sizing before touching the limit.
Decide whether this class of incident is prevented by policy — a fleet-wide default, a start-up log contract, a throttling metric every service exports — rather than rediscovered service by service.
## The setup A sidecar runs beside every pod in the fleet. It hashes fixed-size chunks of data — CPU-bound, embarrassingly parallel work. Its CPU limit is deliberately small, half a core, because it should be cheap. The nodes have 64 CPUs. The symptom is not a high average: it is bursts. Throughput is fine when measured over a minute, and the p99 is dreadful. ## Step 1: what does the runtime think it has? The first two numbers cost nothing and settle most of the question: - `runtime.NumCPU()` — the logical CPUs the process can see. Here, 64. - `runtime.GOMAXPROCS(0)` — the Ps the runtime chose. Here, also 64. Against a 0.5-CPU grant, that is a 128-fold mismatch between what the runtime will attempt and what the kernel will pay for. Log both at start-up in every service; a value you have to go and measure during an incident is a value nobody checks. If GOMAXPROCS is 64 on a modern toolchain, ask why the container-aware default did not apply. The three usual answers: the `go` directive in `go.mod` still names a pre-1.25 version, so the old default is in force; the code calls `runtime.GOMAXPROCS(runtime.NumCPU())` somewhere in start-up; or the deployment injects a `GOMAXPROCS` environment variable computed from something other than the limit. ## Step 2: confirm the kernel is throttling A large P count is a strong hypothesis, not proof. The cgroup keeps the evidence itself. Under cgroup v2, `cpu.stat` inside the container reports `nr_periods` (how many quota periods have elapsed), `nr_throttled` (how many of those ended with the group frozen) and `throttled_usec` (how long it spent frozen in total). What you are looking for is a ratio, not an absolute: `nr_throttled` close to `nr_periods` means the group is running out of allowance in nearly every period. Sampling those counters over a minute and dividing gives you throttled time per second of wall clock — a number you can put next to the latency spikes and show that they move together. ## Step 3: rule out the other explanations The reason this diagnosis is worth practising is that throttling *looks like* other things: - **Slow code.** A genuine hot spot degrades latency smoothly with load; throttling produces a cliff aligned with the quota period, and the process is idle — not busy — during the stall. - **Lock contention.** Contention shows up as work in progress that is not advancing; throttling shows up as no work in progress at all, with the CPU accounting frozen. - **A dependency being slow.** Throttling stalls hit CPU-bound stretches, including the stretch between an upstream response arriving and it being processed, which makes downstreams look slow when they are not. The distinguishing observation is always the same: the process is not running, and the reason it is not running is that its cgroup has no CPU allowance left in this period. ## Why a large GOMAXPROCS makes it worse rather than better With N Ps, N goroutines run in parallel and the quota drains roughly N times faster. Halving the period's useful work into the first few milliseconds and freezing for the remainder is worse for latency than spreading the same work across the whole period, even though the total CPU consumed is identical. Two secondary costs pile on: - The garbage collector sizes its dedicated mark workers as a fraction of GOMAXPROCS (roughly a quarter), so an inflated value means a large slice of an already tiny allowance goes to collection whenever a cycle runs. - Each P carries per-P runtime state and caches, so an inflated value costs memory in a container that is usually memory-limited too. ## The fix, in order of preference 1. Let the runtime do it — a toolchain and a `go.mod` `go` directive recent enough for the container-aware default, and no code or environment override fighting it. 2. If something must pin the value, derive it from the container's CPU limit rather than from `runtime.NumCPU()`, and log it. 3. Only then consider changing the limit itself — but that is a capacity conversation, and it should follow the runtime being sized correctly, not substitute for it. Verify the same way you diagnosed: after the change, GOMAXPROCS should be a small number, and the throttled-period ratio should fall towards zero with the tail latency.
- Why does raising GOMAXPROCS make tail latency worse here rather than better?The quota is a budget per period, not a rate limit per goroutine. More Ps spend the same budget sooner, so useful work is compressed into the head of each period and the group is frozen for the rest. Total CPU is unchanged; the distribution of stalls is much worse.
- Would you expect the same symptom if the sidecar were I/O-bound rather than hashing?Much less. A goroutine blocked in a system call is not consuming CPU time, so an oversized GOMAXPROCS costs little beyond memory and scheduling overhead. Quota exhaustion needs sustained on-CPU work, which is exactly what fixed-size chunk hashing provides.
- GOMAXPROCS is already correct and the container still throttles. What does that tell you?That the workload genuinely wants more CPU than it was granted, which is a capacity question rather than a runtime-tuning one. Sizing the runtime to the limit removes self-inflicted throttling; it cannot manufacture allowance that was never given.
- What do you add to the service so this is obvious next time?A start-up log line carrying runtime.NumCPU, runtime.GOMAXPROCS(0) and the CPU limit the deployment requested, plus a periodically sampled throttled-period ratio from the cgroup counters exported as a metric. Both are cheap and turn a multi-hour investigation into a glance.
saying these in an interview costs you the question
- Blames slow code without checking whether the process was even running
- Assumes high latency with low CPU usage rules out a CPU problem
- Raises GOMAXPROCS to make a throttled service faster
- Reads CPU counts from the node instead of from inside the container
- Treats a single throttled period as proof of a problem instead of the ratio