skip to content

In Go's scheduler, what does a P own, and why must an M acquire one to run Go code?

level: middleimportance: should knowfreq 42%

answer

  1. a permit plus a private toolbox
  2. two counts that should be allowed to differ
  3. single owner means no lock needed
  4. the allocator has a per-P cache too

basics

~20 s

A P owns the state needed to run goroutines: a queue of ready goroutines and per-P caches such as the memory-allocation cache. Requiring an M to hold a P keeps Go-code parallelism at exactly GOMAXPROCS while the thread count varies freely.

solid answer

~40 s

A P is a scheduling context, not a thread and not a core. It owns a local run queue of ready goroutines, a per-P memory-allocation cache so most small allocations need no global lock, and assorted per-P bookkeeping the runtime wants to touch without synchronising. Requiring an M to hold a P before executing Go code buys two things. First, it fixes the parallelism budget: there are exactly `GOMAXPROCS` Ps, so at most that many goroutines run Go code at any instant, no matter how many threads exist. Second, it gives whatever thread is currently running Go code a private working set — queue and caches are owned, not shared, so the hot path is lock-free. Decoupling the permit (P) from the worker (M) is what lets those two counts differ.

go deeper

for a junior

You mainly need the shape: a P is a runtime object an OS thread must hold to run Go code, and there are GOMAXPROCS of them. Do not call it a core.

for a middle

Be ready to list what a P owns — a queue of ready goroutines and per-P caches including the allocator's — and to explain that single ownership is what removes locking from the hot path.

for a senior

Show you can use the distinction: the P count is the Go-code parallelism budget you configure, while the thread count is managed by the runtime for its own reasons, and conflating them leads to bad capacity decisions.

for a principal

The tradeoff to own is that the permit count is a policy input: it decides how much CPU a process will try to use in parallel, so it belongs in your deployment contract rather than being left to whatever a host happens to expose.

## What a P actually contains A P — the runtime's `p` structure — is a *scheduling context*. Think of it as a permit plus a toolbox. The permit half says "the holder may execute Go code"; the toolbox half is the state that makes running goroutines fast: - **A local run queue** of goroutines that are ready to run. The P holding it can push and pop without coordinating with anyone else. - **A per-P memory-allocation cache.** Go's allocator gives each P its own `mcache`, so allocating a small object usually touches only P-private storage instead of a shared, locked structure. This is a large part of why allocation-heavy Go code scales at all. - **Per-P runtime bookkeeping** — GC-related state, reusable structures, counters — all kept per P for the same reason: no lock on the hot path. A P is deliberately *not* a thread and *not* a CPU core. It is an accounting object that exists only inside the Go runtime. ## Why an M must hold one Execution is the triple M + P + G. The M is the OS thread that actually executes instructions; the G is the goroutine's stack and saved state; the P is what says this thread is currently one of the program's Go-code execution slots, and supplies the queue that decides which G runs next. The requirement buys two properties that are hard to get any other way. **1. A parallelism budget that does not depend on the thread count.** There are exactly `GOMAXPROCS` Ps. Because Go code runs only while holding one, at most `GOMAXPROCS` goroutines execute Go code at the same instant — regardless of how many Ms the runtime happens to have created. If instead the thread count *were* the budget, the runtime would have to choose between hurting parallelism and letting the budget drift. Splitting permit from worker lets the two numbers be different numbers, chosen for different reasons. **2. Ownership instead of sharing.** Whatever state the scheduler and allocator need on the hot path can be hung off the P. Since only one M holds a given P at a time, that state is single-owner and needs no locking. A design without P has to keep this state either per-thread (and then it is stranded when a thread is idle, and it grows with the thread count) or global (and then every thread contends on it). Per-P is the compromise: bounded in number, always attached to something that is actively running Go code. ## What you can observe from it The P count is the number you actually control: `runtime.GOMAXPROCS(0)` reads it, `runtime.GOMAXPROCS(n)` changes it, and the `GOMAXPROCS` environment variable sets it before the program starts. Resizing the set of Ps is a global runtime operation, not a cheap toggle, so it belongs at startup or at a deliberate configuration change — not on a request path. Setting `GOMAXPROCS` to 1 is a useful thought experiment. There is then one P, so exactly one goroutine executes Go code at a time. Your program is still *concurrent* — goroutines still interleave, still block, still resume, and the scheduler still switches between them at safe points — but it has no Go-code *parallelism*. Crucially this does not make data races impossible: two goroutines that interleave without synchronisation still race, and the race detector will still flag them. "I set GOMAXPROCS to 1 so I don't need a mutex" is wrong, and it is wrong for a reason worth stating out loud: correctness comes from synchronisation, not from how many execution slots exist. ## How to answer this in an interview Lead with what a P holds — run queue plus per-P caches — then give the one-sentence reason for the permit: it separates *how many goroutines may run Go code at once* from *how many OS threads exist*. Finish by naming the knob: the number of Ps is `GOMAXPROCS`. If you can also say that the per-P allocation cache is why small allocations avoid a global lock, you have shown that you understand P as a memory-system object as well as a scheduling one, which is the part most candidates miss.

  • Besides its place in the parallelism budget, what does a P give the memory allocator?
    A per-P allocation cache (`mcache`). Because only one M holds a P at a time, that cache has a single owner and needs no lock, so the common case of allocating a small object stays on a fast, uncontended path. Without per-P caches, allocation-heavy Go code would contend on shared structures on every allocation.
  • If you set GOMAXPROCS to 1, does that make your concurrent code race-free?
    No. One P means one goroutine executes Go code at a time, but goroutines still interleave — they yield at safe points, block, and resume — so unsynchronised access to shared state is still a data race and the race detector will still report it. Correctness comes from synchronisation, never from the size of the execution budget.
  • Can the number of Ps change while the program is running?
    Yes, but only deliberately: `runtime.GOMAXPROCS(n)` resizes the set of Ps, and on recent Go on Linux the runtime may itself revise the default when the container's CPU limit or the usable CPU count changes. Resizing is a global runtime operation, so treat it as startup or configuration work rather than something to do per request.

saying these in an interview costs you the question

  • Describes a P as a CPU core or as a thread
  • Thinks P holds no state, just a counter
  • Says GOMAXPROCS=1 removes the need for mutexes
  • Believes each M permanently owns one P
  • Cannot name anything a P owns