When you start a goroutine with go f(), where does the Go runtime put it to wait?
answer
- three places, not one queue
- one fast slot, then a ring buffer
- the local ring holds 256
- overflow spills half to the shared queue
- the newest goroutine jumps the line
basics
~20 sThe new goroutine goes into the runnext slot of the P that ran the go statement, bumping any goroutine already there onto that P's 256-entry local run queue. When that queue is full, half of it spills to the global run queue.
solid answer
~50 sGo's runtime keeps runnable goroutines in three places: a one-goroutine `runnext` fast slot per P, a fixed 256-entry local run queue per P, and one global run queue shared by all Ps. A `go` statement creates the goroutine and puts it in the current P's `runnext`, displacing the previous occupant to the tail of that P's local queue. If the local queue is full, the runtime moves half of it plus the new goroutine to the global queue as one batch, so the global lock is touched once per batch rather than once per goroutine. The parent keeps running — starting a goroutine does not yield — and the runtime may wake a thread to pick the work up on an idle P. Local queues exist so that the common case needs no lock; the global queue is overflow and a fairness backstop.
code
go · 11 linesdone := make(chan struct{})
// f is now runnable in this P's runnext slot,
// but this goroutine keeps running
go func() {
fmt.Println("child")
close(done)
}()
fmt.Println("parent")
<-done // the only thing that actually orders the twogo deeper
Be ready to say that a goroutine is queued, not run, and that each P has its own queue alongside one shared global queue. That the go statement returns instantly is the point of the question.
Explain the three-place layout — the runnext slot, the 256-entry local ring, the global queue — and why overflow moves a whole batch rather than a single goroutine.
Expect to connect placement to observed behaviour: work created on one P stays there until another P steals it, which is why a burst can look unevenly spread across CPUs.
Frame it as a contention tradeoff the runtime already made: per-P queues buy lock-free scheduling at the price of short-term imbalance, and application-level goroutine pools rarely improve on it.
## Where a runnable goroutine actually sits Go's runtime multiplexes goroutines onto operating-system threads using three runtime objects: a **G** is a goroutine, an **M** is an OS thread, and a **P** is a scheduling context — a token an M must hold in order to run Go code. The P is the piece that matters for this question, because **run queues hang off Ps, not off threads.** A thread reaches a queue only through the P it currently holds. At any instant a *runnable* goroutine — ready to execute, not currently executing, not blocked on anything — sits in exactly one of three places: 1. **A P's `runnext` slot.** Exactly one goroutine, no more. It is a fast slot in front of the queue holding the goroutine that this P most recently made runnable, and the P will run it ahead of everything in its queue. 2. **A P's local run queue.** A fixed-size circular buffer of 256 entries, owned by that P. 3. **The global run queue.** One list for the whole process, protected by the scheduler's global lock. ## What `go f()` actually does The `go` statement compiles into a call to the runtime's `newproc`. That allocates a goroutine (usually reusing a dead one from a free list), copies the call's arguments onto the new goroutine's stack, sets its program counter to `f`, marks it runnable, and hands it to `runqput` with the "put it next" flag set. So the new goroutine lands in the **current P's `runnext` slot**, and whatever goroutine was sitting there is bumped to the tail of that P's local run queue. Then the runtime calls `wakep`, which — if there is an idle P and no thread is already hunting for work — wakes or starts a thread to go and find something to run. On a multi-core machine that means the goroutine you just created can start running on a different P almost immediately. What does **not** happen is a context switch. `go f()` returns and the parent goroutine keeps running its time slice. The child is merely *queued*. ## What happens when the local queue is full `runqput` cannot push into a local queue that already holds 256 goroutines. The slow path then takes the first half of that queue — 128 goroutines — plus the new one, links them into a batch, and pushes the whole batch onto the global run queue in a single acquisition of the scheduler lock. Two things fall out of that design. First, a program that creates goroutines faster than it runs them does not grow an unbounded per-P queue; the excess drains into the shared one. Second, the global lock is touched roughly once per 128 goroutines rather than once per goroutine, so the overflow path stays cheap even under a burst. ## Why per-P queues exist at all Contention. A single shared queue of runnable goroutines would be hit on every goroutine creation, every wake and every scheduling decision on every core — a textbook lock convoy. A P's local queue is instead a single-producer, multi-consumer ring: the owning P pushes and pops its own end with ordinary atomic loads and stores and no lock, while other Ps stealing from it compare-and-swap at the far end. The overwhelmingly common operations therefore need no lock at all. The global queue exists for overflow, for goroutines readied when no P is available to take them, and as a fairness backstop. ## How goroutines come back out When a P needs something to run, the scheduler looks in this order: on every 61st scheduling tick it deliberately takes one goroutine from the global run queue first (so local work cannot starve the shared queue forever); otherwise it takes `runnext` if it is set; otherwise the head of the local queue; and if all of that is empty it goes hunting — a batch from the global queue, a poll of the network poller, and then stealing from other Ps. ## Practical consequences for code you write - **`go f()` is cheap and never blocks.** It is a few tens of nanoseconds plus stack setup. Even the overflow path is a batch move, not a stall. - **The child does not run first.** Nothing orders parent and child, which is why a test that assumes the goroutine has already run by the next statement is flaky. Synchronise with a channel or a `sync.WaitGroup`. - **Work has affinity to the P that created it.** A burst of goroutines created on one P begins life on that P and reaches other Ps only when they steal it. - **There are no priorities.** You cannot mark a goroutine as important. The only ordering preference in the system is the `runnext` slot, and the runtime decides who gets it.
- Does the new goroutine start running immediately when the go statement executes?No. `go f()` makes the goroutine runnable and returns; the parent keeps its time slice. The child runs when its P reaches it — often very soon, since it sits in `runnext`, and sooner still if an idle P's thread is woken to take it. Nothing orders parent and child, which is why a test assuming the goroutine has already run by the next line is flaky.
- Why is a P's local run queue a fixed 256 entries rather than a growable one?A fixed circular buffer is manipulated with plain atomics and no allocation: the owner pushes and pops its end, thieves compare-and-swap at the other. Growing it would mean allocating and locking on the hottest path in the scheduler. Overflow is handled instead by moving half the queue plus the new goroutine to the global run queue in one batch, which bounds how often the global lock is taken.
- What reaches the global run queue besides overflow from a full local queue?Goroutines readied when no P is available to take them, and the queues of Ps that are retired when the runtime changes how many Ps it has. Nothing routes work there deliberately. In the other direction, a P drains a batch from the global queue whenever its local queue empties, and periodically even when it does not.
saying these in an interview costs you the question
- Says the runtime keeps one global queue of runnable goroutines
- Thinks go f() runs the function immediately
- Claims each OS thread owns the run queue
- Says the local run queue grows without bound
- Believes the OS assigns goroutines to processors