skip to content

How much stack does a new goroutine start with, and what happens when it needs more?

level: middleimportance: should knowfreq 55%

answer

  1. measured in kilobytes, not megabytes
  2. the compiler checks before each call
  3. it is moved, not extended in place
  4. pointers into it are rewritten on the move

basics

~20 s

About 2 KB. When a call would overflow it, the Go runtime allocates a larger stack, copies the existing frames across and adjusts pointers into them, so stacks grow on demand instead of being reserved up front like a thread's.

solid answer

~50 s

A goroutine starts with a stack of roughly **2 KB**, allocated by the runtime rather than the kernel. The compiler puts a small bounds check in each function's prologue; when a call would not fit, the runtime allocates a new, larger stack — typically double — copies the live frames into it and rewrites pointers that referred into the old one, then resumes. Stacks can also be shrunk during garbage collection when a goroutine turns out to be using far less than it holds. The contrast with an OS thread is the whole point: a thread reserves megabytes of address space at creation, so tens of thousands of them is expensive, while `go check(h)` in a loop over 100,000 hostnames is an ordinary Go program. Runaway recursion still ends badly — the runtime enforces a maximum stack size, settable with `runtime/debug.SetMaxStack`, and exceeding it is a fatal stack overflow.

code

go · 3 lines
go
for _, h := range hosts { // hosts holds 100,000 names
	go check(h) // each starts with roughly 2 KB of stack, grown on demand
}

go deeper

for a junior

Know the headline: a goroutine starts with a couple of kilobytes of stack, not the megabytes an OS thread reserves, which is why starting thousands of them is normal in Go.

for a middle

Explain the growth mechanism — the prologue check, allocating a bigger stack, copying frames and fixing pointers — and mention that stacks can shrink again during garbage collection.

for a senior

Reason about the consequence in production: cheap stacks make blocking cheap, but long-lived goroutines keep their grown stacks alive and give the collector more to scan, so goroutine count is still a resource you watch.

for a principal

Weigh the design tradeoff: growable copied stacks buy massive concurrency at the cost of precise pointer metadata and a copy on growth. Be ready to say when per-item goroutines stop being the right shape for a workload.

## Why the number matters The reason Go programs can start goroutines so freely is that a goroutine is cheap in exactly the place threads are expensive: **stack memory**. A new goroutine begins with a stack of about **2 KB**, plus a small runtime structure holding its scheduling state. A typical OS thread reserves stack address space measured in megabytes when it is created. That difference is what turns "start one goroutine per unit of work" from a stunt into an idiom. ```go for _, h := range hosts { // hosts holds 100,000 names go check(h) // ~2 KB of stack each to begin with } ``` This is a normal thing to write in Go (whether it is a *good* thing depends on what `check` does downstream, which is a separate question about bounding concurrency). ## Growing Goroutine stacks are **contiguous and growable**. The mechanism has two halves: - **The check.** The compiler emits a small comparison at the top of most functions: does the frame this function needs fit in the remaining stack? Leaf functions with tiny frames can skip it. - **The copy.** If it does not fit, the runtime allocates a new stack — larger, usually double the old size — copies the existing frames into it, and then fixes up every pointer that pointed into the old stack so it points at the corresponding place in the new one. The goroutine then resumes as if nothing had happened. The fix-up step is why Go can move stacks at all: the runtime knows, from compiler-generated metadata, exactly which words in a frame are pointers. A language that cannot identify pointers precisely cannot relocate a stack, which is why this design is not universal. ## Shrinking The growth is not permanent. During garbage collection the runtime may **shrink** a goroutine's stack if it is using far less than it currently holds, copying the frames down into a smaller allocation. A goroutine that once recursed deeply and now sits in a loop does not keep paying for the peak forever. ## The ceiling Growth is not unbounded. The runtime enforces a maximum stack size per goroutine; `runtime/debug.SetMaxStack` reads and sets it. Infinite recursion therefore does not consume all of memory — it hits the ceiling and the program dies with a fatal `stack overflow` error and a dump. Note this is a *fatal* runtime error, not a panic you can recover from. ## What this buys the scheduler Cheap stacks are what make blocking cheap. Because a parked goroutine costs only its (often still small) stack plus bookkeeping, the runtime can afford to have a great many goroutines blocked on channels, mutexes and network I/O at once, and to park them rather than blocking OS threads. That is the practical reason a straightforward blocking style — "write it as if it were sequential, and start a goroutine per unit of work" — is idiomatic Go rather than a performance trap. ## What it does not mean Two cautions worth voicing in an interview: - **2 KB is a floor, not a budget.** A goroutine whose work allocates large buffers, or recurses, uses far more. Estimating memory as "goroutines times 2 KB" is wrong for anything but idle goroutines. - **Cheap does not mean free.** Each goroutine still costs a stack that survives as long as the goroutine does, plus scheduler bookkeeping. A million goroutines blocked forever is still a large amount of live memory that the garbage collector must scan, which is precisely why goroutines that never finish are a problem worth taking seriously. ## Answering crisply "About 2 KB, heap-allocated by the runtime and grown by copying when a prologue check finds it too small, shrunk again during GC, with a hard ceiling you can set through `runtime/debug.SetMaxStack`. That is why a `go` statement in a loop over a hundred thousand items is unremarkable, whereas a hundred thousand OS threads is not."

  • How does the runtime know a goroutine's stack is about to be too small?
    The compiler inserts a bounds check in the function prologue comparing the frame the function needs against the space left. If it does not fit, the call traps into the runtime, which grows the stack before the function body runs. Small leaf functions can have the check elided.
  • Why can Go move a goroutine's stack when C cannot?
    Moving a stack means rewriting every pointer that points into it. The Go compiler emits metadata describing precisely which words in each frame are pointers, so the runtime can find and update them all. Without that precise information, relocating a stack would leave dangling pointers behind.
  • What happens to a goroutine that recurses without a base case?
    It grows its stack repeatedly until it reaches the runtime's maximum stack size, then the program dies with a fatal stack overflow error and a goroutine dump. It is a fatal runtime error, not a recoverable panic. The ceiling can be adjusted with `runtime/debug.SetMaxStack`.
  • Is estimating memory as number of goroutines times 2 KB reasonable?
    Only for goroutines that are idle and shallow. 2 KB is the starting size; any goroutine that recurses or holds buffers uses more, and its stack stays alive as long as it does. Treat 2 KB as the floor that makes goroutines cheap to start, not as a per-goroutine budget.

saying these in an interview costs you the question

  • Says a goroutine gets a megabyte-sized stack like a thread
  • Thinks the stack is extended in place into adjacent memory
  • Believes goroutine stacks only ever grow, never shrink
  • Claims stack overflow can be recovered like a panic