skip to content

Compare stackful and stackless coroutines: how does each remember where it was, and what can one do that the other cannot?

level: middleimportance: should knowfreq 44%

answer

  1. stackful = its own stack; stackless = state machine object
  2. stackful suspends from any depth, even foreign callees
  3. stackless size = live locals only, known at compile time
  4. stackless suspension propagates up the call chain
  5. stack growth: copying breaks pointers, segments can thrash

basics

~20 s

A stackful coroutine owns a real stack, so it can suspend from any depth, including inside a function that knows nothing about coroutines. A stackless one is compiled into a state machine holding only its own locals, so it can suspend only at explicit points in itself, but costs far less memory.

solid answer

~50 s

Both need to remember "where was I", and they store it differently. **Stackful** (fibers, green threads): the coroutine gets its own call stack. Switching means swapping stack pointer and registers. Because the whole stack is preserved, it can suspend from *any* frame — even deep inside a library that was never written with coroutines in mind. Cost: each coroutine reserves a stack, so runtimes use small growable or segmented stacks, and moving a stack can invalidate raw pointers into it. Interop with native/foreign frames is delicate. **Stackless**: the compiler rewrites the coroutine body into a state machine whose live locals become fields of one heap object sized exactly for that function. There is no separate stack, so memory per coroutine is tiny and predictable. But only frames the compiler transformed can suspend: suspension must be declared and propagates up the call chain, giving you the sync/async split of the API. Rule of thumb: stackful buys transparency and interop, stackless buys density and predictable cost.

go deeper

for a junior

Give the one-line contrast: stackful has a real stack and can pause anywhere; stackless is a compiler-built state machine that pauses only at declared points.

for a middle

Explain the memory and interop tradeoff — stack-sized versus exactly-the-live-locals — and why stackless suspension has to propagate to callers.

for a senior

Discuss stack growth strategies and their hazards, pinning during foreign calls, and why visible suspension points help reviewers reason about interleaving.

for a principal

Frame it as a platform decision: stackful lets an existing blocking codebase scale without rewrite; stackless gives predictable per-task footprint and no runtime dependency, at the price of a bifurcated API surface across every library you own.

## The shared problem Suspending means preserving everything needed to continue: the instruction to resume at, and every local variable still alive. The two families differ only in where that lives. ## Stackful coroutines A stackful coroutine (also fiber, green thread, or user-mode thread) is allocated its own call stack, separate from the carrier thread's. Suspending is a context switch performed entirely in user space: save the callee-saved registers and stack pointer, load the scheduler's, jump. Resuming reverses it. No kernel transition is involved, so a switch is typically a few dozen nanoseconds — far cheaper than a kernel thread switch, though more than a stackless resume. Because the entire stack survives, the coroutine can suspend from arbitrary depth. Function `a` calls `b` calls `c`, and `c` yields — none of `a` or `b` need to know coroutines exist. This is what people mean by **transparent** or **colorless** suspension: an ordinary-looking function can suspend, so libraries do not need async variants. The costs: - **Memory**: every coroutine reserves a stack. Runtimes shrink this with small initial stacks that grow on demand — either by *copying* to a bigger stack or by chaining *segments*. Copying invalidates any raw pointer into the stack (a real problem when interoperating with native code); segmenting can cause "hot split" thrash when a loop crosses a segment boundary repeatedly. - **Interop**: if a coroutine calls into foreign code that owns frames the runtime cannot relocate or does not understand, it usually must be pinned to a carrier thread for the duration, losing the benefit. - **Implementation**: it needs runtime or platform support; you cannot express it in portable source-level code alone. ## Stackless coroutines A stackless coroutine is a compile-time transform. The compiler splits the body at each suspension point and rewrites it as a state machine: a heap object with a state field and one field per local that is live across a suspension point. Resuming means calling a resume function which switches on the state and jumps to the right block. Consequences: - **Size is exact and known statically**: the object holds only the variables that must survive; typically tens of bytes. Millions of suspended coroutines are feasible. - **Resume is a plain call**: no register/stack-pointer juggling, no separate stack, so it plays well with existing calling conventions and needs no runtime magic. - **Only transformed frames suspend**: a suspension point can appear only in a function the compiler transformed. If such a function is called by an untransformed one, the caller cannot wait for it without blocking — so the property propagates up the call chain. That is the structural origin of the sync/async API split. - **The transform must be able to see the code**: you cannot suspend across a frame the compiler did not transform, which includes most foreign/native frames. ## Comparison | | Stackful | Stackless | |---|---|---| | Where state lives | Its own stack | Heap state-machine object | | Memory per coroutine | Stack-sized (KBs, growable) | Exactly the live locals (bytes) | | Suspend from any depth | Yes | No — only in transformed frames | | Splits the API into two kinds of function | No | Yes | | Needs compiler support | No (needs runtime) | Yes | | Native/foreign interop | Awkward (pinning, pointer stability) | Cannot suspend across foreign frames | | Switch cost | User-space register/stack swap | Ordinary function call/return | ## Choosing Retrofitting concurrency onto a large existing codebase full of blocking libraries favours stackful: existing code keeps working and becomes suspendable for free. Building for maximum density and predictable footprint — millions of connections, embedded budgets, or environments where you cannot install a runtime — favours stackless. A third consideration is auditability: stackless suspension points are syntactically visible, which makes it easy to see where another task can interleave; stackful suspension is invisible, which is convenient but hides those interleaving points from reviewers.

  • Why can a stackless coroutine not suspend inside a helper function it calls?
    Because only the transformed function has a state machine to record a resume point; the helper's frame lives on the ordinary call stack and would be destroyed by returning to the scheduler. To suspend there, the helper itself must be transformed, which is why the property spreads to every caller in the chain.
  • What goes wrong when a stackful runtime grows a coroutine's stack by copying it?
    Any raw pointer or reference that pointed into the old stack now points at freed or stale memory. Managed runtimes can rewrite such pointers because they track them, but pointers held by foreign or native code cannot be found and rewritten, so those coroutines usually must be pinned to a stack that will not move.

Stackful is packing your entire desk, chair and all, into storage so you can return to it exactly as it was. Stackless is writing a sticky note listing only the few things you still need and which step you were on.

saying these in an interview costs you the question

  • Saying stackless coroutines have no state at all — they have exactly the live locals, on the heap
  • Claiming stackful coroutines are just OS threads with a different name — the switch is user-space, not kernel
  • Asserting stackful coroutines are always slower; their switch is cheap, the real cost is stack memory
  • Believing the sync/async API split is a language style choice rather than a consequence of the stackless transform
  • Assuming either kind gives you parallelism without multiple carrier threads

context