skip to content

A machine can comfortably run a few thousand operating-system threads, yet some runtimes host a million lightweight (runtime-scheduled) threads on the same hardware. What actually accounts for that difference, and where is the new ceiling?

level: middleimportance: should knowfreq 52%

answer

  1. Contiguous reserved stack + guard page vs heap continuation
  2. Kernel stack ~8–16 KB, non-pageable, per thread
  3. Growable / copying / segmented stacks
  4. Switch: syscall + cold cache vs register swap
  5. New ceiling = live frames × count, not core count

basics

~20 s

An OS thread carries a large fixed stack reservation plus non-pageable kernel bookkeeping, and every switch is a trip through the kernel. A lightweight thread is a heap object with a tiny growable stack switched in user space — kilobytes, not megabytes. The new ceiling is heap and live stack depth.

solid answer

~50 s

Three costs separate them. **Stack.** An OS thread gets one contiguous stack reserved up front — typically 512 KB to 8 MB of address space, physically committed as touched — plus guard pages. A lightweight thread's stack lives on the heap and starts at hundreds of bytes to a few kilobytes, growing by copying, segmenting, or by storing only the live frames as a continuation when it parks. **Kernel bookkeeping.** Each OS thread adds a task structure, a kernel stack that cannot be paged out, and an entry in the kernel's scheduling data structures. A lightweight thread costs the kernel nothing; it is invisible. **Switch cost.** Kernel switches cross a privilege boundary and disturb cache and TLB state — roughly microseconds. Swapping a continuation in user space is tens of nanoseconds. The ceiling moves rather than disappearing: memory is now heap for live stacks, so deep call stacks or large per-thread buffers still bound the count. Compute throughput is unchanged — it is still bounded by cores.

code

text · 10 lines
text
OS threads:        10,000 threads
  kernel stack     ~12 KB each, resident, non-pageable  -> ~120 MB
  user stack        1 MB reserved, ~16 KB touched       -> ~160 MB committed,
                                                            10 GB address space reserved
  switch cost      ~1-3 us + cache/TLB refill

Lightweight:    1,000,000 threads
  parked frames    ~1-4 KB of heap each                 -> ~1-4 GB heap
  switch cost      ~tens of ns, user space
  kernel cost      0 (kernel sees only the N host threads)

go deeper

for a junior

Name the two big items: OS threads reserve a large fixed stack and cost the kernel real bookkeeping; lightweight threads are small heap objects switched without entering the kernel.

for a middle

Explain reserved versus committed memory, the growable/copying/continuation stack strategies, and that the ceiling becomes heap and live stack depth rather than kernel resources.

for a senior

Reason numerically — in-flight count times realistic per-thread memory — and identify which per-thread patterns in an existing codebase stop being affordable when thread counts jump three orders of magnitude.

for a principal

Treat thread cost as a capacity model: what the cheap representation unlocks, what implicit limits it removes (pool-as-admission-control, connection ceilings), and what must be added back explicitly.

## What one operating-system thread costs When a process asks the kernel for a thread, three separate resources are consumed. **A contiguous stack.** The thread needs a stack it can grow downward without moving, so the system reserves a fixed contiguous range of the process's address space — commonly 512 KB to 8 MB, depending on platform defaults — plus a guard page below it to convert overflow into a fault instead of corruption. Reservation is virtual: physical pages are committed only as the stack is touched, so a thread that uses 8 KB of stack does not consume 1 MB of RAM. But the address-space reservation is real and, on a 32-bit space or with aggressive overcommit limits, it binds. The mistake in both directions is common: candidates either claim each thread costs a megabyte of RAM (usually false) or that stack size does not matter at all (also false, because commitment follows actual depth and every thread pays *something*). **Kernel bookkeeping.** Each thread is an entry in the kernel's scheduler: a task structure, credentials, signal state, and a small kernel-mode stack — typically 8–16 KB — which is *not* pageable. That is genuine, permanently resident memory per thread, and it scales linearly. **Scheduler pressure.** More runnable threads mean longer run queues, more load-balancing work across cores, and more preemption. Each preemption crosses the user/kernel boundary and leaves the incoming thread with a cold cache and TLB, so the real cost of a switch is usually dominated by the working set it destroys rather than the register save itself. Add those up and a few thousand threads is comfortable, tens of thousands is uncomfortable, and hundreds of thousands is not a design that works. ## What one lightweight thread costs A lightweight thread — the general concept behind "virtual threads", green threads and goroutine-style threads — is an ordinary object on the heap owned by the language runtime. Creating it is an allocation, not a system call. It holds scheduling state and a stack, and the stack is where all the savings live. Three strategies are used, sometimes in combination: - **Small growable stacks.** Start at a few kilobytes and grow when a check at function entry detects insufficient room. - **Copying/moving stacks.** On growth, allocate a larger block and copy the frames over, fixing up pointers. This requires the runtime to know the layout precisely, which is why it is practical in managed runtimes and awkward in C. - **Segmented stacks.** Chain additional chunks instead of copying — cheap to grow, but a call that repeatedly crosses a segment boundary can thrash. - **Continuation capture.** When the thread parks, the runtime copies only its *live* frames to the heap and releases the host stack entirely; on resume, it copies them back onto whatever host is free. Switching is a user-space operation: save a few registers, swap to another continuation, jump. No privilege transition, no TLB shootdown, no kernel run queue. So the per-thread cost falls from "a large virtual reservation plus non-pageable kernel memory" to "a few hundred bytes to a few kilobytes of heap", and a million of them becomes arithmetically ordinary. ## Where the new ceiling is The limit does not vanish; it changes units. **Live stack depth.** A lightweight thread parked twenty frames deep with big locals is not cheap. A framework that runs deep call chains per task can push the average from 1 KB to 20 KB, and a million threads then means gigabytes. **Per-thread data.** Anything the program allocates per thread — buffers, connection wrappers, thread-local context, tracing state — multiplies by the thread count, and that pattern was invented in an era when threads were scarce. Copying a 64 KB I/O buffer per thread is fine at 200 threads and fatal at 200,000. **Garbage-collection and allocation pressure.** Stacks on the heap are objects the collector must scan or manage; a million short-lived lightweight threads is a million allocations. **Downstream resources.** A million threads that each want a database connection or a file descriptor hit those limits long before memory. Cheap threads remove the accidental admission control that a bounded thread pool used to provide. **CPU throughput is unchanged.** Compute capacity is set by cores. Lightweight threads let you *represent* a million mostly-waiting activities affordably; they do not let you *compute* faster. ## How to reason about it numerically Estimate with concurrency = throughput × latency. Twenty thousand requests per second each holding a thread for 200 ms of mostly waiting means 4,000 in-flight threads. At an OS thread's committed cost that is uncomfortable but survivable; at 200 ms of waiting with 100,000 in flight it is not, and the lightweight model is the only affordable representation. That framing — count the in-flight activities, multiply by realistic per-thread memory — is what separates a memorised answer from a designed one.

  • If each OS thread reserves 1 MB of stack, why does 1,000 threads not consume 1 GB of RAM?
    Because the reservation is address space, not committed physical memory. Pages are backed only as the stack is actually touched, so a thread whose deepest call chain uses 12 KB commits roughly that plus a page or two. The reservation still matters for address-space exhaustion and for the guard-page layout, and the non-pageable kernel stack is genuinely resident, but conflating reserved with resident overstates the cost by orders of magnitude.
  • What patterns from the scarce-thread era become dangerous once threads are cheap?
    Anything sized per thread: pooling by thread identity, per-thread I/O buffers, per-thread caches, and treating the pool size as the system's concurrency limit. At 200 threads a 64 KB buffer each is 13 MB; at 200,000 it is 13 GB. Most importantly, a bounded pool was implicit backpressure, so unbounded lightweight threads require explicit admission control — a semaphore or queue — to replace it.

An OS thread is a hotel room booked for the whole stay whether you sleep there or not; a lightweight thread is a locker that holds only what you are actually carrying, and you rent one only while you are waiting.

saying these in an interview costs you the question

  • Claiming each OS thread consumes its full stack size in physical RAM
  • Saying lightweight threads increase CPU throughput
  • Believing lightweight threads have no stack at all
  • Assuming a million lightweight threads is always safe, ignoring per-thread buffers and downstream connection limits
  • Attributing the savings purely to switch speed and ignoring stack allocation strategy

context