skip to content

When a daemon's thread count rises from 8 to 800, which regions of its memory map multiply and which do not?

level: middleimportance: should knowfreq 48%

answer

  1. split shared term from per-thread term
  2. seven regions, two carry the count
  3. one heap, many stacks
  4. thread-local block scales with stacks too
  5. reservation per thread times thread count

basics

~20 s

Only the per-thread regions multiply: each thread gets its own stack reservation and its own thread-local block. Code, constants, static data and the heap stay single and shared, so thread count scales just one term of the footprint.

solid answer

~40 s

Split the footprint into a shared term and a per-thread term. The shared term - code, read-only constants, both static-data regions and the one heap - does not change when threads are added, because every thread uses the same mapped copy. The per-thread term does: each thread is created with its own **stack reservation** and its own **thread-local storage** block, so it is multiplied by the thread count. At 1 MB of stack and 8 KB of thread-local storage per thread, 8 threads reserve about 8 MB and 800 threads about 806 MB of address space, with the shared term unchanged in both. That is why a thread-per-connection design hits an address-space wall long before its heap looks interesting, and why the two knobs are thread count and per-thread stack size.

code

pseudocode · 10 lines
pseudocode
per_thread_stack_reservation = 1024 KB
per_thread_local_block       =    8 KB
per_thread                   = 1032 KB

function footprint(threads):
    shared = code + constants + static_data + heap
    return shared + threads * per_thread

footprint(8)   = shared +   8 * 1032 KB  ~ shared +   8 MB
footprint(800) = shared + 800 * 1032 KB  ~ shared + 806 MB

go deeper

for a junior

Remember the split: one heap, one copy of the code and static data, but one stack and one thread-local block for every thread that exists.

for a middle

Do the arithmetic out loud - reservation times thread count - and say which term of the footprint a given change actually moves.

for a senior

When a footprint tracks concurrency instead of data volume, look at the per-thread regions first, and report reserved address space separately from physical cost.

for a principal

Decide the concurrency model with this term in mind: a thread per unit of waiting work buys simplicity and pays for it linearly in reserved address space.

## Split the footprint into two terms The memory map of a running process divides cleanly along one line: regions that exist once for the process, and regions that exist once per thread. Every question about how a process scales with concurrency is a question about which side of that line a region sits on. | Region | Copies | Multiplied by thread count | |---|---|---| | Code | one, shared | no | | Read-only constants | one, shared | no | | Initialized static data | one, shared | no | | Zero-filled static data | one, shared | no | | Heap | one, shared | no | | Stack | one per thread | yes | | Thread-local storage | one per thread | yes | Two of seven regions carry the thread count. That is the whole answer, and everything else is a consequence of it. ## Why the shared regions cannot multiply - **Code and constants** are never written, so one mapped copy serves every thread. There is nothing a second thread would need a private copy of. - **Static data** is reserved once, before the first instruction runs, and its whole purpose is that all threads see the same variables. Giving each thread a copy would defeat it - and would also remove the reason such variables need synchronising. - **The heap** is one region per process. Every thread requests blocks from it and blocks in it can be reached from any thread. Allocators commonly keep per-thread bookkeeping to reduce contention, but that is bookkeeping inside one region, not a second heap region. ## Why the per-thread regions must multiply A thread is a call chain. Two threads execute two different chains at the same instant, so each needs its own range to hold its own active calls - they cannot interleave in one range. Thread-local storage follows the same logic one level up: it exists precisely to hold a value that is one-per-thread rather than one-per-process, so a per-thread block is the definition of the feature, not an implementation detail. ## The arithmetic With a per-thread stack reservation of 1 MB and a thread-local block of 8 KB, each thread costs about 1032 KB of address space: 1. At 8 threads the per-thread term is about 8 MB - invisible next to almost any real shared term. 2. At 800 threads it is about 806 MB - usually larger than everything else in the map put together. 3. The shared term is byte-for-byte identical in both cases. The practical shape of that curve is why a design that dedicates a thread to every connection runs into a wall well before its data volume suggests any problem: the wall is the per-thread term, and the data is in the shared term. ## What the arithmetic does not say Reserving a range is not the same as consuming physical memory. A stack reservation is an address range set aside so the call chain has somewhere to grow; how much of it is ever backed by hardware, and how the operating system accounts for that, is a separate subject with its own owner. Keep the two claims apart when you report this: the multiplication is a fact about the map, and the physical cost is a question you answer with operating-system tools. ## The two knobs - **Thread count.** Handling more concurrent work with fewer threads - a pool with a bounded size, or a model where waiting work does not hold a thread at all - reduces the multiplier directly. - **Per-thread stack size.** Chosen when a thread is created, and often left at a generous default. Halving it halves the per-thread term, at the price of less room for the call chain, which is a real trade rather than free money. A third move changes the shape rather than the size: if waiting work does not need to hold a call chain at all, the per-thread term stops tracking the number of concurrent operations and starts tracking the number of workers, which is a much smaller and much flatter number. Turning either of the two knobs moves the same term. Adding heap capacity moves a different term and will not help at all, which is the diagnosis mistake this question is really testing: a footprint that tracks concurrency rather than data volume is pointing at the per-thread regions, and no amount of attention to the heap will explain it.

  • Does raising the thread count raise heap usage at all?
    Indirectly, and for a different reason: more threads means more work in flight at once, so the peak live data can rise. But the heap region itself is not duplicated - there is still one of it. If the footprint tracks thread count rather than work volume, the per-thread regions are the explanation.
  • Which knob changes the per-thread term without changing the thread count?
    The stack reservation requested when each thread is created. It is frequently left at a generous default, and lowering it scales the whole per-thread term down proportionally. The cost is a shallower call chain before the stack is exhausted, so the figure has to be chosen against the deepest work the threads actually do.

saying these in an interview costs you the question

  • Thinks each thread gets its own heap region
  • Says code and constants are duplicated per thread
  • Believes static variables are copied for every thread
  • Treats a stack reservation as physical memory consumed
  • Assumes only the heap can dominate a process footprint