skip to content

Why can a ThreadPoolExecutor of 8 workers end up with hundreds of live threads in one process?

level: seniorimportance: nice to knowfreq 22%

answer

  1. Concurrency levels do not simply add
  2. The default is chosen twice, independently
  3. Outer workers times inner workers
  4. A native runtime starts its own thread team
  5. Budget threads per process, not per pool

basics

~20 s

Because pools nest and multiply. If each of the 8 tasks calls code that starts its own pool sized to the machine, the process runs 8 times that many threads, and a process pool repeats it inside every child.

solid answer

~50 s

Thread counts compose multiplicatively. Eight outer workers each calling a helper that builds its own `ThreadPoolExecutor`, or each calling into a compiled extension whose runtime starts one thread per logical CPU on first use, gives 8 × N threads. With a `ProcessPoolExecutor` it is worse: every child interpreter repeats the inner sizing, so a 16-core host can reach 16 × 16 runnable threads for 16 cores' worth of work. The cost is real — stack reservations per thread, scheduler churn, cache thrashing between threads fighting for the same cores, and GIL contention for any inner work that is pure Python. The fixes are to budget the *product* rather than each pool, pin an inner native runtime to one thread through its own thread-count environment variable before it is imported, and create long-lived pools once instead of one per call.

code

python · 15 lines
python
import threading, time
from concurrent.futures import ThreadPoolExecutor

def leaf(_):
    time.sleep(0.05)
    return threading.active_count()

def task(_):
    with ThreadPoolExecutor(max_workers=4) as inner:   # inner fan-out per task
        return max(inner.map(leaf, range(4)))

with ThreadPoolExecutor(max_workers=8) as outer:       # ... 8 outer workers
    peak = max(outer.map(task, range(8)))

print("peak live threads while 8 workers run:", peak)   # about 41, not 8

go deeper

for a junior

Know that creating a pool inside a function that runs many times is a mistake, and that a pool's max_workers only counts the threads that pool starts, not the ones anything it calls may start.

for a middle

Be able to compute the product: outer workers times whatever inner concurrency each task triggers, and explain why oversubscribed threads cost scheduling, memory and, for pure-Python work, GIL contention.

for a senior

Show how you would find it in a live process, using the live thread count against the configuration you declared, and how you would fix it by choosing which level owns the parallelism and pinning the other to one.

for a principal

Own it as a process-wide budget rather than a bug. Decide where parallelism lives in the stack, make that a stated convention libraries and services follow, and require that thread totals be observable so a dependency upgrade cannot quietly double them.

### Pools compose by multiplication Every sizing decision in a process is usually made in isolation: a helper module builds a pool sized to the machine because that is a sensible default for a helper module, and the caller builds a pool sized to the machine because that is a sensible default for a caller. Composed, the defaults multiply. Eight outer workers, each entering a function that creates a four-worker inner pool, means thirty-two worker threads plus the eight outer ones, doing eight cores' worth of useful work at best. The multiplication has three common sources. The first is an explicit inner `ThreadPoolExecutor` created inside the task function, often as a convenience for parallelising a small fan-out. The second is a compiled extension: a native numeric or compression runtime typically starts a thread team on first use, sized to the logical CPU count, and it does this once per *process*, invisibly, with no Python object you would notice. The third is the same helper called through a `ProcessPoolExecutor`, where each child interpreter independently repeats whatever the parent did — so a per-process inner pool becomes cores × cores threads across the machine. ### Why the extra threads cost more than they look Oversubscription is not free idling. Each OS thread reserves stack space, so a few hundred threads is a measurable and sometimes surprising memory line even when they are asleep. Runnable threads beyond the CPU count force the scheduler to time-slice, and every slice ends with cold caches for the thread that resumes; for memory-bound inner loops that alone can halve throughput. Threads doing pure-Python work also serialise on the GIL, so multiplying them adds contention on the one lock without adding parallelism. And in a nested arrangement the failure is self-reinforcing: inner pools created per call are also created and destroyed constantly, paying thread start-up over and over. What makes it a diagnosis question rather than a design question is that nothing in the code says 256. The outer pool says 16, the inner default says 16, and the product exists only at run time. `threading.active_count()` in a health endpoint or a log line at start-up is the cheapest way to notice, and it is worth having before you need it. ### Fixing it The rule is to budget the product, not each pool. Decide how many runnable threads the process should have in total — for CPU-bound work, roughly `os.process_cpu_count()` — and then divide it across the nesting levels rather than letting each level claim it. In practice that means one of three shapes. Pin the inner level to one. A native runtime that reads a thread-count environment variable will honour a value of 1 if it is set **before the extension is imported**, because the team is usually sized once at load or first use; setting it afterwards from Python has no effect. This is the standard arrangement when the outer pool is the one you want to be wide: many single-threaded inner computations, parallelism owned at the top. Or pin the outer level to one and let the inner runtime spread. That is the right answer when the inner work is compiled, releases the GIL and parallelises well internally: one caller, one wide native thread team, no Python-level contention at all. Or keep both, sized deliberately, when the levels genuinely serve different purposes — an outer pool of a few threads for blocking I/O, an inner pool for compute — and check that the product is still the number you intended. Alongside the sizing, hoist pools out of task functions. An executor created inside a function that runs thousands of times is both a multiplication and a churn problem; one module-level executor, or an executor passed in as an argument, fixes both and makes the total count auditable by reading the code. If a library you call insists on managing its own pool and gives you no way to size it, that is a genuine reason to run it in a process pool with a single-threaded inner configuration rather than inside your thread pool. ### The interview shape This is a differentiator rather than a screening question: many strong Python engineers have never hit it, because it only bites when compute-heavy libraries meet a hand-rolled pool. But it is the concrete meaning of the word oversubscription, and the answer an interviewer is listening for is the multiplication itself — that concurrency levels compose by product, and that a sizing policy has to be stated for the whole process rather than pool by pool.

  • Why is a process pool worse for this than a thread pool?
    Because each child is a fresh interpreter that repeats the inner sizing decision independently. A per-process native thread team sized to the machine is created once in a thread pool and once per child in a process pool, so the total becomes workers times cores. Children also cannot see each other, so nothing coordinates the total; the only place the product can be capped is in the environment the children inherit or in the initializer that runs inside each one.
  • How would you detect this in a running service rather than by reading the code?
    Log threading.active_count() at start-up and expose it on a health endpoint, then compare it with the number your configuration implies. A count far above the sum of your declared pool widths means something below you is starting threads. Operating-system per-thread views confirm it and usually name the culprit, since native runtimes tend to give their threads recognisable names.
  • When is it right to let the inner level be the wide one?
    When the inner work is compiled, releases the GIL and parallelises well internally. Then a single caller driving one wide native thread team gets full use of the cores with no GIL contention and no Python-level scheduling, and the outer pool exists only to overlap I/O. Reversing that, many outer workers each with a single-threaded inner computation, is the right shape when the parallel unit is a whole task rather than the maths inside one.

Two managers each told to use the whole team do not share it; they each hire a full team, and the payroll is the product of their two reasonable decisions.

saying these in an interview costs you the question

  • Assumes total threads equal the outer pool's max_workers
  • Creates a new executor inside a function called repeatedly
  • Sets a native thread-count variable after the extension is imported
  • Sizes each pool to the core count independently
  • Thinks idle threads cost nothing at all
  • Expects a process pool to share one thread budget

context