skip to content

Why does every `concurrent.futures.InterpreterPoolExecutor` worker re-import your modules?

level: seniorimportance: should knowfreq 22%

answer

  1. Isolation is not free at startup
  2. sys.modules is per interpreter
  3. Caches and config duplicated per worker
  4. Runtime mutation in the parent is invisible
  5. Some extensions refuse to load at all

basics

~20 s

Each worker runs its own interpreter with its own sys.modules, so no import from the parent is visible and every worker re-executes the import graph. Caches and configuration are per worker, and some native extensions refuse to load at all.

solid answer

~50 s

An `InterpreterPoolExecutor` worker is a separate interpreter, and `sys.modules` is per-interpreter state. Nothing the parent imported carries over, so each worker re-executes every import your tasks touch — on top of a few milliseconds to create the interpreter itself. Three consequences bite in production. Startup cost scales with worker count and dependency depth, so short tasks with a heavy import graph spend their time importing. Module-level state is duplicated: caches, lookup tables and connection pools exist once per worker, multiplying memory, and any configuration the parent mutated **at runtime** is absent — the worker sees the module's import-time defaults. And a native extension that has not declared subinterpreter support raises `ImportError` in a worker, which can rule the pool out entirely. Pass configuration explicitly per task, keep worker-side imports lean, and verify the dependency set before committing.

code

python · 15 lines
python
from concurrent.futures import InterpreterPoolExecutor

SEEN = 0


def count_row():
    global SEEN
    SEEN += 1
    return SEEN


if __name__ == "__main__":
    with InterpreterPoolExecutor(max_workers=1) as pool:
        print([pool.submit(count_row).result() for _ in range(3)])
    print("main interpreter still sees:", SEEN)

go deeper

for a junior

Take away the core fact: each worker interpreter has its own sys.modules, so imports and module-level variables are not shared with the parent or with other workers.

for a middle

Be able to explain what that costs — a fresh import graph per worker, one copy of every module-level cache per worker — and to show it with a module global that a worker increments while the parent still reads its original value.

for a senior

An interviewer expects a diagnosis story: low CPU with poor speedup points at import time, linear memory growth points at duplicated caches, and output that is wrong only under the pool points at configuration the parent mutated at runtime and workers never saw.

for a principal

Own the adoption call. Whether an interpreter pool is viable is largely a dependency-compatibility and memory-budget question, so decide what evidence gates the switch and set the convention — configuration as task input, never as ambient module state — before teams build on it.

### The mechanism A `concurrent.futures.InterpreterPoolExecutor` worker is an OS thread running its own interpreter, and `sys.modules` is per-interpreter state. The mapping starts nearly empty — a fresh interpreter carries only the frozen bootstrap modules — so the first task that touches a module triggers a full import in that worker: find the source, execute it top to bottom, bind the resulting module object into *that* interpreter's `sys.modules`. Do it again for the next worker. Nothing the parent imported is visible, and nothing one worker imports is visible to another. That is the same isolation that gives the pool its parallelism. It is not a bug to route around; it is the deal. ### Cost one: startup Creating an interpreter is cheap in absolute terms — single-digit milliseconds — but the import graph on top of it is not bounded by anything except your dependencies. A worker that ends up importing a deep tree of packages can spend hundreds of milliseconds before its first task runs, and it pays that once per worker, not once per process. For long-lived pools this amortises away. For a pool created per request, or a pool of many workers running short tasks, it can dominate the wall clock and quietly undo the parallelism you switched for. The executor accepts an `initializer` callable that runs once in each worker interpreter, which is the designated place for per-worker warm-up rather than doing it lazily inside every task. ### Cost two: duplicated module state, and the configuration trap Every module-level object exists once per worker. A cached lookup table, a compiled pattern set, a memoised reference map: N workers, N copies, N times the memory. That changes how you size `max_workers` — with a thread pool the cache is shared and worker count is nearly free memory-wise; here each worker carries the full weight. The sharper failure is configuration. Consider a payment reconciliation job that, at startup in the parent, reads an environment setting and rebinds a module-level amount-format constant to a grouped, locale-dependent form. On a thread pool that just works: one interpreter, one module object, everybody sees the mutation. Move to an interpreter pool and the workers import that module fresh and get its **import-time default** instead. Nothing raises. The job runs, and at a 1,200-request-per-minute peak it writes hundreds of thousands of amounts in the wrong format, discovered downstream by whoever consumes the file. ```python from concurrent.futures import InterpreterPoolExecutor SEEN = 0 def count_row(): global SEEN SEEN += 1 return SEEN if __name__ == "__main__": with InterpreterPoolExecutor(max_workers=1) as pool: print([pool.submit(count_row).result() for _ in range(3)]) print("main interpreter still sees:", SEEN) ``` The worker counts 1, 2, 3 in its own copy of the module global; the parent still reads 0. Any design that treats a module global as shared mutable state is broken by the move, and it is broken silently. The fix is a discipline, not a workaround: **make configuration an explicit argument**, part of the task's input rather than ambient module state. Anything that must be established per worker goes in the executor's `initializer`. Anything that is genuinely global — a feature flag, a format setting — is read from process-level state such as `os.environ`, which *is* shared, or passed with every task. ### Cost three: extensions that refuse to load A native extension module has to opt in to being loaded in more than one interpreter. One that has not simply refuses, and the import raises an error saying that the module does not support loading in subinterpreters. In the stdlib, `readline` is an example. In the wider ecosystem, plenty of compiled packages are still in the process of adding support. This is the constraint that decides adoption, and it is discovered late if you do not go looking: the pool works fine in a unit test that imports nothing heavy, and fails the first time a real task pulls in the dependency. Test the actual import graph in a real subinterpreter before you commit. ### Diagnosing it in a running service Symptoms cluster. Throughput that improves far less than the worker count suggests, with CPU low: suspect import time per worker — have a task report how long its first import took, or the size of its `sys.modules`. Memory that scales linearly with `max_workers`: suspect duplicated module-level caches. Results that are subtly wrong only under the pool, never under a thread pool or in tests: suspect configuration that the parent mutated at runtime and the workers never saw. Import errors under load but never in development: suspect an extension without subinterpreter support on a code path only production reaches. ### The framing to offer Per-worker module state is the price of the per-interpreter GIL, and it is the same price a process pool charges — you were already paying it there. What is new is the temptation to forget, because the pool lives in one process and *looks* like a thread pool. The senior habit is to treat every task as if it ran in a fresh interpreter, because it does.

  • How would you keep per-worker import cost from eating the parallelism you gained?
    Measure it first: have a task report how long its imports took and how large its `sys.modules` is. Then keep the worker-side graph lean — import inside the task function rather than pulling a package's whole surface, and split heavy optional dependencies out of the module the tasks touch. Reuse one long-lived pool instead of creating one per request, and do warm-up once per worker through the executor's `initializer` rather than lazily inside every task.
  • A job produces subtly wrong output only under the interpreter pool. Where do you look first?
    At any state the parent mutates at runtime and the workers are assumed to inherit: a module-level constant rebound at startup, a registry populated by decorators that only run on an import path the workers never take, a cache seeded before the pool was created. Workers see import-time defaults. Confirm by having a task return the value it actually sees, then move that configuration into the task's arguments or the `initializer`.
  • What process-level state do interpreter pool workers still share?
    Everything below Python: the PID, open file descriptors, the current working directory, and `os.environ` — an environment variable set in the parent before a worker starts is visible there. That is a legitimate escape hatch for genuinely global configuration, and it is also why there is no fault isolation: a segfault in one worker's native code takes down the entire process, unlike a process pool where the parent survives.
  • How do you find out whether your dependencies work in a subinterpreter before committing?
    Exercise the real import graph inside one. Create an interpreter with `concurrent.interpreters.create()` and run `Interpreter.exec()` on the actual imports your tasks make, in CI, on the same platform and wheels production uses. Modules without support fail loudly with an import error naming the module, so a short test that imports each top-level dependency catches the blockers before a service depends on the pool.

saying these in an interview costs you the question

  • Assuming workers inherit the parent's imports
  • Treating a module-level cache as shared across workers
  • Sizing max_workers without counting duplicated module state
  • Configuring by mutating a module global at startup
  • Believing every compiled dependency loads in a subinterpreter
  • Expecting a worker crash to leave the parent process alive

context