skip to content

Interpreter Pools and Queues

Multiple interpreters landed in the stdlib in 3.14: each keeps its own state, they hand data across a queue, and a pool executor drives them. Interviewers compare them to threads and processes.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How does `concurrent.futures.InterpreterPoolExecutor` differ from `ThreadPoolExecutor` for CPU-bound work?

level: middleimportance: must knowfreq 38%

answer

  1. Same Future API, different runtime underneath
  2. Threads share one lock; interpreters do not
  3. One process, several GILs
  4. Arguments and results must cross a boundary
  5. Closures are rejected outright

basics

~20 s

Both run workers as OS threads in one process, but ThreadPoolExecutor workers share one interpreter and one GIL, so pure-Python CPU work serializes. Each InterpreterPoolExecutor worker gets its own interpreter and its own GIL, so CPU work runs genuinely in parallel.

solid answer

~40 s

`concurrent.futures.InterpreterPoolExecutor`, added in Python 3.14 (PEP 734), still backs each worker with an OS thread in a single process — but each of those threads runs its **own** interpreter, and since 3.12 each interpreter holds its own GIL. So where a `ThreadPoolExecutor` serializes pure-Python CPU work behind one lock and only wins on I/O, an interpreter pool executes bytecode on several cores at once without paying for process creation. The price is isolation: the callable, its arguments and its result must be able to cross an interpreter boundary, so a closure over free variables is rejected with `NotShareableError`; workers share no module-level state, so each one re-runs its own imports; and communication goes through `concurrent.interpreters.create_queue()` rather than shared objects. `ProcessPoolExecutor` still wins where a dependency refuses to load in a subinterpreter.

code

python · 14 lines
python
from concurrent.futures import InterpreterPoolExecutor
import os


def settle(batch):
    import os
    return sum(batch), os.getpid()


if __name__ == "__main__":
    batches = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]
    with InterpreterPoolExecutor(max_workers=3) as pool:
        for total, pid in pool.map(settle, batches):
            print(total, "same process as main:", pid == os.getpid())

go deeper

for a junior

Learn the headline contrast: thread pool workers share one interpreter and one GIL, interpreter pool workers each get their own, so only the second one speeds up pure-Python CPU work.

for a middle

Be ready to explain the mechanics — one process, one OS thread per worker, one interpreter per thread, per-interpreter GIL since 3.12 — and to name what breaks: closures, shared module state, and objects that cannot cross the boundary.

for a senior

Show the judgement: characterise the workload first, then weigh per-worker import cost, memory multiplied by worker count, the absence of fault isolation, and whether every dependency actually loads in a subinterpreter before you commit a production service to it.

for a principal

Own the migration story. Argue when the operational simplicity of one process beats a mature process pool, what evidence would justify the switch, and how you would keep task boundaries clean enough that the choice stays reversible.

### The three executors, in one shape each `concurrent.futures` exposes the same `submit()` / `map()` / `Future` API over three very different runtimes: - **`ThreadPoolExecutor`** — N OS threads, one process, **one interpreter, one GIL**. Only one thread executes Python bytecode at a time. It wins on I/O-bound work, where the GIL is released while a thread waits on a socket or a file, and on calls into native code that release the GIL around a long computation. - **`ProcessPoolExecutor`** — N OS processes. Real parallelism, complete isolation, but process creation, a separate address space, and every argument and result crossing a serialization boundary. - **`InterpreterPoolExecutor`** (new in 3.14) — N OS threads in **one** process, each running its **own** interpreter with its **own** GIL. Real parallelism, no second process. ### Why the GIL boundary moved Until Python 3.12 the GIL was a single process-wide lock, so extra interpreters bought nothing for CPU-bound code. PEP 684 made the GIL per-interpreter in 3.12; PEP 734 then shipped the public `concurrent.interpreters` module and this executor in 3.14. The effect is directly measurable: a pure-Python arithmetic loop mapped over four workers finishes several times faster on an interpreter pool than on a thread pool of the same size, because the thread pool's workers are taking turns holding one lock while the interpreter pool's are not. ```python from concurrent.futures import InterpreterPoolExecutor def burn(n): return sum(i * i for i in range(n)) if __name__ == "__main__": with InterpreterPoolExecutor(max_workers=4) as pool: print(sum(pool.map(burn, [2_000_000] * 4))) ``` ### What you give up The isolation that buys the parallelism also constrains the API. **Everything must cross a boundary.** The callable, its arguments and its return value all move between interpreters. Simple immutables (`int`, `str`, `bytes`, tuples of those) cross efficiently; dicts, lists and ordinary instances are copied, so the worker gets an equivalent new object; and anything tied to interpreter or OS state — a `threading.Lock`, an open file, a generator, a socket — is refused with `concurrent.interpreters.NotShareableError`. A **closure over free variables** is refused too, which is the first thing people trip over when porting from a thread pool: `pool.submit(make_handler(rate), row)` works on threads and fails here. Pass the configuration as an argument instead. **No shared mutable state.** On a thread pool, workers naturally share a module-level cache, a counter, a connection pool. On an interpreter pool each worker has its own copy of every module's globals, so a counter incremented by workers is invisible in the main interpreter, and a warm cache is warmed N times. Cross-worker communication is explicit, through `concurrent.interpreters.create_queue()`. **Per-worker startup.** Creating an interpreter costs single-digit milliseconds, but each one then re-runs the entire import graph your tasks touch. For short tasks with a heavy dependency tree, that dominates. **Extension compatibility.** A native extension module that has not opted into subinterpreter support refuses to import in a worker, raising `ImportError`. That is the constraint most likely to rule the pool out for a given dependency set today. ### Choosing, concretely Take a payment reconciliation job that normalises, validates and hashes rows, peaking at 1,200 requests per minute. The work is pure-Python CPU: on a `ThreadPoolExecutor` the peak simply queues, because four workers share one GIL. `ProcessPoolExecutor` fixes throughput, at the cost of a full interpreter per process, per-process memory for every cached table, and every row batch serialized across a pipe. `InterpreterPoolExecutor` gets the same parallelism inside one process — one PID to supervise, one memory ceiling to reason about, cheaper handoff for the immutable payloads — provided every library the job imports loads cleanly in a subinterpreter and no task closes over shared objects. The decision rule is short. I/O-bound, or native code that releases the GIL: `ThreadPoolExecutor`. CPU-bound Python with dependencies that may not be subinterpreter-safe, or where a hard fault must not take the parent down: `ProcessPoolExecutor`. CPU-bound Python with a clean, self-contained task boundary and a dependency set you have verified: `InterpreterPoolExecutor`. ### The honest caveat This is new surface. `ProcessPoolExecutor` has been in the stdlib since 3.2 and every third-party library has been tested against it; subinterpreter support is still being added across the ecosystem. Say that out loud in an interview — it is the part a senior answer includes and a memorised answer does not.

  • When would you still pick `ProcessPoolExecutor` over `InterpreterPoolExecutor` on 3.14?
    When a dependency has not opted into subinterpreter support and refuses to import in a worker; when a task can hard-crash the runtime and you need the parent to survive, since interpreters share one process and no fault isolation; when tasks legitimately need separate OS-level state such as their own signal handling or working directory; and when you value the maturity of a pool that every library has been tested against.
  • Your tasks are I/O-bound. Does an interpreter pool help?
    Not meaningfully. I/O-bound work already releases the GIL while waiting, so a `ThreadPoolExecutor` scales fine and costs almost nothing per worker. An interpreter pool would add per-worker interpreter creation and a full re-import of the dependency graph for no throughput gain, and would take away the shared connection pools and caches that thread workers use naturally.
  • How do two workers in an `InterpreterPoolExecutor` share data with each other?
    They do not share objects — each worker has its own module globals, so there is no shared cache or counter. Explicit channels are the mechanism: create a queue with `concurrent.interpreters.create_queue()` and pass it in, since a queue is itself a cross-interpreter object. Anything else goes back through the `Future` returned by `submit()`, or through process-level state such as a file or a database.
  • Does an interpreter pool avoid serialization entirely?
    No. Immutable scalars and tuples cross efficiently and a `memoryview` genuinely shares its underlying buffer, but a dict, a list or an ordinary instance is copied — the receiving interpreter gets an equivalent new object, not the original. The saving over processes is that the copy happens in one address space rather than through a pipe, not that copying disappears.

saying these in an interview costs you the question

  • Saying InterpreterPoolExecutor spawns processes
  • Claiming a ThreadPoolExecutor parallelises pure-Python CPU work
  • Assuming workers share module-level caches or counters
  • Submitting a closure and expecting it to work
  • Treating it as a drop-in replacement for every process pool
  • Believing data crosses interpreters with no copy at all

context

open as a page

Which objects can cross a `concurrent.interpreters` queue, and which cannot?

level: middleimportance: should knowfreq 26%

basics

~20 s

Immutable scalars and tuples of them cross efficiently, and a memoryview genuinely shares its buffer. Dicts, lists and ordinary instances are copied, so the receiver gets an equivalent new object. Locks, sockets, generators, modules and closures raise NotShareableError.

open as a page

Why does every `concurrent.futures.InterpreterPoolExecutor` worker re-import your modules?

level: seniorimportance: should knowfreq 22%

basics

~20 s

Each worker runs its own interpreter with its own sys.modules, so no import from the parent is visible and every worker re-executes the import graph. Caches and configuration are per worker, and some native extensions refuse to load at all.

open as a page

What does `concurrent.interpreters.create()` return, and where does that interpreter run?

level: juniorimportance: nice to knowfreq 18%

basics

~20 s

It returns an Interpreter object: an additional Python interpreter living inside the same OS process, not a new process. It has its own sys.modules, its own module globals and its own GIL, and you drive it with Interpreter.exec().

open as a page