skip to content

Worker Process Models

Pre-fork, threaded and async workers as a design choice: how much concurrency each holds, what each costs in memory, and which one a blocking database driver or a CPU-heavy handler forces on you.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why do pre-fork Python application servers run several worker processes instead of one?

level: juniorimportance: must knowfreq 70%

answer

  1. One process, one core
  2. The interpreter lock is per process
  3. Every worker owns its own interpreter
  4. Parent listens, children accept
  5. Isolation paid for in duplicated memory

basics

~20 s

Because one CPython process executes Python bytecode on one core at a time under the global interpreter lock. Forking several workers, each with its own interpreter and its own lock, puts real work on every core and isolates a crash.

solid answer

~50 s

A pre-fork server binds and listens on one socket, then forks N worker processes that each accept connections on that same inherited listening socket; the kernel hands each arriving connection to one of them. The reason for N rather than 1 is CPython's global interpreter lock: in the standard 3.14 build only one thread per process runs Python bytecode at a time, so one worker saturates at most one core no matter how much traffic arrives. Separate processes each carry their own interpreter and their own lock, so N workers can use N cores. Processes also buy fault isolation — a crash in a C extension, or one worker killed for using too much memory, takes down that worker only and the supervisor replaces it. The price is memory: every worker holds its own interpreter state, its own imported module objects and its own connection pool.

code

python · 23 lines
python
import os
import socket

listener = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
listener.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
listener.bind(("127.0.0.1", 0))
listener.listen(16)
port = listener.getsockname()[1]

for _ in range(3):
    if os.fork() == 0:
        conn, _addr = listener.accept()
        conn.sendall(str(os.getpid()).encode())
        conn.close()
        os._exit(0)

for _ in range(3):
    with socket.create_connection(("127.0.0.1", port)) as client:
        print("served by worker", client.recv(64).decode())

for _ in range(3):
    os.wait()
listener.close()

go deeper

for a junior

Be ready to say in one sentence that a standard CPython process runs Python bytecode on one core at a time, and that a pre-fork server therefore runs several worker processes. Know that each worker is a separate interpreter with separate memory.

for a middle

Explain the mechanics: the parent binds and listens, the children inherit that socket and all call accept on it, and the kernel picks the winner. Be able to state what the model costs in RAM and why copy-on-write does not stay free in CPython.

for a senior

Show you have operated it: worker count against downstream connection limits, resident memory per worker measured rather than guessed, and the reasoning for adding threads or coroutines inside each worker instead of adding more workers for I/O-bound handlers.

for a principal

Own the tradeoff between blast radius and memory budget across a fleet. Be ready to argue when process isolation is worth several times the RAM, and to give a considered position on whether the free-threaded 3.14 build changes that calculus for your stack yet.

## The constraint that forces the shape CPython protects its interpreter state with a **global interpreter lock**, and in the default build of Python 3.14 exactly one thread per interpreter executes Python bytecode at any instant. That is a **per-process property**. A worker process can start twenty threads, and those threads will still hand the lock back and forth rather than run bytecode side by side. So a single worker converts, at best, one core's worth of Python work per unit of time. On a machine with many cores, one worker leaves almost all of them idle for handlers that actually compute. A pre-fork server sidesteps the lock the crudest and most reliable way available: it **replicates the process**. Each forked worker is a full CPython instance with its own object graph, its own reference counts and its own interpreter lock, so N workers really do execute Python on N cores. Nothing has to be thread-safe across workers, because nothing is shared at the Python level — the workers cannot see each other's objects at all. ## How a request finds a worker The distinctive move in the pre-fork model is that the *parent* creates the listening socket — bind and listen — *before* forking. - Each child inherits that same open socket, so every worker sits in `accept()` on one shared queue of pending connections. - The kernel picks which waiting worker gets each connection; the application does not route anything, and there is no dispatcher process in the data path to become a bottleneck. This is also why the model **degrades gracefully**: if one worker is busy for a second, the others keep accepting, and the connection backlog set by `listen()` absorbs the burst. ## What the extra workers cost **Memory**, mostly, and more of it than people expect. Right after `fork()` the child shares the parent's pages copy-on-write, so a freshly forked worker looks nearly free. But CPython writes a reference count into the header of every object it touches, and that write dirties the page — so the shared, read-only-looking pages holding your imported modules get privatised over the first minutes of traffic. Budget closer to "a full copy of the application heap per worker" than to "one copy plus deltas". On top of that, each worker opens its own database connections, its own caches and its own file handles, so downstream connection counts scale with worker count too. ## What the extra workers buy besides cores **Blast radius.** A segmentation fault in a native extension, an unrecoverable memory blow-up, or a wedged handler kills exactly one worker; the supervising parent notices the exit and forks a replacement, and the other workers never observed the failure. A threaded server in one process has no such boundary — a native crash takes every in-flight request with it. That isolation is often worth more in practice than the core count, especially for services that decode untrusted input. ## Where the reasoning stops applying If the handlers spend their time **waiting** — on a socket, a database, an object store — they are not holding the interpreter lock while they wait, because CPython releases it around blocking I/O calls. For that workload the constraint is not cores at all, it is how many waiting requests one process can hold, and adding threads or an event loop inside each worker is the cheaper answer than adding whole processes. Real deployments usually combine the two: - a modest number of pre-forked workers for cores and isolation, - and threads or coroutines inside each for I/O concurrency. ## Two version-specific notes for 3.14 1. First, the **free-threaded build** (PEP 779) became officially supported in 3.14 and removes the interpreter lock entirely, which changes this arithmetic — but it is a distinct build, it costs roughly 5–10% single-threaded performance, and it requires every native extension you load to support it, so it does not make pre-forking obsolete today. 2. Second, if you spawn workers through the `multiprocessing` module rather than a raw `os.fork`, note that 3.14 changed the default start method on Unix platforms other than macOS to `forkserver`; macOS and Windows default to `spawn`, and plain `fork` must now be requested explicitly. A server that relied on children silently inheriting the parent's already-initialised state needs to know which start method it is actually getting.

  • If the interpreter lock is the reason, why not just start more threads inside one worker process?
    Threads inside one process share that one lock, so CPU-bound handlers still serialise — they take turns rather than run together. Threads do help when handlers wait on I/O, because CPython releases the lock around blocking calls, and they are far cheaper in memory than processes. So threads solve the waiting problem, not the core-count problem, which is why the two are usually combined rather than substituted.
  • What does each additional worker actually cost in memory?
    Close to a full copy of the application. Pages are shared copy-on-write immediately after the fork, but CPython writes a refcount into every object header it touches, so the pages holding imported modules get privatised as traffic flows. Each worker also holds its own connection pool, caches and file handles — which means downstream connection counts scale with worker count, not just RAM.
  • With several workers blocked in accept on the same socket, what decides which one gets a given connection?
    The kernel does, not your code. Modern kernels wake a single waiting acceptor per connection rather than every one of them, so a shared listening socket does not produce a stampede. The consequence for the application is that you cannot route a request to a chosen worker or assume any affinity — per-worker in-memory state is unreachable to the other workers and must not be treated as a cache the whole service shares.

It is a ticket counter with several clerks sharing one queue: nobody assigns customers to clerks, whichever clerk is free takes the next person — but each clerk needs their own desk, terminal and stack of forms.

saying these in an interview costs you the question

  • Says more worker processes always mean more throughput, whatever the workload
  • Thinks the interpreter lock is shared between separate processes
  • Claims workers keep sharing Python objects after the fork
  • Assumes ten workers cost about as much memory as one
  • Believes one crashed worker takes the whole service down
  • Thinks a dispatcher process hands each connection to a chosen worker

context

open as a page

How do Python's pre-fork, threaded and async worker models differ in concurrency held and memory cost?

level: middleimportance: must knowfreq 65%

basics

~20 s

A pre-fork worker holds one request and costs a whole interpreter. A thread holds one request and costs a stack. An async worker holds thousands of waiting connections in one thread — but only while nothing blocks.

open as a page

A chat-transcript archiver on async workers spikes latency on every open connection whenever one upload hits a slow synchronous decode. Why?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each async worker runs its coroutines on one thread with cooperative scheduling. A synchronous decode never yields, so the event loop cannot resume anything else until it finishes, and every connection that worker holds waits behind it.

open as a page

How would you choose a worker model for a new Python service, and what actually forces the decision?

level: principalimportance: should knowfreq 45%

basics

~20 s

Characterise the handlers — waiting or computing, how many concurrent connections, how large the memory budget — then let the driver stack decide. A synchronous driver rules out async workers; heavy computation rules out a single shared loop or thread pool.

open as a page