skip to content

Long-Lived Service Patterns

What it takes to run the same interpreter for weeks: the calling convention a server speaks to your app, the worker model underneath it, restarts that lose no work, and one instance of a cron job.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

20

Why do pre-fork Python application servers run several worker processes instead of one?

level: juniorimportance: must knowfreq 70%

answer

  1. One process, one core
  2. The interpreter lock is per process
  3. Every worker owns its own interpreter
  4. Parent listens, children accept
  5. Isolation paid for in duplicated memory

basics

~20 s

Because one CPython process executes Python bytecode on one core at a time under the global interpreter lock. Forking several workers, each with its own interpreter and its own lock, puts real work on every core and isolates a crash.

solid answer

~50 s

A pre-fork server binds and listens on one socket, then forks N worker processes that each accept connections on that same inherited listening socket; the kernel hands each arriving connection to one of them. The reason for N rather than 1 is CPython's global interpreter lock: in the standard 3.14 build only one thread per process runs Python bytecode at a time, so one worker saturates at most one core no matter how much traffic arrives. Separate processes each carry their own interpreter and their own lock, so N workers can use N cores. Processes also buy fault isolation — a crash in a C extension, or one worker killed for using too much memory, takes down that worker only and the supervisor replaces it. The price is memory: every worker holds its own interpreter state, its own imported module objects and its own connection pool.

code

python · 23 lines
python
import os
import socket

listener = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
listener.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
listener.bind(("127.0.0.1", 0))
listener.listen(16)
port = listener.getsockname()[1]

for _ in range(3):
    if os.fork() == 0:
        conn, _addr = listener.accept()
        conn.sendall(str(os.getpid()).encode())
        conn.close()
        os._exit(0)

for _ in range(3):
    with socket.create_connection(("127.0.0.1", port)) as client:
        print("served by worker", client.recv(64).decode())

for _ in range(3):
    os.wait()
listener.close()

go deeper

for a junior

Be ready to say in one sentence that a standard CPython process runs Python bytecode on one core at a time, and that a pre-fork server therefore runs several worker processes. Know that each worker is a separate interpreter with separate memory.

for a middle

Explain the mechanics: the parent binds and listens, the children inherit that socket and all call accept on it, and the kernel picks the winner. Be able to state what the model costs in RAM and why copy-on-write does not stay free in CPython.

for a senior

Show you have operated it: worker count against downstream connection limits, resident memory per worker measured rather than guessed, and the reasoning for adding threads or coroutines inside each worker instead of adding more workers for I/O-bound handlers.

for a principal

Own the tradeoff between blast radius and memory budget across a fleet. Be ready to argue when process isolation is worth several times the RAM, and to give a considered position on whether the free-threaded 3.14 build changes that calculus for your stack yet.

## The constraint that forces the shape CPython protects its interpreter state with a **global interpreter lock**, and in the default build of Python 3.14 exactly one thread per interpreter executes Python bytecode at any instant. That is a **per-process property**. A worker process can start twenty threads, and those threads will still hand the lock back and forth rather than run bytecode side by side. So a single worker converts, at best, one core's worth of Python work per unit of time. On a machine with many cores, one worker leaves almost all of them idle for handlers that actually compute. A pre-fork server sidesteps the lock the crudest and most reliable way available: it **replicates the process**. Each forked worker is a full CPython instance with its own object graph, its own reference counts and its own interpreter lock, so N workers really do execute Python on N cores. Nothing has to be thread-safe across workers, because nothing is shared at the Python level — the workers cannot see each other's objects at all. ## How a request finds a worker The distinctive move in the pre-fork model is that the *parent* creates the listening socket — bind and listen — *before* forking. - Each child inherits that same open socket, so every worker sits in `accept()` on one shared queue of pending connections. - The kernel picks which waiting worker gets each connection; the application does not route anything, and there is no dispatcher process in the data path to become a bottleneck. This is also why the model **degrades gracefully**: if one worker is busy for a second, the others keep accepting, and the connection backlog set by `listen()` absorbs the burst. ## What the extra workers cost **Memory**, mostly, and more of it than people expect. Right after `fork()` the child shares the parent's pages copy-on-write, so a freshly forked worker looks nearly free. But CPython writes a reference count into the header of every object it touches, and that write dirties the page — so the shared, read-only-looking pages holding your imported modules get privatised over the first minutes of traffic. Budget closer to "a full copy of the application heap per worker" than to "one copy plus deltas". On top of that, each worker opens its own database connections, its own caches and its own file handles, so downstream connection counts scale with worker count too. ## What the extra workers buy besides cores **Blast radius.** A segmentation fault in a native extension, an unrecoverable memory blow-up, or a wedged handler kills exactly one worker; the supervising parent notices the exit and forks a replacement, and the other workers never observed the failure. A threaded server in one process has no such boundary — a native crash takes every in-flight request with it. That isolation is often worth more in practice than the core count, especially for services that decode untrusted input. ## Where the reasoning stops applying If the handlers spend their time **waiting** — on a socket, a database, an object store — they are not holding the interpreter lock while they wait, because CPython releases it around blocking I/O calls. For that workload the constraint is not cores at all, it is how many waiting requests one process can hold, and adding threads or an event loop inside each worker is the cheaper answer than adding whole processes. Real deployments usually combine the two: - a modest number of pre-forked workers for cores and isolation, - and threads or coroutines inside each for I/O concurrency. ## Two version-specific notes for 3.14 1. First, the **free-threaded build** (PEP 779) became officially supported in 3.14 and removes the interpreter lock entirely, which changes this arithmetic — but it is a distinct build, it costs roughly 5–10% single-threaded performance, and it requires every native extension you load to support it, so it does not make pre-forking obsolete today. 2. Second, if you spawn workers through the `multiprocessing` module rather than a raw `os.fork`, note that 3.14 changed the default start method on Unix platforms other than macOS to `forkserver`; macOS and Windows default to `spawn`, and plain `fork` must now be requested explicitly. A server that relied on children silently inheriting the parent's already-initialised state needs to know which start method it is actually getting.

  • If the interpreter lock is the reason, why not just start more threads inside one worker process?
    Threads inside one process share that one lock, so CPU-bound handlers still serialise — they take turns rather than run together. Threads do help when handlers wait on I/O, because CPython releases the lock around blocking calls, and they are far cheaper in memory than processes. So threads solve the waiting problem, not the core-count problem, which is why the two are usually combined rather than substituted.
  • What does each additional worker actually cost in memory?
    Close to a full copy of the application. Pages are shared copy-on-write immediately after the fork, but CPython writes a refcount into every object header it touches, so the pages holding imported modules get privatised as traffic flows. Each worker also holds its own connection pool, caches and file handles — which means downstream connection counts scale with worker count, not just RAM.
  • With several workers blocked in accept on the same socket, what decides which one gets a given connection?
    The kernel does, not your code. Modern kernels wake a single waiting acceptor per connection rather than every one of them, so a shared listening socket does not produce a stampede. The consequence for the application is that you cannot route a request to a chosen worker or assume any affinity — per-worker in-memory state is unreachable to the other workers and must not be treated as a cache the whole service shares.

It is a ticket counter with several clerks sharing one queue: nobody assigns customers to clerks, whichever clerk is free takes the next person — but each clerk needs their own desk, terminal and stack of forms.

saying these in an interview costs you the question

  • Says more worker processes always mean more throughput, whatever the workload
  • Thinks the interpreter lock is shared between separate processes
  • Claims workers keep sharing Python objects after the fork
  • Assumes ten workers cost about as much memory as one
  • Believes one crashed worker takes the whole service down
  • Thinks a dispatcher process hands each connection to a chosen worker

context

open as a page

Why does preloading a Python app before forking workers save less memory than copy-on-write suggests?

level: middleimportance: must knowfreq 55%

basics

~20 s

Copy-on-write shares pages only until something writes to one. CPython stores a reference count inside every object header, so merely touching a preloaded object dirties its page and the sharing quietly decays over the worker's lifetime.

open as a page

How do you use fcntl.flock to guarantee only one copy of a Python job runs?

level: middleimportance: must knowfreq 40%

basics

~20 s

Open a fixed lockfile with os.open, then call fcntl.flock(fd, fcntl.LOCK_EX | fcntl.LOCK_NB). A BlockingIOError means another copy holds it, so exit. Keep the descriptor open for the whole run; the kernel releases the lock when the process dies.

open as a page

Why should a long-lived Python service run in the foreground instead of double-fork daemonising itself?

level: middleimportance: must knowfreq 52%

basics

~20 s

A supervisor manages a process by being its parent. Staying in the foreground keeps the supervisor as the parent, so it knows the pid, reads the exit status, and captures whatever the process writes to sys.stdout and sys.stderr.

open as a page

How do Python's pre-fork, threaded and async worker models differ in concurrency held and memory cost?

level: middleimportance: must knowfreq 65%

basics

~20 s

A pre-fork worker holds one request and costs a whole interpreter. A thread holds one request and costs a stack. An async worker holds thousands of waiting connections in one thread — but only while nothing blocks.

open as a page

What does a PEP 3333 WSGI application callable receive, and what must it return?

level: middleimportance: must knowfreq 50%

basics

~20 s

A WSGI application is any callable of two arguments: environ, a dictionary of CGI-style request variables, and start_response, which the application calls with a status string and a list of header pairs. It returns an iterable of bytes.

open as a page

Why is checking for an existing pidfile with os.path.exists a racy single-instance guard?

level: juniorimportance: should knowfreq 30%

basics

~20 s

Between the os.path.exists check and the write, another copy can run the same check. Both see nothing, both create the file, both start. That gap is the race, and the file also outlives a killed process.

open as a page

How does a Python program set the exit status its supervisor sees when it fails?

level: juniorimportance: should knowfreq 46%

basics

~20 s

sys.exit(n) raises SystemExit, and the interpreter exits with that integer status. Falling off the end of the program exits 0; an uncaught exception prints a traceback and exits 1; sys.exit('message') prints the text to stderr and exits 1.

open as a page

What does a forked worker inherit from a preloaded parent, and which of it is unsafe to reuse?

level: middleimportance: should knowfreq 50%

basics

~20 s

The child gets a duplicate of the parent's memory and of every open file descriptor: connected sockets, database sessions, buffered file objects and any generator state built at import. The listening socket is meant to be shared; connected descriptors and pre-seeded generators are not.

open as a page

Why does PEP 3333 put latin-1 native strings in the WSGI environ but bytes in the body?

level: middleimportance: should knowfreq 32%

basics

~20 s

PEP 3333 makes environ values, the status line and headers native str decoded from the wire with latin-1, while request and response bodies stay bytes. Latin-1 round-trips every byte value, so no information is lost in the detour.

open as a page

Your preloaded metrics-scraper workers grow past their memory budget every day; how do you bound it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Cap each worker's lifetime. After a set number of scrapes, or above a resident-memory ceiling, the worker stops taking new work, finishes the batch in flight, and exits so the supervisor forks a replacement. Jitter the threshold, and remember recycling bounds a leak rather than fixing it.

open as a page

A document-conversion worker holding an fcntl.flock lock is SIGKILLed — what does the next run find?

level: seniorimportance: should knowfreq 24%

basics

~20 s

It finds the lock free. The kernel closes every descriptor of a dying process, which releases the advisory lock whether the exit was clean or a SIGKILL. The lockfile itself stays on disk; that is normal and needs no cleanup.

open as a page

How should a Python service under a supervisor signal that it is ready to take traffic?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Do the slow start-up work first and only then become reachable: create the listening socket last, or hold a readiness flag that the health path consults, or send the notification the supervisor is waiting for. Started is not the same as ready.

open as a page

Why should a supervised Python service leave restart backoff to the supervisor instead of retrying internally?

level: seniorimportance: should knowfreq 35%

basics

~20 s

A process that catches everything and sleeps stays alive while doing nothing, so the supervisor sees a healthy service and restart-based alerting never fires. Exiting non-zero makes the failure visible and lets the supervisor apply its own backoff.

open as a page

A chat-transcript archiver on async workers spikes latency on every open connection whenever one upload hits a slow synchronous decode. Why?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each async worker runs its coroutines on one thread with cooperative scheduling. A synchronous decode never yields, so the event loop cannot resume anything else until it finishes, and every connection that worker holds waits behind it.

open as a page

What does ASGI's scope, receive and send contract allow that a WSGI callable cannot?

level: seniorimportance: should knowfreq 44%

basics

~20 s

ASGI replaces WSGI's one synchronous call with an async callable taking scope, a dict describing the connection, plus receive and send awaitables for events. Those events allow awaiting I/O, incremental streaming both ways, WebSockets, and startup and shutdown hooks.

open as a page

How would you set preload and worker-recycling policy across a fleet of long-lived Python services?

level: principalimportance: should knowfreq 35%

basics

~20 s

Standardise the mechanism and the contract - what a preload may construct, how a worker drains before exiting, which metrics it emits - and leave the numbers per service, because a job limit follows that service's measured growth per job and its tolerance for a restart-time capacity dip.

open as a page

How would you choose a worker model for a new Python service, and what actually forces the decision?

level: principalimportance: should knowfreq 45%

basics

~20 s

Characterise the handlers — waiting or computing, how many concurrent connections, how large the memory budget — then let the driver stack decide. A synchronous driver rules out async workers; heavy computation rules out a single shared loop or thread pool.

open as a page

How do you run and conformance-check a WSGI application with only the standard library?

level: juniorimportance: nice to knowfreq 16%

basics

~10 s

The stdlib package wsgiref is the reference implementation: wsgiref.simple_server.make_server runs a single-threaded HTTP server around your callable, and wsgiref.validate.validator wraps an application so any PEP 3333 violation raises an assertion.

open as a page

What is the difference between fcntl.flock and fcntl.lockf in Python?

level: seniorimportance: nice to knowfreq 14%

basics

~20 s

fcntl.flock takes a whole-file lock owned by the open file description, so duplicated and inherited descriptors share it. fcntl.lockf takes a POSIX byte-range lock owned by the process, dropped when it closes any descriptor to that file.

open as a page