How do Python's pre-fork, threaded and async worker models differ in concurrency held and memory cost?
answer
- Ask what one worker holds at once
- Price per held request differs hugely
- Process, stack, or a coroutine frame
- Waiting overlaps; bytecode takes turns
- Cooperative model forbids blocking calls
basics
~20 sA pre-fork worker holds one request and costs a whole interpreter. A thread holds one request and costs a stack. An async worker holds thousands of waiting connections in one thread — but only while nothing blocks.
solid answer
~50 sThe three models buy concurrency at very different unit prices. **Pre-fork**: one in-flight request per worker, and each worker is a full CPython process — own interpreter, own module objects, own pools — so concurrency lands in the tens and memory scales roughly linearly. **Threaded**: one in-flight request per thread, at a stack plus per-thread state, so hundreds to low thousands are practical; but the interpreter lock means only one thread runs Python bytecode at a time, so threads overlap *waiting*, not computing. **Async**: one thread runs an event loop over many coroutines, each costing kilobytes rather than a stack, so tens of thousands of idle connections are realistic — provided every handler yields at every wait. One synchronous call stalls the whole worker, so a blocking driver rules the model out. Production stacks usually nest these: a few pre-forked workers, each threaded or async inside.
code
python · 17 linesimport asyncio
import time
async def other_connection():
start = time.perf_counter()
await asyncio.sleep(0.1)
return time.perf_counter() - start
async def slow_handler():
time.sleep(0.5)
return "archived"
async def main():
waited, _ = await asyncio.gather(other_connection(), slow_handler())
print(f"a 0.1s await actually took {waited:.2f}s")
asyncio.run(main())go deeper
Know the three names and the one-line difference: a process per request, a thread per request, or many coroutines sharing one thread. Remember that an async handler must never call something that blocks.
Be able to price each model: what a worker, a thread and a suspended coroutine each cost, and roughly what concurrency each reaches. Explain why threads help waiting handlers but not computing ones, and why a synchronous driver rules out the async model.
Demonstrate that you measure rather than assume: resident memory per worker, in-flight requests per worker under real traffic, and where a synchronous call sneaked into an async path. Justify the nested shape you run — several pre-forked workers, each threaded or async inside.
Own the choice as an architectural commitment. Weigh what the async model forbids across the whole dependency set, the migration cost of changing model later, and whether free-threading or multiple interpreters in 3.14 shift the arithmetic enough to bet on.
Three models, three answers to one question: *how many requests can one worker hold at once, and what does each held request cost?* ## Pre-fork — one request per process A worker is a whole CPython interpreter. Concurrency per worker is exactly one; total concurrency is the worker count. The unit cost is the largest of the three: a private interpreter, a private copy of every imported module object, private connection pools and caches. Pages start shared copy-on-write after the fork and get privatised as refcount writes touch object headers, so the honest budget is close to a full application heap per worker. In exchange you get the only model that puts Python bytecode on more than one core in the default 3.14 build, and the only one with a **hard fault boundary**: a native crash or a runaway allocation kills one worker and nothing else. - **Concurrency ceiling:** tens. - **Cost per unit of concurrency:** highest. ## Threaded — one request per thread Inside one worker, a pool of threads each handle a request end to end. A thread costs a stack (megabyte-scale virtual reservation, less resident) plus its per-thread interpreter state, so a few hundred to a couple of thousand is the practical range before context switching and memory dominate. The decisive detail is what threads actually overlap. CPython releases the global interpreter lock around blocking I/O calls and inside native code that chooses to release it, so threads waiting on sockets or a database overlap perfectly — that is the whole point. Threads running *Python* do not overlap; they take turns. So a threaded worker multiplies **I/O concurrency** and does nothing for CPU-bound handlers. The compensating advantage is enormous in practice: **threads accept ordinary blocking code**. Any synchronous driver, any library that was never written with concurrency in mind, works unchanged. - **Concurrency ceiling:** hundreds to low thousands. - **Cost per unit:** moderate. ## Async — many requests per thread One thread runs an **event loop**; each in-flight request is a coroutine suspended at an `await`, holding only its own frame and locals. That is kilobytes, not a stack, and there is no kernel context switch to resume one, so tens of thousands of mostly-idle connections in a single worker are ordinary. This is the right model for long-lived, low-traffic-per-connection sockets — streaming, websockets, fan-out to slow upstreams — where the other two models would be paying a process or a stack per idle peer. The cost is a hard rule about the code you may run: **every wait must be an `await`**. The loop is cooperative, so a coroutine keeps the thread until it suspends voluntarily. A synchronous call in a handler — a blocking driver, a `time.sleep`, a large synchronous decode, a CPU-heavy loop — freezes *every* connection that worker holds, not just its own. This is what "a blocking driver rules out the async model" means concretely: if the only driver available for your datastore is synchronous, either you wrap every call in `asyncio.to_thread` (and are then paying for threads anyway, with an event loop bolted on top) or you should not have chosen async workers. - **Concurrency ceiling:** tens of thousands. - **Cost per unit:** lowest — and paid for in a constraint on the code rather than in memory. ## Why real deployments combine them None of the three covers both axes. **Processes** give cores and isolation but no cheap concurrency; **threads and coroutines** give cheap concurrency but no extra cores. So the standard shape is a small number of pre-forked workers — enough to occupy the cores and to contain a crash — with threads or an event loop *inside* each worker for concurrency. Choosing the inner model is mostly a question about your drivers, not about theory. ## What 3.14 changes - The **free-threaded build** is officially supported (PEP 779) and drops the interpreter lock, so threads in one process can execute Python on several cores; it costs roughly 5–10% single-threaded and requires native extensions that support it, so treat it as an option to evaluate rather than a default. - The stdlib also gained **multiple interpreters** (PEP 734): `concurrent.interpreters` and an `InterpreterPoolExecutor` give per-interpreter isolation with a lighter footprint than a process, a fourth point on this spectrum that is still young. - And if you spawn workers via `multiprocessing`, the default start method on Unix other than macOS is now `forkserver` rather than `fork`.
- If threads cannot run Python bytecode simultaneously, why does a threaded worker still improve throughput?Because most request time is waiting, not computing. CPython releases the interpreter lock around blocking I/O calls, so a thread parked on a socket read holds nothing and the others proceed. Throughput rises in proportion to how much of a request is wait time; for a handler that is pure Python computation it does not rise at all, and the extra threads only add switching overhead.
- When would you pick async workers over threaded ones even though both target I/O-bound work?When the concurrency you need exceeds what stacks can pay for — tens of thousands of mostly-idle long-lived connections, heavy fan-out to slow upstreams, or streaming where each peer costs almost nothing between messages. The precondition is a fully non-blocking stack: async-native drivers and no synchronous work inside handlers. Without that, threads are the safer choice at a fraction of the migration cost.
- What does a CPU-heavy handler do to each of the three models?In a pre-fork worker it occupies that worker and nothing else — the other workers keep serving on other cores. In a threaded worker it holds the interpreter lock and starves the other threads in that process. In an async worker it is the worst case: it blocks the loop, so every connection that worker holds stalls. Heavy CPU work belongs in a process pool, not in a request handler.
Three ways to serve eleven diners: eleven separate kitchens, one kitchen with eleven cooks sharing one stove, or one cook who starts every dish and returns whenever something needs stirring — fine until one dish demands undivided attention.
saying these in an interview costs you the question
- Says async workers make CPU-bound handlers faster
- Claims threads in one process run Python bytecode in parallel by default
- Treats a coroutine as costing about the same as a thread
- Thinks a blocking call only delays its own request
- Assumes you must pick exactly one model rather than nesting them
- Quotes an async connection ceiling without checking the drivers are non-blocking