skip to content

How would you choose a worker model for a new Python service, and what actually forces the decision?

level: principalimportance: should knowfreq 45%

answer

  1. Measure the handler before ranking models
  2. Simultaneous connections, not requests per second
  3. The dependency stack usually decides
  4. Blast radius argues for some pre-forking
  5. Changing the choice later is a rewrite

basics

~20 s

Characterise the handlers — waiting or computing, how many concurrent connections, how large the memory budget — then let the driver stack decide. A synchronous driver rules out async workers; heavy computation rules out a single shared loop or thread pool.

solid answer

~50 s

I start from three measurements, not from preference: the ratio of wait time to CPU time in a typical handler, the number of simultaneous connections the service must hold, and the memory budget per instance. Those pick a shape — pre-forked workers for cores and blast radius, threads inside them for ordinary blocking code, an event loop inside them when concurrency must reach tens of thousands. But the decision is usually *forced* by the dependency stack: if any driver on the request path is synchronous, async workers give you cooperative scheduling's constraints without its benefit. I also weigh what the choice costs the team — an async codebase colours every function and every library it may use, and reversing it later is a rewrite. My default is the boring one, threaded workers under a small pre-fork.

go deeper

for a junior

Know that the choice exists and that it is driven by what handlers do: mostly waiting, or mostly computing. Remember that a blocking library on the request path is the thing that rules async out.

for a middle

Be able to argue a concrete recommendation from a handler profile: wait-to-CPU ratio, simultaneous connections, memory per instance. Explain why the usual answer nests a thread pool or an event loop inside a few pre-forked processes.

for a senior

Demonstrate that you have checked the whole dependency set for blocking calls, sized worker counts against downstream connection limits, and instrumented the model you chose — loop lag or pool saturation — so its failure mode is visible before customers find it.

for a principal

Own the decision as a long-lived commitment: what it forbids across the codebase, what reversing it would cost, when to split a mixed workload into two deployments instead, and how you would pilot free-threading or multiple interpreters rather than bet the fleet on them.

## Start by measuring the handler, not by ranking the models Three numbers decide most of it. 1. First, the **split between waiting and computing** in a typical request: profile it, because intuition is wrong here surprisingly often, and a handler that looks like pure I/O frequently spends a third of its time serialising and validating. 2. Second, the **peak number of *simultaneous* connections** — not requests per second, which is a different quantity: a service with modest throughput but tens of thousands of long-lived idle sockets has a completely different answer from a high-throughput service with short requests. 3. Third, the **memory budget per instance**, since that sets how many processes you can afford before you are buying capacity rather than choosing an architecture. ## Then admit what actually forces the decision: the drivers The theory says async workers hold the most concurrency per byte. Reality says a model is only usable if every library on the request path cooperates. One synchronous datastore driver, one blocking client for an internal service, one file-format library with a synchronous read — and the async model degrades into an event loop that offloads to a thread pool, which is the threaded model with extra constraints and a new failure mode when the pool saturates. So the first question is not "which model is best" but "which models can my dependency set actually support today". Frequently only one can, and the design is settled before any tradeoff is weighed. ## Blast radius is the second forcing function Multiple processes are the only structure here with a **hard boundary**: a native crash, an unbounded allocation, or a wedged handler takes one worker and the supervisor replaces it. For services that decode untrusted input, or that load native extensions of uncertain quality, that boundary is worth paying several times the memory for — and it is the reason almost every production Python deployment has *some* pre-forking in it, whatever runs inside each worker. The corollary is that the models are not alternatives to be selected among; the real decision is **what to nest inside a small number of processes**. ## Weigh the costs that do not appear in a benchmark - An **async** codebase is a commitment that spreads: async colours call chains, constrains which libraries may ever be added, and demands a discipline every future contributor must hold. - **Threads** demand a different discipline — shared mutable state, per-request context, and lock ordering — which is more familiar but easier to get subtly wrong. - **Processes** demand the least discipline and the most memory, and they push state you might have kept in a local cache into a shared store instead. Also count the cost of *changing* the choice: moving a mature service from threaded to async is a rewrite of every handler and every dependency decision, so this is a decision to make with the ten-year version of the service in mind, not the prototype. ## How I would decide, concretely - If the handlers wait and the drivers are ordinary blocking libraries: pre-forked workers with a **thread pool** inside each, sized against downstream connection limits. - If the service holds very many mostly-idle connections and the whole stack is non-blocking: pre-forked workers with an **event loop** inside each, plus a loop-lag metric from day one so the first blocking call is caught in a deploy rather than an incident. - If handlers do real computation: keep it **off the request path** entirely — a queue and separate consumers — rather than trying to make any worker model absorb it. Mixed workloads should be split into separate deployments before they are made to share a model, because a service that must be both is a service where every model is the wrong one. ## What is genuinely new in 3.14, and how much to bet on it **Free-threading** is officially supported (PEP 779), which for the first time makes threads inside one process a real answer to CPU concurrency — at roughly 5–10% single-threaded cost and conditional on every native extension supporting it. **Multiple interpreters** reached the stdlib (PEP 734) with `concurrent.interpreters` and an `InterpreterPoolExecutor`, offering isolation closer to a process at a lighter footprint. Both are real and both are young. The principled position is to design so that either could be adopted — keep handlers free of process-global mutable state, keep the worker model a deployment concern rather than a shape baked into every module — and to run one of them behind a flag on a non-critical service before making a fleet-wide bet.

  • A team proposes rewriting a threaded service as async to reduce its memory footprint. How do you evaluate that?
    By asking what the memory is actually spent on. If it is thread stacks for thousands of mostly-idle connections, the argument is real. If it is application heap replicated across processes, or caches, async changes nothing and the rewrite buys a constraint for free. I would also price the migration honestly — every handler and every dependency — and check that no driver on the path is synchronous before the work starts.
  • How do you keep the worker-model choice reversible as the codebase grows?
    Keep handlers free of process-global mutable state and of assumptions about which worker serves a request; put shared state in an external store rather than a per-worker cache; keep blocking and computing work behind a small number of well-marked seams rather than scattered through business logic. Then changing model is a change to those seams and to deployment configuration, not to every module.
  • What would make you split one service into two deployments rather than pick a single worker model?
    A genuine mix: long-lived idle connections alongside CPU-heavy handlers. One model cannot serve both without the heavy work degrading the light path — in an async worker it freezes the loop, in a threaded one it holds the interpreter lock. Splitting lets each half take the model that suits it and sizes them independently, at the cost of one more deployable and a queue between them.

saying these in an interview costs you the question

  • Picks async by default because it sounds more modern
  • Ignores whether the available drivers are blocking
  • Sizes on requests per second when connections are what is held
  • Treats the models as exclusive rather than nested
  • Assumes a threaded-to-async migration is a small refactor
  • Standardises the fleet on free-threading without piloting it

context