skip to content

Picking an Execution Model

Classify the workload first — CPU-bound, I/O-bound or connection-bound — then defend threads, processes, asyncio or a GIL-free build. Interviewers want a measurement behind the answer, not a slogan.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

5

Why does running a CPU-bound Python loop on threading.Thread workers not make it faster?

level: juniorimportance: must knowfreq 80%

answer

  1. One interpreter, one bytecode stream
  2. Threads take turns, not seats
  3. Waiting is fine; computing is not
  4. Extra cores need extra interpreters
  5. Free-threaded build supported in 3.14

basics

~20 s

In a standard CPython build the global interpreter lock lets only one thread execute Python bytecode at a time, so CPU-bound threads take turns rather than running in parallel. Move that work to separate processes instead.

solid answer

~40 s

CPython protects its internal state, including every object refcount, with one global interpreter lock (GIL), and a thread must hold it to run bytecode. Four `threading.Thread` workers spinning through a pure-Python loop therefore interleave on one core: total wall time stays roughly serial, plus a little overhead from handing the lock back and forth. Threads are still the right model when the work spends its time *waiting* — sockets, files, subprocesses — or inside a C extension that releases the lock around a long computation. To actually use several cores with pure Python you need several interpreters: a `ProcessPoolExecutor`, or the free-threaded build that Python 3.14 supports officially. The interview point is that the model follows the measurement: prove the loop is burning CPU in Python before blaming or replacing anything.

code

python · 22 lines
python
import time
from concurrent.futures import ThreadPoolExecutor


def burn(n):
    total = 0
    for i in range(n):
        total += i * i
    return total


if __name__ == "__main__":
    n = 3_000_000
    start = time.perf_counter()
    for _ in range(4):
        burn(n)
    print(f"serial    {time.perf_counter() - start:.2f}s")

    start = time.perf_counter()
    with ThreadPoolExecutor(max_workers=4) as pool:
        list(pool.map(burn, [n] * 4))
    print(f"4 threads {time.perf_counter() - start:.2f}s")

go deeper

for a junior

Recall the one-sentence rule and its exception: one thread runs Python bytecode at a time, but threads still overlap waiting. Be ready to say what you would use instead for CPU work.

for a middle

Explain the mechanics: the lock protects interpreter state, it is handed over on a switch interval, and blocking calls release it. Name processes, the free-threaded build and subinterpreters as the ways to get parallel bytecode.

for a senior

Show the measurement before the fix — serial versus threads versus processes on one unit of work — and be honest about what the process model costs in startup, memory and serialization once you recommend it.

for a principal

Own the policy: which of your services are allowed to assume the default build, what it would take to qualify the free-threaded build, and how you keep teams from reaching for processes reflexively where the real limit is elsewhere.

## What the lock actually is CPython keeps a lot of mutable interpreter state — reference counts on every object, the small-integer cache, interned strings, module dicts — and it protects all of it with a single mutex, the global interpreter lock. A thread must hold that lock to execute Python bytecode. The interpreter hands it around: the running thread drops it every few milliseconds (the switch interval, readable with `sys.getswitchinterval` and adjustable with `sys.setswitchinterval`) so another thread gets a turn, and it also drops it around calls that will block. The consequence is narrow but decisive. *Concurrency* — many things in flight — works fine with threads. *Parallel execution of Python bytecode* does not. Four threads each summing millions of integers do not finish in a quarter of the time; they finish in about the same wall time as running them one after another, sometimes slightly worse, because the handoffs are not free. ## Why threads are still worth having A thread that is waiting is not holding the lock. Every blocking call in the standard library releases it before it parks: socket reads and writes, `time.sleep`, file I/O, `subprocess` waits. So a `ThreadPoolExecutor` running fifty blocking HTTP-style requests genuinely overlaps them, and the speedup is real. The same is true of a C extension that releases the lock around a long computation — a compression codec or a numeric kernel written in C can run on several cores from several Python threads, because the heavy loop is not executing bytecode. That is why the honest classification question comes first. *Threads do not help* is wrong as a slogan; *threads do not help pure-Python CPU work in the default build* is the accurate version. ## What does give you cores There are three ways to get more than one Python bytecode stream running at once: 1. **More processes.** `concurrent.futures.ProcessPoolExecutor` or `multiprocessing` gives each worker its own interpreter and its own lock. The cost is startup and data transfer: arguments and results are pickled and pushed through a pipe, and each worker imports your modules again and pays for its own memory. 2. **The free-threaded build.** The experimental build first shipped in 3.13 and became officially supported in 3.14 (PEP 779). It is a separate interpreter binary with the lock removed, so plain threads scale — at the cost of roughly 5–10% single-threaded slowdown, extension compatibility work, and data races that the lock used to paper over. 3. **Multiple interpreters in one process.** Python 3.12 gave each subinterpreter its own lock, and 3.14 exposed the feature in the standard library through `concurrent.interpreters` and `concurrent.futures.InterpreterPoolExecutor`. Isolation is stronger than threads and startup is cheaper than a process, but objects still cannot be shared freely. ## Proving it before you change anything The cheap experiment is three timings of the same unit of work: serial, N threads, N processes. If threads are flat and processes scale close to linearly, the lock is your ceiling and the process model is the answer. If none of them help, the bottleneck is somewhere else — a slow algorithm, a remote service, or memory bandwidth — and swapping execution models will only add overhead. A second useful signal is the ratio of `time.process_time` to `time.perf_counter` for one unit of work: near 100% means the code is burning CPU in this process, near 0% means it is waiting and threads or asyncio are the right tools. ## The trap on the other side Juniors who learn this rule often over-rotate and reach for processes everywhere. Processes multiply memory, serialize every argument and result, and can turn a fast job into a slow one when the payload is large and the per-item work is small. The correct instinct is: classify the workload, then pick the cheapest model that removes the measured bottleneck.

  • When is threading.Thread still the right model in a standard CPython build?
    Whenever the work waits rather than computes: network calls, disk reads, subprocess waits, database round trips. Also when the expensive part lives in a C extension that releases the lock around a long call, since the extension is not running bytecode while it works. In both cases threads overlap real elapsed time and cost far less than processes.
  • What is the cheapest way to prove the lock is the bottleneck rather than something else?
    Time the same unit of work three ways: serial, N threads, N processes. Flat with threads and near-linear with processes points squarely at the lock. If processes do not help either, the limit is elsewhere — algorithmic cost, a remote dependency, or memory bandwidth — and changing execution model will only add overhead.
  • Does the free-threaded build simply make threaded Python faster?
    No. In 3.14 it is a separate, officially supported build that carries roughly a 5–10% single-threaded penalty, needs extensions built for it, and exposes real data races that the lock previously hid. It pays off when threads share a large in-memory structure that would be expensive to copy into processes; otherwise the default build plus processes is still the safer default.

The lock is a single microphone in a meeting room: any number of people can be in the room, but only one can speak at a time. Extra speakers only help if most of them are listening rather than talking.

saying these in an interview costs you the question

  • Says the lock makes Python single-threaded for everything, including I/O
  • Claims adding more threads will eventually beat the lock
  • Thinks asyncio parallelizes CPU-bound Python code
  • Believes multiprocessing is a free drop-in speedup with no transfer cost
  • Assumes free-threading is on by default in Python 3.14
  • Cannot name any workload where threads genuinely help

context

open as a page

How do you measure whether a Python workload is CPU-bound, I/O-bound or connection-bound?

level: middleimportance: must knowfreq 60%

basics

~10 s

Time one unit of work and compare CPU time to elapsed time: time.process_time over time.perf_counter near 100% means CPU-bound, near zero means it is waiting. Thousands of simultaneous idle waits means connection-bound.

open as a page

A search-index rebuilder moved from threads to ProcessPoolExecutor and got slower while memory climbed — which costs of the process model explain that?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Processes pay for interpreter startup, per-worker memory and pickling every argument and result through a pipe. Large payloads and per-worker copies of module-level caches can cost more than the parallelism gains, and unconsumed results pile up in memory.

open as a page

How do you decide between asyncio and a thread pool for 50,000 idle connections, and what does the choice cost the codebase?

level: principalimportance: should knowfreq 38%

basics

~20 s

At that concurrency the deciding number is cost per waiter: an operating-system thread reserves a large stack, an asyncio task is a small object. Asyncio wins on capacity, but only if nothing on the path blocks.

open as a page

When would you deploy the free-threaded CPython build instead of a process pool?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

Use it when threads must share a large in-memory structure that processes would copy or rebuild per worker, and every extension on the path supports the build. Python 3.14 supports it officially, at a 5-10% single-threaded cost.

open as a page