What is CPython's GIL, and what does it do to CPU-bound threads?
answer
- One interpreter, one runnable thread
- A mutex you never see from Python
- Waiting overlaps, computing does not
- About a 5 ms switch interval
- Processes or compiled code for CPU work
basics
~20 sThe GIL is a single interpreter-wide mutex that only one thread at a time may hold to execute Python bytecode. CPU-bound threads therefore take turns instead of running in parallel, so extra threads add no speedup.
solid answer
~40 sCPython's Global Interpreter Lock is one mutex per interpreter, and a thread must hold it to execute bytecode. Because only one thread holds it at a time, four threads running the same pure-Python numeric loop finish in roughly the time four sequential runs would take, sometimes a little worse because of the cost of handing the lock back and forth. The interpreter forces those handoffs: a running thread is asked to drop the GIL after about 5 ms, the value reported by `sys.getswitchinterval()`. Threads still pay off for I/O-bound work, because a thread blocked on a socket, a file read or `time.sleep()` releases the GIL for the whole wait. For CPU-bound Python the real answers are separate processes, pushing the hot loop into compiled code that releases the lock, or the free-threaded build.
code
python · 25 linesimport threading
import time
def spin(n):
total = 0
for i in range(n):
total += i
return total
N = 5_000_000
start = time.perf_counter()
spin(N)
spin(N)
serial = time.perf_counter() - start
start = time.perf_counter()
workers = [threading.Thread(target=spin, args=(N,)) for _ in range(2)]
for w in workers:
w.start()
for w in workers:
w.join()
threaded = time.perf_counter() - start
print(f"serial={serial:.2f}s threaded={threaded:.2f}s")go deeper
Be ready to state the definition in one sentence and give the consequence: one thread executes bytecode at a time, so CPU-bound threads take turns while I/O-bound threads genuinely overlap their waiting.
Explain the mechanics: the switch interval that forces handoffs, the fact that blocking calls release the lock before they block, and why processes rather than threads are the fix for pure-Python compute.
Show that you check whether a workload is actually CPU-bound before reaching for processes, and that you can measure the difference rather than assume it. Know the cost of serialising data across a process boundary.
Own the framing that the GIL is an interpreter property, not a language one, and that the choice between threads, processes, multiple interpreters and the free-threaded build is a deployment and dependency-compatibility decision as much as a performance one.
## The one-sentence definition The Global Interpreter Lock is a mutex inside the CPython interpreter that a thread must hold in order to execute Python bytecode. It is not a lock you create, acquire or see from Python code; it is part of the interpreter's own machinery, and every OS thread that wants to run your `.py` code competes for it. The consequence follows directly from the definition: **at any instant, one interpreter runs bytecode on exactly one thread.** Your process may have eight OS threads and eight idle cores; if all eight threads are executing Python-level loops, seven of them are parked waiting for the lock. ## What that means for CPU-bound work CPU-bound means the thread is doing arithmetic, string building, object churn — work expressed as bytecode with no waiting in it. Splitting such work across `threading.Thread` workers gives you concurrency (the work is interleaved) but not parallelism (it is never simultaneous). Total wall-clock time stays about the same as running the pieces one after another, and it is common to measure it as slightly *worse*, because you have added: * the cost of releasing and reacquiring the lock at every switch, * CPU-cache disruption as work migrates between cores, * and occasional latency spikes when a thread that wants the lock has to wait out another thread's turn. ## What that means for I/O-bound work The picture inverts the moment a thread stops needing the interpreter. When a thread calls into a blocking operation implemented in C — reading a socket, waiting on a file, `time.sleep()`, waiting on a subprocess — CPython releases the GIL before it blocks and reacquires it after. During that window other threads run bytecode freely. Ten threads each waiting one second on a remote service finish in about one second, not ten. **Threads overlap waiting; they do not overlap computing.** That single sentence explains almost every real observation people have about threaded Python. ## How the turn-taking works A thread does not hold the lock until it finishes. The interpreter's evaluation loop periodically checks a request flag: when another thread has been waiting longer than the *switch interval* — 5 ms by default, readable with `sys.getswitchinterval()` — the running thread drops the GIL at the next safe point between bytecode instructions and lets a waiter in. That safe point matters: a single bytecode instruction that happens to be very expensive (a huge integer multiplication, a large regular-expression match) runs to completion while holding the lock, so the interval is a request, not a hard preemption deadline. ## It is CPython, not Python The GIL is an implementation choice of CPython, the reference interpreter almost everyone runs, not a rule of the language. Other implementations have made different choices. Saying "Python cannot do parallelism" is therefore imprecise on two counts: it names the language when you mean one interpreter, and it ignores every way CPython does achieve real parallelism. ## The ways out, in the order you should consider them 1. **Is the work actually CPU-bound?** Most services are not. If your threads spend their lives on network and disk, the GIL is already not your bottleneck and adding processes only adds cost. 2. **Separate processes.** `concurrent.futures.ProcessPoolExecutor` or `multiprocessing` give each worker its own interpreter, its own GIL and a genuine core. The price is that arguments and results are serialised between processes, so this wins when each work item is chunky and the data crossing the boundary is small. Note that on 3.14 the default start method on Unix other than macOS is now `forkserver`; macOS and Windows use `spawn`. 3. **Push the hot loop into compiled code that releases the lock.** A compiled extension that wraps a long computation in the release/reacquire sequence lets other Python threads run while it works, which is why some numeric workloads *do* scale across threads. 4. **Multiple interpreters in one process.** Since 3.12 each interpreter has its own GIL (PEP 684), and 3.14 exposes them from the stdlib as `concurrent.interpreters` plus an interpreter-backed executor (PEP 734). 5. **The free-threaded build**, officially supported as of 3.14 (PEP 779), which runs without the GIL at a modest single-threaded cost. ## The trap to avoid in an interview Do not overclaim in either direction. "Threads are useless in Python" is wrong — the entire ecosystem of threaded network clients depends on them. "The GIL makes my code thread-safe" is also wrong: the lock is dropped between bytecode instructions, so an operation you think of as one step can still be interrupted halfway. The accurate framing is narrow and easy to say: one lock, one thread running bytecode, released whenever a thread waits.
- Why is threaded CPU-bound code sometimes measurably slower than the same work run serially?You have added cost without adding capacity. Every switch interval the running thread must release the GIL and a waiter must wake and reacquire it, which costs syscalls and cache locality as the work bounces between cores. Contention also produces a convoy effect, where a thread that releases the lock and immediately wants it back keeps losing to a compute thread. None of that buys parallelism, so the overhead shows up as a net loss.
- How do you get genuine CPU parallelism out of CPython today?Four options. Run separate processes, each with its own interpreter and GIL, via `concurrent.futures.ProcessPoolExecutor` or `multiprocessing`. Move the hot loop into compiled code that releases the GIL around the computation. Use multiple interpreters in one process, which have had their own GIL since 3.12 and are exposed as `concurrent.interpreters` in 3.14. Or run the free-threaded build, officially supported in 3.14. Choose by how much data crosses the boundary and how well your dependencies cooperate.
- Is the GIL part of the Python language definition?No. It is an implementation detail of CPython, the reference interpreter. The language specification says nothing about a global lock, and other implementations have made different choices. This matters in practice because it means the GIL is something a build or an interpreter can change, and CPython itself now ships a build without it, rather than something inherent to the code you write.
It is a single microphone in a conference room: everyone can be present and everyone can be waiting on a phone call, but only the person holding the microphone is actually speaking, and a timer makes them pass it along.
saying these in an interview costs you the question
- Says the GIL is part of the Python language specification
- Claims threads are useless in Python for every workload
- Believes the GIL makes all Python operations thread-safe
- Thinks adding threads always adds CPU throughput
- Describes the GIL as a lock you acquire from Python code
- Treats threading and multiprocessing as interchangeable