skip to content

Which operations cause a CPython thread to release the GIL?

level: seniorimportance: should knowfreq 52%

answer

  1. Three ways out, one way in
  2. Waiting on the OS drops it
  3. Compiled code may choose to drop it
  4. The switch interval forces a handoff
  5. Pure-Python loops never volunteer

basics

~20 s

Three things: blocking calls in the standard library release it before they wait on the OS, compiled extension code can release it around a long computation, and the evaluation loop forces a handoff when the switch interval expires. Pure-Python computation never releases it voluntarily.

solid answer

~50 s

A thread gives up the GIL in three situations. First, when it enters a blocking operation implemented in C — a socket or file read, `time.sleep()`, waiting on a subprocess, waiting to acquire a `threading.Lock` — CPython releases the lock before blocking and reacquires it afterwards, which is why threads overlap I/O well. Second, compiled extension code may release it around a long computation that touches no Python objects; the standard library does this in places such as `zlib` and `hashlib`, and third-party compiled code may or may not. Third, the evaluation loop drops it at the next instruction boundary once another thread has waited out the switch interval, about 5 ms. Everything else holds it: a pure-Python loop, and any C call that neglects to release it, blocks every other thread for its full duration.

code

python · 28 lines
python
import os
import threading
import time
import zlib

payload = os.urandom(4_000_000)

def compress():
    zlib.compress(payload, 6)

def spin():
    total = 0
    for i in range(3_000_000):
        total += i

def timed(fn, n):
    start = time.perf_counter()
    workers = [threading.Thread(target=fn) for _ in range(n)]
    for w in workers:
        w.start()
    for w in workers:
        w.join()
    return time.perf_counter() - start

for fn in (compress, spin):
    one = timed(fn, 1)
    four = timed(fn, 4)
    print(f"{fn.__name__}: 1 thread {one:.2f}s, 4 threads {four:.2f}s")

go deeper

for a junior

Recall the simplest version: calls that wait on the network, on disk or on a sleep let other threads run, while ordinary Python computation does not.

for a middle

Be able to list the three release points and say why the switch interval exists at all, and explain why an expensive single operation can hold the lock past that interval.

for a senior

Show the diagnostic method: scale-test each stage at one and four workers, classify call sites by whether they wait or compute, sample thread stacks, and use a process pool as a control before committing to a redesign.

for a principal

Own the architectural consequence — that a mixed pipeline usually wants its waiting stages on threads and its computing stages elsewhere — and be able to justify the added serialisation and operational cost of that split against a latency budget.

## The three release points Knowing *when* the lock changes hands is what turns the GIL from trivia into a diagnostic tool. There are exactly three ways a thread stops holding it. **1. Blocking calls in C.** When a thread is about to wait on the operating system, the interpreter releases the GIL, performs the blocking call, then reacquires it. This covers essentially all of the waiting a program does: socket reads and writes, file reads and writes, `time.sleep()`, waiting on a subprocess, waiting on a `threading.Lock` or `threading.Event`, and the select-style calls underneath them. This is the mechanism behind the whole "threads are fine for I/O" story — the wait happens outside the lock. **2. Compiled code that chooses to release it.** An extension module written in C can drop the GIL around a stretch of work that touches no Python objects, then take it back before returning. The standard library does this for genuinely long computations — compressing a large buffer with `zlib`, hashing a large buffer with `hashlib` — and a compiled numeric library will typically do it around a heavy kernel. Crucially it is a *choice*: an extension that forgets, or that cannot because it manipulates Python objects throughout, holds the lock for the entire call and stalls every other thread. **3. The switch interval.** While executing bytecode, a thread periodically checks whether another thread has requested the lock. Once a waiter has been waiting longer than the switch interval — 5 ms by default — the running thread drops the GIL at the next instruction boundary. This is what stops one compute thread from starving everything else, and it is why threaded CPU-bound code makes progress on all threads rather than none. Everything outside those three is a hold. A pure-Python loop holds the lock except for the forced handoffs. A single expensive bytecode instruction holds it for its whole duration, because the check only happens between instructions. ## Diagnosing it in practice Consider a genome-annotation pipeline whose per-record path runs several stages: parse a record, look up features over the network, score them, then serialise the result. It runs on a `concurrent.futures.ThreadPoolExecutor`, and the team has a 92nd-percentile latency budget it keeps missing. Someone raises the worker count from eight to twenty-four and throughput does not move. The question to ask is which stage holds the lock. A structured way to find out: * **Scale-test one stage at a time.** Run a stage alone with one worker, then with four. A stage that speeds up releases the lock; a stage whose wall-clock time stays flat holds it. This is the cheapest and most conclusive test, and it needs no tooling. * **Classify each call site.** Network and disk waits release. Compression and hashing in the standard library release. Pure-Python parsing, scoring and object construction do not. A compiled dependency is the interesting case: assume nothing and measure it. * **Sample where threads actually are.** `sys._current_frames()` gives you every thread's current frame, and `faulthandler.dump_traceback_later()` with `repeat=True` will dump all thread stacks on a timer. If the same pure-Python frame appears across samples while others wait, you have found the holder. * **Use processes as a control.** Running the suspect stage under a `concurrent.futures.ProcessPoolExecutor` at the same worker count tells you immediately how much of the ceiling was the lock. It is a diagnostic even when it is not the shipping answer. The usual outcome is a split pipeline: the I/O-heavy stages stay on threads, where they overlap perfectly, and the CPU-heavy scoring stage moves to processes or into compiled code that releases the lock. Tail latency in particular tends to improve more than mean throughput, because the tail was dominated by requests queued behind a compute thread's turn. ## The subtleties worth stating Releasing the lock is not free and not instantaneous. A thread coming back from a blocking call must queue for the GIL behind whoever holds it, so an I/O thread can be delayed by up to a full switch interval by a compute thread — the classic convoy effect, and a real source of tail latency in mixed workloads. And a release point is not a preemption point in the middle of an operation: a single long-running C call or an expensive bytecode instruction runs to completion regardless. One more caution when moving compute off threads: if the stage does floating-point reduction, changing how the work is partitioned changes the order of additions, and floating-point addition is not associative. A small rounding drift in results after a threads-to-processes migration is usually chunking, not a bug in the new code. ## The compact answer Blocking calls release it, compiled code may release it, the switch interval forces it, and pure Python holds it. Every threading decision you make in CPython follows from that list.

  • A compiled dependency is on the hot path. How do you find out whether it releases the GIL?
    Measure rather than read documentation. Run the call alone in one thread, then the same call in four threads, and compare wall-clock time; if four threads take roughly as long as one, the call releases the lock, and if they take four times as long, it holds it. Sampling thread stacks while the test runs confirms where the other threads are parked. Do this per call, because a library can release the lock in one function and hold it in another.
  • Why can an I/O thread still see latency spikes even though its waits release the GIL?
    Because returning from a blocking call means queueing for the lock again. If a CPU-bound thread currently holds it, the I/O thread waits until that thread reaches a release point, which can be up to a full switch interval, and under contention it can lose repeatedly to compute threads. This convoy effect shows up as tail latency rather than mean latency, which is why a percentile budget catches it and an average does not.
  • Does a thread hold the GIL while it waits to acquire a `threading.Lock`?
    No. Blocking on a `threading.Lock` releases the GIL, otherwise a thread waiting for application-level synchronisation would freeze the whole interpreter and the lock could never be released by its holder. The same applies to `threading.Event.wait()`, `threading.Condition.wait()` and joining a thread. Only the busy-waiting you write yourself in Python keeps the interpreter lock occupied.

saying these in an interview costs you the question

  • Thinks every C call releases the GIL automatically
  • Believes a thread holds the lock for its entire lifetime
  • Says pure-Python loops release the lock on function calls
  • Cannot name blocking I/O as a release point
  • Assumes threads waiting on a lock still block the interpreter
  • Concludes threads are broken without measuring a single stage

context